Becky Allen is one of the most thought-provoking and thoughtful education researchers. Her substack ‘Falsifiable’ explores whether learning can become a science. We discussed why it’s been so hard to achieve a science of learning, why there is no orthodoxy for learning science to argue against and whether attempts to automate teaching with AI could help us build the evidence base we need.
Becky is Professor of Education at the University of Brighton and co-founder and Chief Analyst at Teacher Tapp. She is also an advisor at Alpha School and previously worked as Director of FFT Education Datalab and Professor at the UCL Institute of Education. She is co-author of The Teacher Gap (with Sam Sims) and The Next Big Thing in School Improvement (with Matthew Evans and Ben White).
We edited the transcript lightly for clarity.
Why education is so hard
Phil: Becky, you’re a Professor of Education, co-founder of Teacher Tapp, previously head of FFT Education Datalab, and co-author of The Teacher Gap and The Next Big Thing in School Improvement. The second of those really inspired me in terms of thinking about school leadership and supporting schools. I think you’re one of the most important thinkers on building a science of learning, which is why I wanted to talk to you today. Thanks so much for being here, and welcome to Chalk Talk.
Becky: Thank you, it’s a pleasure to be here.
Phil: Primarily what I wanted to talk about is your recent blog posts on your Substack, Falsifiable, about how to develop a field of learning science. Through your recent writing and your work on teacher development and assessment, it seems to me that we often, in the education community, don’t fully appreciate how hard it is to design learning, to assess it, or to lead schools. Why is education so hard?
Becky: I’m mostly not going to talk about all of education today, but specifically about instruction and learning, because I think they’re different things. And I think it’s so hard because the things we actually care about are somewhat invisible, really slow and really contested. So we’re working in an environment with very poor feedback loops.
It’s really hard for us — whether we’re instructional designers online or teachers in the classroom — to ever know whether the decisions we’re making are good, productive decisions or not.
I sometimes use that phrase, “learning is invisible”. I don’t think that’s quite right, and I don’t think it’s always a useful way to think about it. It’s more that many different invisible things can cast the same visible shadow. We can observe things, but it’s really hard to extract the meaning from what we’re observing. When we see, for example, that a child successfully answers the question “seven eights are 56”, there are all kinds of things that could have gone on in their head that led them to give that answer. And if we’re a maths teacher or a primary teacher, it matters quite a lot that we know which of those things led to the response.
These are the reasons why education is so hard. In my blog posts I sometimes write that education is a diet problem, not a Sudoku problem. What I mean by that is that it’s very easy to propose instruction. It’s very easy to write a lesson plan — and with LLMs, it’s now very easy to write a hundred different lesson plans. The hard part is verifying whether the lesson plan has worked or not. That’s the same as the diet problem: we can all think of diets we could follow, and there are hundreds of them out there, but we know very little about which of them work, because when we put diets out into the wild, actually observing their effectiveness — or the mechanisms that determine whether they work — is really difficult to do. So those are some of the things we’re battling against.
Phil: Some of your phrases have so much meaning packed into them. “Learning can be invisible but it casts a visible shadow” is a really productive phrase. I hadn’t heard the diet-and-Sudoku comparison before, but I have read you describing education as more comparable to nutrition than pharmacology.
Becky: Pharmacology — drug research. And I think that was the mistake when we set up the Education Endowment Foundation: we drew on people who had written a lot about drug studies and said, we can do this, we can just test whether the widgets work or not. We failed to recognise at the time the extent to which we were nothing like pharmacology. We were operating in an environment where we’re trying to learn about good nutrition. And actually, the world of scientific research on nutrition is in the same complete, messy, terrible state that the world of education is in, because it suffers from exactly the same problems we suffer from in trying to work out what works.
Phil: Dylan Wiliam famously gave a talk at researchED in 2014 where he said teaching will never be a research-based profession. Obviously his argument was more nuanced than that — he described how research can inform but not determine practice. Do you think that’s still largely right, or have things changed since then?
Becky: Firstly, I think the classroom — at least in England — has become more research-informed. I also think that most of the decisions teachers have to make about how to teach aren’t research-informed, and they can’t be, because the research operates with a set of principles that isn’t going to tell them how to answer specific questions about how to teach particular things.
Do I think he’s right that it will never be? Well, we have to decide what we mean by teaching. If we mean teaching with humans in a classroom, that’s necessarily a very complex endeavour, because we’re having to translate the learning science through a teacher who has a personality, and hopes and desires for themselves, which then inform how they want to teach. And then we have to teach thirty children at once, with very poor feedback loops in the classroom — necessarily poor.
Distinguish that from online learning situations, where we can control the instruction very precisely. We can deliver it to thousands of students in exactly the same way, vary one element only, and look at the impact of varying one element at a time, trying to work out whether the variations we’re putting in place are material. Now, I don’t want to understate how difficult it’s going to be to do good learning science research online. It’s still incredibly difficult, because we still have all the same problems we have in the classroom: complex outcomes that we’re trying to target, very long time horizons for working out whether instruction has been effective or not. But we do at least have a model where, in theory, there’s some prospect of moving forward and working out what effective instruction looks like.
To me the interesting question is: what is it that we learn? Do we learn new, high-level theoretical ideas? Or do we actually learn something else — that we don’t necessarily build grand theories, we build worked solutions to how to teach specific things, or specific types of things, in specific situations? That’s a much slower, less satisfying process. But it is the reality of how teachers learn to teach. They’re not largely relying on grand theories. They’re relying on individual knowledge of how particular concepts are best taught.
Does learning science need an orthodoxy?
Phil: I wonder if there’s an analogy here. There are many examples in the history of educational debates over whether instruction should be more top-down or bottom-up — whether you generate examples and develop schema that way, or support students to develop a schema first and then think about the examples. And similarly with machine learning: LLMs are basically pretty bottom-up, they learn from examples, whereas old-fashioned AI techniques like scripts and schemas were more top-down. So it feels like a debate that’s recurred over time.
I really want to get to your idea about how digital platforms and AI could potentially solve this verification problem. But first I wanted to ask whether there’s, to some degree, an academic solution. Economics has its own rigorous sub-discipline, econometrics; medicine has biostatistics. These are sub-disciplines that are bespoke to the way data is generated in those domains. Does education have an equivalently bespoke set of methodologies — a whole sub-discipline — yet? And would it be helpful to have one?
Becky: Let’s be specific about what those other disciplines do and don’t have. Since I’m an economist, let’s use that example — you used the word econometrics, not economics, but the two go hand in hand. Economics has a set of theories, and the only reason econometrics can exist as a discipline — which is essentially statistics as applied to economics — is because it has a set of theories that it’s willing to test. The goal is that you have a theoretical foundation, and then a means to challenge and develop that foundation. In the case of social studies like economics, the means to challenge and develop is with reference to empiricism, the real world, because that’s the target: explaining the real world.
So what are we missing in education — or in learning and instruction? Let’s throw away the rest of education and just talk about learning and instruction. Do we have a theory? We have something weird. We have layers of theories. We have people observing instruction and learning at different layers: the layer of neuroscience, of cognitive psychology, of behavioural psychology, of educational psychology — which often looks quite different — curriculum theorists, assessment theorists.
What’s the issue with that? Well, you’re developing theories that find it really hard to speak to each other, and they do that for two reasons. One is just that they have different language, and in a sense that’s potentially resolvable: you can learn to talk to each other and develop a common language. But something else is going on, I think, which is that they perceive things within the learning process at different layers and different levels of detail. And when you talk about things at different levels of detail, it can be quite hard to reconcile theories that are just talking at completely different layers of meaning.
I often think that even among teachers, what they argue about is often the layer of meaning they’re talking at. Think of teachers who have incredibly bottom-up ways of perceiving something like teaching science. They perceive scientific knowledge as atomised knowledge: you teach each part, and then you teach the way the parts connect to each other. They have a whole language that talks about knowledge at that level. Compare that to science teachers who perceive at the level of grand ideas and grand meaning — that’s the thing they’re constantly using to ask, does a student get it or not?
Out in the world of education I often see conversations that can’t align because people are talking at different layers of meaning — whether it’s the difference between talking about schema theory and talking about neuroscience, or the difference between talking about perceiving and understanding versus atomised knowledge. I see these as intrinsic problems of meaning that are very hard to resolve. And if we don’t resolve them, we can’t develop a unified theoretical idea of what we’re trying to understand, what we’re trying to falsify, and what we’re trying to develop.
One of the nice things about economics — a discipline that has lots and lots of problems — is that it does have an orthodoxy at its centre. Everybody has to learn and fully understand orthodox economic thinking before they’re allowed to become heterodox and have mad ideas about how we should think about economic theory-making. So everything happens with reference to an agreed orthodoxy. We don’t have that.
Maybe this is a very long way round to getting back to the idea of empiricism — how do we learn? To my mind, we can’t learn anything as a discipline unless we know what the thing is that we’re challenging. That’s why we get very stuck. We can run individual RCTs, we can do better psychometrics and better measurement. But if we’re only thinking about it one research question at a time, it becomes widget-testing. We can say: does the widget work? But it never helps us develop our own theoretical ideas of what anything means in the world of learning and instruction.
So that’s the bit I’m hesitant to say we’re going to make progress on. I genuinely think we’re going to make progress on working out the widgets — how to teach canonical ideas and lessons really well to students. I think we can throw away some ineffective ideas and establish some ways that work better than others. But I’m more cautious about the extent to which we’ll manage to connect that back to theory-building, which is the thing that I personally think we need.
Regularities all the way down
Phil: Roughly, what do you think the layers should be? In your piece “Regularities All the Way Down”, you talk about — if I’m understanding rightly — the bottom layer being cognitive science, where cognitive load theory is perhaps the dominant aspect. Then there’s another layer which, depending on your persuasion, might be Engelmann and Carnine’s direct instruction. And then there’s pedagogical content knowledge. Is that roughly what you think the layers are? I know the exact contents might differ.
Becky: Whenever I’m thinking about these things, I’m always asking: am I working from the point of trying to get from knowledge to making decisions about instruction? Or am I trying to say something about cognition — the part that goes from the instruction into the child’s brain? I know it’s all one process, but at the two ends of it you have things that are quite different in their structure.
When I wrote “Regularities All the Way Down”, I was trying to describe the idea that there are principles of how to teach. Some of them are unifying principles that are always true, and cognitive load theory would be an example. Then there are some principles that only apply to particular knowledge types. This was Engelmann’s insight: you work out what type of knowledge it is, and once you know the type, it tells you something about how you should teach it. And then I say that there’s a layer below that, which is the thing itself, the idea itself. We call that pedagogical content knowledge. What do we think it is? It’s the knowledge of how to teach a particular thing to a child or a student. It’s knowledge of what kinds of misconceptions tend to arise, and what kinds of examples tend to be most productive when we try to teach a concept.
The weird thing is: why should that exist at all? That’s what I’ve been writing about more recently. What is it? Is it just idiosyncratic ideas, and all we can ever do is draw up the list of them and learn them over time as a teacher — or ask whether the LLM has decent ideas about them? In the last blog post I wrote, I was positing that a lot of these idiosyncratic rules of thumb about how to teach individual concepts arise because we’ve got higher-level principles we’re trying to apply, but they conflict with each other.
In the post I gave the example, for primary teachers, of first introducing negative numbers. Negative numbers are a highly abstract concept that for most of human history we’ve managed to live without, and they don’t really have a great real-world meaning. Teachers disagree wildly about the best ideas to use to teach them: submarines, diving boards, debt and borrowing money, working only on the abstract number line, two-sided counters that you turn over to represent different things, or an embodied experience of children walking up and down a number line — walking backwards and forwards, turning around to represent the arithmetic operators.
The problem is that you’re trading off immediate comprehensibility — can the student in the moment make sense of what you’re saying? — against whether the example is going to be generalisable in the long run. Stage one is just: does the student have some sense in their head of what this abstract number line looks like? To get there in the quickest way, you use the most comprehensible example, which is usually the submarine or the diving board. But those examples aren’t useful to draw analogies against when you start using arithmetic operations. So you run out of headroom. If the students have developed a particular way of visualising negative numbers that runs out of usefulness very quickly, then you have a problem as a teacher. It was optimal in the moment, but in six months’ time it could be worse than a more difficult example — one you can keep returning to and extending as the topic develops.
So I was discussing the idea that these ideas about how we teach are universal — it’s just that the principles conflict all the time, and that’s why teaching is so tough.
The limits of cognitive load theory
Phil: It’s such an interesting point about people having slightly different conversations because they’re thinking about different layers. This might be a slightly abstract question, but cognitive load theory is a dominant theoretical foundation for a lot of the way teachers think about teaching, and I think it’s really important. For listeners: cognitive load theory splits cognitive load into intrinsic, extraneous and germane load. But I was reading an interview with John Sweller where he said there was a bit of a problem, because it isn’t really possible to measure intrinsic, extraneous and germane cognitive load. People have used eye-tracking and things like that, but it’s really hard to measure. He described all these research papers coming out that explained their experimental results — students learning more, or some other effect — in terms of cognitive load theory, and he said that it rendered the theory unfalsifiable.
Is that one of the core issues with learning science? We talked earlier about pharmacology and physiology — those theories enable you to create hypotheses that are falsifiable. Is that one of the issues with turning learning into a science?
Becky: Cognitive load theory isn’t really connected to any neuroscientific mechanism or explanation of what’s taking place. And it’s not their fault: the theory was developed quite a long time ago, when neuroscience — educational neuroscience — wasn’t really in a place where it had much to say about these ideas. I think neuroscience has moved forward a lot, but it still isn’t in a place where we can experimentally measure levels of different types of cognitive load. That’s why I think the theory in itself has limits.
It tells us as teachers one really useful thing: don’t put unnecessary load on students when they’re learning complex ideas — get rid of what you can. That’s super easy. And it has another general insight: worry about how many ideas you’re trying to load onto a child and ask them to hold in their head at one time. But it doesn’t say much about what to do about the trade-offs. Because there are trade-offs. There are times when we do want to load lots of ideas and ask a student to reason about multiple things at once, because it’s the only way we can teach more complex reasoning. And that’s where, for teachers, I think it runs out of headway in terms of what it’s able to do.
You say it’s a dominant way for teachers to think. I think it’s a dominant way for some teachers to think. I think it’s a theory that’s really useful for some subjects, in particular maths. Dylan Wiliam was a maths teacher, and that’s why to him it was incredible to learn about this theory. But if you’re an English or history teacher, I don’t think it really helps you with any of the most complex decisions about how to teach in your classroom, because that’s not the level at which the complexity of instructional decisions is being made. Whereas in maths, it does send very clear signals about being quite controlled in the ways you ask students to think and reason at any particular point in time, and how you can structure ideas. What was the question?
Phil: Is one of the issues with learning science that we don’t have a unifying theory that can generate falsifiable hypotheses? I don’t know if it’s correct to call pharmacology or physiology unifying theories, but is that one of the problems?
Becky: The question is: should there be, and will there be, unifying theories? I spend a lot of time talking to people who think about these ideas, and at the moment I’m at the end of the scale that believes there are not grand unifying theories of instruction. The reason I think that is that we are fundamentally trying to take knowledge — a complex domain, structured in quite idiosyncratic ways across different parts of the knowledge domain, which we call the different subjects — and translate and transfer it somehow into a brain, into cognition, which we are certain is complex. It’s not just that we don’t know all kinds of things about cognition. It’s that when you’re doing this translation between complex domains with a lot of variation in their texture, into a complex environment, I suspect there aren’t grand unifying theories. Every time you try to think of one, you can always think of a situation where it doesn’t apply, or where it would be bad advice.
But I’m willing to be proven wrong, and that’s why my blog is called Falsifiable. I’m always trying to ask myself: how can we work out that this is a bad idea? At the moment that’s where I sit. That might be quite unsatisfying, but it doesn’t mean we can’t get better at instruction. It also doesn’t mean we can’t learn a huge amount. We can learn a huge amount about productive ways of talking about knowledge, and about productive ways of talking about cognition. It’s just that whenever we try to put the two together, what we end up with is a little bit messy.
What AI can and can’t do about verification
Phil: You’ve written extensively, and in quite a nuanced way, about how AI could potentially help learning become more scientific — help us understand what optimal instruction looks like at a particular scale. Could you speak a bit about that?
Becky: In the short run, what LLMs know right now is what typical human knowledge about good instruction looks like. I say “good instruction” — really I mean written-down and accepted instruction, which I think is better than the average instruction of teachers, because I think good teachers write stuff down. So what we can do very quickly is this: if we try to develop a lesson on how to teach a single idea, we have the means, via an LLM, to develop it, and the chances are that what’s developed is something an average teacher thinks is okay, actually.
The question is how we get better than that. We get better than that in part by developing multiple different ways of teaching things according to different principles, and trying to test which types of learning are effective. Online platforms allow us to do that in a way we were never going to manage in the classroom, because we just can’t run trials in the way you need to run trials. It’s really, really expensive to run trials, and you’ve got to get randomisation working at the level of the individual student, so that every student is experiencing something different. If you can’t do that, and you have to randomise at the level of the school, you’re having to recruit thousands of children just to find out whether a tiny widget works. That’s a non-starter. With online platforms, not only can you randomise what individual children see, you can randomise what a single student sees across different topics and lessons, and look at the variation in how they respond to different types of experience. That’s the positive side.
The negative side is that if we leave engineers to do this, we will end up in an almighty mess, because they will optimise on the things they naively think are sensible proxies for learning — which includes the end-of-topic test. There are lots of situations where, if we end up optimising on end-of-topic tests, we don’t allow students to develop useful and generalisable mental models within subjects — the ones that allow them to become good mathematicians, good economists, good historians. So this only works if people who think really hard about what the engineers would call verification, and what we would just call assessment, are involved in getting it right.
How far can AI go in helping us? I’m pretty sceptical about some other things I see going on at the moment, and I’m pretty clear that I’m right to be sceptical. One thing I see quite a bit is people going on X and saying they’ve used an AI to simulate a third-grader — they’ve run an AI simulation where a thousand third-grade students have already taken their course, and they’ve identified what’s right and wrong about it. Now, there’s a nice body of literature that’s grown up in the past couple of years about this, and it’s absolutely unambiguous that LLMs cannot simulate students. Once you understand what LLMs are and how their knowledge is structured, you can understand why: you can’t take knowledge away. You can’t take a really smart LLM, take knowledge away from it, and say, “pretend you don’t know those things”, or “pretend you can act as a third-grader”. They suffer from all the same problems that, frankly, humans suffer from when we try to imagine we’re a third-grader, look at a lesson plan, and ask when the student is going to get stuck. I’ll come back to that in a moment, because I’m going to write something about it soon — I think teachers do something that’s quite unique.
What LLMs can do is this: if you ask them to simulate being an expert teacher — look at this lesson as an expert teacher and tell me when the students are likely to get stuck — they’re able to do that in an okay way. So the people who say “just solve the verification problem with AI” are definitely wrong. We don’t solve the verification problem. We just develop a platform that potentially gives us a route to get through the verification problem in a reasonable way.
Coming back to simulating students. One of the interesting things about teachers is: when they look at a lesson plan and decide whether they like it or not, and give feedback, what are they doing? What is their taste? Is it a set of principles that in principle we could codify, so that an LLM could just copy what they do? Or is it something that in the literature is typically called tacit knowledge — stuff we can’t codify into principles, but that is a whole set of knowledge about what works and doesn’t work, built up over time and hard to codify? It probably is. But if it is, there’s an element of it that’s correlated with effective learning, and an element that’s just aesthetic taste. Teachers like to teach in particular ways, and that part isn’t interesting to us if we’re trying to build online platforms.
The literature is quite troubling in general, because it talks about all the problems: if you put a single lesson plan in front of a hundred teachers, they will all have different views on what they think of it. That’s troubling if you’re trying to build an online platform and use teachers as the human-in-the-loop to improve the initial versions of the lesson plans. But the one thing it does say is that teachers seem to develop a skill that normal humans can’t do: they put themselves in the place of the novice. When they look at a lesson plan, they’re able to simulate being a student. They seem to be uniquely good at this, because that’s really what the job is — the job of being an effective teacher is to imagine being a student. And normal humans are pretty bad at doing this, because they don’t have a life where they need to simulate being a student.
These are all things I’m mulling over. When I’m thinking about how we build systems that produce better instruction, what the role of humans is as experts, and what the role of AI is — it matters quite a lot that we think hard about who the humans are, and whether they’re giving us feedback that’s likely to be correlated with improving the instruction from the perspective of students learning more.
Phil: That’s a really interesting point. I hadn’t thought about it in terms of teachers developing that skill of simulation. I suppose in a lot of other professions you develop equivalent simulation skills — a product manager might develop the skill to simulate the user, a salesperson the skill to simulate the buyer — and that’s correlated with being successful.
I hadn’t actually seen people publicly talking about simulating students. I did see a research paper — and I’m not an expert in this — on simulating marketing campaigns, or something like that, where they had a labelled dataset that apparently showed LLMs were quite close to accurate. I’m not saying that has any bearing on whether LLMs can simulate students, because there’s obviously a difference in terms of training data — there’s probably loads of information on the internet about that sort of thing.
Becky: To be clear, there are exactly equivalent papers in the education space that say LLMs are almost as good as using the underlying attribute data and a statistical model to predict whether something’s going to be successful or not. But in that case, why not just use the statistical model? That’s not the basis on which we judge whether something is good. We need to be better than the tools we already have. And you end up in an internal problem where you’re verifying yourself with the data you already have, so you can never get better than it.
Which subjects are safest to optimise?
Phil: Going back to using digital platforms in conjunction with learning designers to help with the verification problem. A very strong version of it would be that you can apply some sort of sampling theory to education, where you have a concept space, and if you’re sampling the concept space to cover the “understanding boundary”, let’s call it, then the student will understand the concept — and if you’re not sampling around that boundary, they may not. But for that to work — and you’ve written about this in your article on knowledge graphs versus curriculum graphs — you need a very well-defined region, and you presumably also need students to have the same prior understanding. So do you think digital platforms and AI can only help in subjects with very structured curricula, like maths, or do they have wider applicability?
Becky: I think they can help any time you can define an outcome. In other words, you can write the assessment, and you can feel confident that the assessment is universal enough that if you optimise on it there won’t be unintended consequences. You have to decide, within a subject, how difficult or easy that is.
Actually, I don’t think maths is the easiest subject to do it in. I think there are lots of problems with maths and science, and the reason is that they’re two subjects where there are often short-run hacks for teaching a student to do the thing you want them to do that are quite unhelpful in the long run. Science is famous for it — it’s writ large with teaching misconceptions and folk science in order to get to the next stage, and you store up problems. Often we need to teach students quite difficult abstract concepts so that in the long run they can think scientifically and not get stuck. I think that’s true in maths as well.
Think of the most obvious example. In the national curriculum in England, and in the US, we decide that quite young children have to be taught what a half is and what a quarter is. Because we do it when they’re so young, and they don’t yet have the idea of the continuous number line, we can only teach one abstraction, which is cutting up the pizza. So we teach children to cut up the pizza, and for the rest of their life they walk through the world with this abstraction of a half where they visualise a pizza. It is the most mathematically useless representation of what a half is. If only we could not teach them that, and delay it until we’re ready to start talking about the continuous number line, and about division and sharing as the representations of fractions — the mathematically productive ones — we would get out of a lot of the troubles we have. But we do it all the time in maths. Everywhere, we teach mathematical ideas in ways that cause children to get stuck, because we try to introduce concepts too young.
I would argue that all the algebra and pseudo-algebraic stuff that goes on in primary school has, to my mind, made algebra a lot more difficult to teach than it was in my day — which was just about pre-national curriculum — when you didn’t do any algebra or pseudo-algebra at primary school. You arrived being able to actually learn what the equals sign is, how to think productively about equality, about two sides of an equation, about balancing. That’s a useful heuristic and visualisation for thinking about what algebra is. But we don’t do that.
So I’m talking myself into something here — and as I say, I change my mind about things a lot — but I would feel more confident running an optimisation programme on an AI platform in almost all subjects other than maths and science. I’d feel far happier doing it in the social studies, or perhaps English, where you don’t have all these abstract dead ends that you can create in subjects like science and maths.
Phil: That’s a really interesting point, and one I hadn’t heard before. I used to teach history, as a disclaimer, and this is slightly tangential, but I always remember that when history teachers teach the rise of Nazi Germany, they talk about proportional representation in the Weimar Republic making it structurally weak. I always think that’s a simplification. And I sometimes wonder: in an alternate universe where we didn’t teach students that, would the UK have a proportional representation system by now?
Becky: And would that be good? I’m not sure.
Phil: Exactly. Maybe useful simplifications can also have unintended benefits.
Becky: They can, yes.
Reward hacking — by students, teachers and LLMs
Phil: I find your optimism about digital platforms helping us understand instruction better at a certain layer — quite a fine-grained layer, I think, from what we’ve discussed — really compelling. But are there still risks? Based on what we’ve just discussed, it might be possible to do some sort of reward hacking, where students are hacking the wrong things to get good results. Transferring that evidence to other settings, like classrooms, might be difficult. And is teaching, as a profession, structured so that this kind of knowledge can be distributed? Those might be reasons to temper enthusiasm. Are those potential concerns as well?
Becky: To be clear, when it comes to reward hacking, I’m not worried about the students. They do that all the time. They are reward hackers; we’re already living in that world. I’m worried about the LLM reward hacking.
Take history. When we teach history, we’ve got the period of time and the facts around what happened — the story of what happened. Do they know the story? Then we’re always layering on, usually, a couple of different things. One is that we have big ideas and big themes that we keep returning to, and we keep trying to develop more sophisticated ways of thinking about them — we return again and again to monarchy and who runs the country, or democracy, and so on. And then we have historical disciplinary practice: a set of things that historians do that students learn to do. They learn a structured and increasingly sophisticated model of thinking about causation, trigger events and so on. You’re constantly building that model. You’re constantly building the model of what primary and secondary sources are like, and productive ways to think about them. You’ve got those three things.
Now say you’re asking an LLM to optimise a set of instruction, and you’ve got an endpoint exam which includes a lovely essay and some other questions, and it’s got to optimise against a rubric or a set of exemplar essays. If you leave it to do that on its own, it is going to find ways to not constantly hold in mind these big, long-term goals you have in history about developing higher-order ways of thinking and tying together ideas. When you’re teaching a new example in history — a new situation — you will spend time as a teacher choosing to say: remember when we learnt about such-and-such? Can we all talk about what we learnt and how we thought about it? Can we apply those ideas? Sort of — but they’re not quite right; let’s develop them. Perhaps we could go back and think about that idea again, and think about whether our new way of thinking could have been useful when we studied that topic. None of that shows up in the test, because the test is about the topic you’re teaching at the moment. But all of it is absolutely critical to developing the thing teachers call the schema. We don’t know what it is in the mind, but it’s the way the neurons have been connected to have productive ways of thinking about history.
That’s what you risk losing when the LLM decides to reward hack. It’s not choosing to think in that meta way — not because it can’t, it can — but it’s not going to think that way if you’re optimising against an exam where there are other routes to passing. There are more mundane routes: just making sure the topic is very, very secure, and that you have sophisticated ways of talking about that topic, but in a kind of verbatim way — let’s just make sure we can deliver the correct essay for that topic.
Phil: This might be an ignorant question, but that time-horizon issue seems like a really big and complicated one to get around. I don’t know how you’d get around the problem of thinking about how this particular instructional move, or this sequence of examples, is going to multiply across time — and also the influence of prior learning. Do you think it’s fair to say the time horizon is the biggest barrier to using AI and digital tools to learn more about learning science?
Becky: Yes. But I also think we can draw analogies with teachers in the classroom, and whether they are reward hackers themselves. That’s a controversial thing to say, but we all know teachers who are amazing at teaching to the test. And frankly, if you’ve got a child about to enter their GCSE season, as I do, you’re pretty happy when they’re reward hackers.
Take a very ill-defined discipline like English, where nobody knows what English is or what the point of the lessons is — it’s the most ill-defined domain. There are teachers who know exactly how to teach a class for the AQA or Edexcel spec, so that whatever students encounter on that question, they know exactly how to write the answer. And they’re completely comfortable with that being the primary mode of instruction for the GCSE course. Then there are other English teachers who teach in a way where that is quite opaque, and only gets discovered through a process of quite sophisticated thinking about English. Now, are they good or bad teachers? Who knows. But what we know is that they’re thinking about the goals of English instruction in slightly different ways. One has very explicit goals to optimise against — a set of examiner reports, past papers and so on — and one is optimising against their long-term notion of what it means to be good at English.
I think this varies across subjects, but we see teachers at GCSE all the time doing things, and teaching students to do things, that are really unhelpful if that student goes on to study further. In the sciences, they encourage students to use formula triangles a lot, which are enormously unhelpful if you want students to go on to study science or maths or anything that involves an equation. Why are they doing it? They’re reward hacking for the exams. And, again, frankly, as a parent — and as a student — you’re like, knock yourself out. My daughter doesn’t want to study science beyond GCSE. I want her to get through that exam, and if that means she’s using a formula triangle, so be it.
So we shouldn’t pretend it’s “teachers good, LLMs bad” — that LLMs are going to reward hack and teachers don’t. We all do it. Everyone’s doing it. But the LLMs will make the trade-offs very, very explicit. And they potentially scale things: an LLM that can teach a thousand students, a hundred thousand, a million. So you can end up with very distorted views of what a subject looks like if you leave an LLM to make those decisions.
Phil: It would be useful if we had something equivalent to life expectancy, where most people agree on the measure — although not everyone agrees; my partner’s a doctor and I know there are debates around quality of life as well.
Becky: I studied a bit of health economics, and I’ll say it’s very contested, because they tend to optimise on quality-adjusted life years, which is not life expectancy. And there’s a very strong human notion about the fairness of a “fair innings”: that we should always intervene on young people and try to save their lives, even if the costs are billions, because they haven’t had their fair innings. But if somebody reaches 70 and gets a disease, we have this very strong internal sense of justice that says, well, we’re not going to drop a billion pounds to save your life, because you’ve had a decent run, mate. That’s not consistent with quality-adjusted life years, but we do those things all the time. So health economics is very contested too, unfortunately.
What automating teaching might teach us about teaching
Phil: The grass is always greener, as they say. I had one more question. For quite a long time — maybe not that long in the grand scheme of things — there have been repeated attempts to automate teaching. B.F. Skinner introduced what he called a teaching machine, which wasn’t a business success. Khan Academy was, in some sense, an attempt to automate teaching with videos. Adaptive platforms like Knewton and others didn’t have the impact they thought they would. With the current attempts to automate parts of teaching, do you think we’ll see failures that are instructive about what teaching actually is? Might attempts to automate teaching teach us things about teaching that we didn’t know?
Becky: Maybe. Let me start by saying why I think the current attempts are different. I think they’re different because they don’t have to be algorithmic. I talk about transforming complicated knowledge into a complex mind, and if there are no grand unifying theories, that does explain why algorithmic approaches to working out how to teach fail. But we don’t have to have algorithmic approaches with LLMs, because, just like us, they have distributed minds, and they can reason about how to teach individual things. You can say to them: these are the principles of good instruction, but they’re going to conflict, and that’s okay — you just reason and figure out what you’d like to do. And they will figure it out in the same way humans will, because they have access to the same knowledge we have: the body of human knowledge about how to teach concepts. That’s why I believe they are capable at the moment of figuring out how to teach things as well as humans can. Admittedly, they have to deliver it on an online platform, and that’s another question — whether that will ultimately be productive or not.
It does raise all these questions that I think we will, over time, be able to reveal quite deeply: what teacher knowledge is, what teacher taste is, and whether it’s all stuff that’s actually useful for learning, as opposed to things that make instruction more enjoyable or more pleasant, or make sense to the teacher.
If attempts to automate teaching keep failing, one of the biggest reasons people say it’s likely to keep failing is that students themselves are so idiosyncratic — both in their prior knowledge and in their own tastes and preferences, what they’re like, how their minds tend to organise things. So people say it’s not going to work, because we have these idiosyncratic minds, and therefore when we run A/B tests to see what we can learn, we find that barely anything has an effect, because within any particular strand of students receiving a particular treatment we have heterodox responses. I would say: well, if that’s the case, then the classroom is an absolute non-starter, because you’ve got thirty minds in there — so stop kidding yourself that you’re managing to meet their needs. If that’s the case, we all face the same problems. And I’m open to the idea that it is the case.
There are all sorts of things we observe about the ways that, over time, students reorganise quite disconnected factual ideas into productive, useful and efficient ways of thinking — both in subjects like history, which are ill-defined and messy, and in subjects like maths. We can observe massive variation in how that happens. We’re pretty sure it’s not related to instruction. I think something like IQ is an insufficient explanation for what’s going on. And we can’t really work out why some students are able to make really useful transformations of knowledge and others aren’t.
One example is something that’s implicit inside Engelmann’s theory of instruction. Engelmann is very good at teaching very well-defined cognitive routines or procedures. In everyday language: when we teach a worked example where a student has to go through ten steps to answer a question, it’s a very useful way of making sure the student understands every step and can do it consistently. And then the book describes the process of going from overtisation — being overt about the ten steps — to covertisation: the idea that the student doesn’t actually think through the ten steps in their head any more. They’ve somehow reorganised the topic, the concept, the procedure, in such a way that it’s become self-evident to them how to do it. And Engelmann just says: this happens. You do practice, and then it happens.
I look at that and think, that’s interesting, because that’s the one time Engelmann really goes for discovery learning. He’s saying: just try, just work it out, and at some point it happens. But speak to any maths teacher and they’ll say the interesting thing in a class is that sometimes it happens and sometimes it doesn’t. And we can’t really work out why, because it’s not happening at the point of explicit instruction — it’s happening at the point of practice. Something is going on inside the heads of some students as they’re doing the same procedure again and again, where they manage to reorganise knowledge in a productive way that means they never have to memorise that procedure again. They’ve got no revision to do for the exam; it’s just obvious to them. And then there’s another load of students who can do the procedure completely fluently, but will always have to revise and memorise it. We call that mathematical understanding, don’t we? And we think it’s important, but we’re not great at working out what’s going on and what it is.
So I use that as an example of things that will potentially remain undiscovered when we work on online platforms. We’ll continue to see this variation, and we may not make much progress on improving things like that, because we don’t have good theoretical ideas about how to manipulate it. And without good theoretical ideas, we can’t design the experiments to work out how to manipulate it. But that doesn’t mean we’ve all got to go back to classroom teaching as the only option — because classroom teaching has all the same problems that the online platform will have.
Phil: Amazing. Thank you so much, Becky. That’s been incredibly thought-provoking for me, and your work in general has been really helpful for thinking through the massive project of making learning into a science. Thank you.
Becky: Pleasure. Good to speak to you.