Unstoppable Learning with Kris Boulton

Unstoppable Learning with Kris Boulton

What if we could design learning so that it was impossible for students not to understand? This is what ​Kris Boulton, CEO of Unstoppable Learning and former teacher, thinks we can aim for. There is no reason, he argues, why almost every student can’t get a 9 in Maths.

Unstoppable learning is based on Engelmann and Carnine’s Theory of Instruction, one of the most rigorously scientific approaches to instruction. The approach involves decomposing maths (or other subjects) into its smallest parts (‘atoms’) before rebuilding through four different elements: categoricals, transformations, facts and subroutines. If teachers identify atoms correctly, sequence and teach them effectively before ‘chaining’ them together then failure becomes impossible for the student.

Engelmann and Carnine’s book is notoriously dense so if you’re looking for other resources on Unstoppable Learning and Kris’s work then check out the Unstoppable Learning Substack or Kris’s episode with Craig Barton.

Theory of Instruction: Principles and Applications by Siegfried Engelmann |  Goodreads

Philip Bell: Kris is the CEO of Unstoppable Learning, and I have personally been extremely excited by Unstoppable Learning since I heard about it because I think it’s a really interesting way of thinking about teaching. The thing that excites me most is just the sheer ambition. So I’m really excited to introduce Kris today, who’s at Unstoppable Learning, previously worked at Uplearn, and then before that worked as a teacher at Kings Solomon Academy and elsewhere. So without further ado, Kris, would you like to take it away?

Kris: Yeah, absolutely. Thanks, Phil. Thanks again, everyone, for having your cameras on. I do appreciate it.

And with respect to the ambition that Phil was just talking about — so first and foremost, everything I talk about here is going to be set in the context of mathematics. I apologize that a couple of the examples are plucked from secondary mathematics, but not really very much. It’ll still have a bit of a primary lens. But Phil’s got a bit of a question about this for me later — it can apply across other subjects as well. I’m just not going to share the concrete examples this time around.

And usually when speaking with people about what our goal is, we’ll describe it in terms of 100%. We’re looking for 100% success for 100% of kids, 100% of the time, in a subject where that success or failure — mathematics — is very unambiguous. You can see if they have succeeded or not. That immediately invites the question: why is it that we can be so confident in shooting for that and committing to that with the teachers and schools that we work with?

Do you mind turning on sharing for me and then I’ll get the presentation going?

Philip Bell: Yes. Good.

Kris: Okay, you should be able to share now. Thank you very much.

So, first things first. What’s the core of everything we do is what’s referred to as logically faultless communication. Very often, we’ve essentially taken this term and given it a bit of a simpler rebrand, terming it unstoppable learning. So if we want to use the academic language from the literature: logically faultless communication. If we want something a bit more colloquial and easy to get our heads around, this is what the goal is — this is the outcome that logically faultless communication provides.

And the best way to explain what is meant by it is to show you. There’s a little bit of interactivity here — if you’ve got access to the chat feature, or maybe we’ll have some mics on and shouting out, we’ll go for that. But the way I’m going to try to demonstrate this is by teaching you the meaning of the made-up word, “glurb.”

[Technical difficulties — filter issue resolved]

So, what is glurb? This is my AFL.

Participant: Waving something.

Kris: Waving something. Thank you very much. Did anybody think it meant something different?

Participant: I thought it might be something green.

Kris: Something green — that comes up quite frequently. Anything else it could have been?

Participant: Green and squidgy.

Kris: Green and squidgy — haven’t had that one before, but I can see where that came from. Last two people — would either of you like to give it a go?

Participant: I was going to say wafting.

Kris: Wafting. Thank you.

[Participant on mute — unable to contribute]

OK, we’ll progress. We’ve got several different ideas there. And so we can look at this from two perspectives. One perspective is that you’re just not a very smart group of people, and that’s why you can’t get what is so obviously the meaning of glurb — my demonstration couldn’t have been clearer. Or the other perspective we can look at is: in many ways, you are surely the best possible group of people I could hope to teach. You’re motivated enough to turn up after work in your own free time. You’re all intelligent, educated adults. Maybe the fault did not lie in you and your brains. Maybe the fault lay somewhere in my teaching.

So on screen now I have various examples of what people have offered previously. Let’s try to do this via the chat feature — might be the easiest way to do this synchronously. Two people on phones is going to make it harder for you. Could all of you go off mute?

OK. I’m just going to ask you to shout some things out in a moment based on what I do next, because I’m going to give this another go and try to do a better job than I did the first time.

So first of all — this is glurb. This is not glurb. Which two number options did we just rule out?

Participants: Five and six.

Kris: Yeah, five and six. Correct. Thank you. Next one — this is glurb. Which two options were just ruled out?

Participants: Three and four.

Kris: Brilliant. Thank you, everybody. Finally, this — this is not glurb. What is the last option that has just been ruled out?

Participants: Number one.

Kris: Number one. Brilliant.

OK. That is logically faultless communication. It’s taking something where instead of having every mind in the room be uncertain or coming up with a range of different ideas about what your meaning might be, every mind in the room converges on the exact same interpretation — the only interpretation that is logically consistent with what was presented.

I’m going to go through one more demonstration and use it to exemplify one of the techniques used in the second sequence. This time I’m going to teach you the meaning of the made-up word, “blurble.”

[Camera adjusted]

OK. First of all, this is blurble. This is blurble. This is blurble. And this is blurble. What must you now logically conclude is the meaning of the word blurble?

Participant: A writing implement?

Kris: Absolutely — it’s a writing implement. I’ve done this demonstration dozens and dozens of times now, with hundreds, thousands of teachers, and every single time I get the same answer. A writing implement, something you write with. Sometimes people say it’s a pen.

I’m going to add one more example to that sequence. And as soon as I do, the meaning of the word blurble is going to transform before your mind’s eye. We’re changing our minds about things all the time — it’s very, very rare that somebody warns you in advance that you’re about to change your mind about something. So this is a very special moment. Watch out for it.

This is blurble. This is not blurble. What now must be the meaning of the word blurble?

Participant: A writing implement with a lid on.

Kris: With a lid on. Exactly correct. So what made all the difference? Because I showed you four examples of blurble and not one of you got it. Nobody ever gets it. What made the difference?

Participant: A non-example.

Kris: Exactly. A non-example. A negative example.

Now, I’m going to show a practical example of how to make use of this. It used to be the case that if I were doing this demonstration 15 years ago I’d often say to people: if nothing else, remember this one thing — use non-examples. We learn what things are not just through examples of what they are, but through examples of what they are not. And 10, 15 years ago, that was a revelation to people. Now that idea, certainly at least in the secondary sector — I suspect in the primary as well — does seem to have proliferated quite a lot. I’m seeing non-examples used quite a bit now, and people talking about it. But you can use non-examples well, and you can use them not so well. So we’re going to go into a bit more detail.

Just before I do that, Phil asked that I explain a little bit more about what the Unstoppable Learning approach is and where the ideas come from. So a lot of it is rooted in Theory of Instruction: Principles and Applications, written by Engelmann and Carnine. It’s an absolute magnum opus — an extraordinary, extraordinary book. I would love to recommend that all of you go and read it, but there are very, very few people I’m aware of who’ve picked it up and tried to read it and have successfully made sense of it. It’s not any fault of theirs. The book is almost written to be illegible — so much so, famously so, that in the foreword to its second edition, Rob Dixon said: “I’m asked many questions about Theory of Instruction. Why this? Why that? Why is it so difficult to read?” — in its own foreword. So it’s a very difficult book to read, but that’s where a lot of it comes from, cross-referenced with a lot of cognitive science and behavioural psychology.

But one particular bit that I think is game-changing in Engelmann and Carnine’s ideas is their analysis of cognitive learning. So this image — if you’ve never seen it — is a portion of the universe of mathematics created by Complete Maths. And I think it’s a pretty good depiction of how we think about mathematics. We think about mathematics as thousands of different ideas that need to be learned. And that is fair — there are thousands of different ideas that do need to be learned, and they have many connections to them as well.

So then we conclude: well, that means there are thousands of different things about mathematics that I need to learn how to teach. And then it’s not really thousands of different things — because to take a couple of examples from the secondary sector, expanding a pair of brackets, or sharing a quantity into a ratio — for each of those I can off the top of my head think of five different ways of teaching kids how to do those things. And then there are people who will say, actually, you should teach multiple methods to students. And a lot of the rationale behind that often seems to be: we think different things will make sense to different kids, we’re not really sure what will work for different children, so we should teach them lots of different methods in the hope that some of it sticks for different people.

So then that means there are tens of thousands of things that we have to learn how to teach. And then — OK, so I’m going to try teaching a thing, and I might not get it right first time. When do I get to try this again? Usually it’s next year, a year later, if we’re lucky. So not even that necessarily, if you’re moving between year groups in the primary sector. And so suddenly, you get 30, 40 chances maybe — tops — with a year or so in between each opportunity, across a three or four decade career, to figure out how to perfectly teach tens of thousands of different things. And quite suddenly, why teaching mathematics in a way that every child can be successful seems like a really impossible task starts to make sense.

And so what’s really brilliant about Engelmann and Carnine’s work — almost more than anything — is they don’t say there are thousands of things to learn. Or more precisely, they say: each one of these things, you can pick it apart. When you pick it apart — what we call atomising — it turns out that every single one of them will be composed of some combination of these same atomic elements, every time, the same form.

And in the book itself, this is their picture of their analysis of cognitive learning. I think a simpler way to structure it is to look at it like this. Step one: analyse the knowledge. Here’s the thing that you’re going to try and teach kids — what are its constituent atoms? What are the elements that make it up? Because once you know that, for each element the way that you communicate it can be effectively the same every time. What the knowledge is informs the type of communication. And it’s basically more or less the same communication structure every single time, which means you don’t have to learn tens of thousands of things to teach in mathematics. You only need to learn about four. And that feels an awful lot more manageable.

And then the final part of the process: analyse the response. Give students something to respond to. Do they respond the way that we expect them to? If they do, great — we can probably move on. If they don’t, then we don’t blame the student or say there’s something wrong with their brain. Instead, we infer there was something wrong with the communication. It doesn’t matter that two-thirds of the group got it right and one-third didn’t — two-thirds of the group got it right despite us, not because of us. Every time I do the glurb demo, some people get it right and some people say the wrong thing, and it’s because of my communication, not because there’s anything wrong with their brains. So we go back, reanalyse the communication, tweak it, and try again.

So now we’ll look briefly at the first of these four elements — categoricals — what that is and how to communicate them.

To identify them effectively: if you can meaningfully ask the question “is this an X?”, then X is a categorical concept. So for example, gradient is not a categorical, because I can’t meaningfully ask “is this a gradient?” And solving an equation is not a categorical concept, because I can’t meaningfully ask “is this a solving an equation?” But equation is a categorical concept because I can meaningfully ask “is this an equation — yes or no?” Prism is a categorical concept because I can meaningfully ask “is this a prism?” And histogram — that’s another categorical because I can meaningfully ask “is this a histogram?” But “find the mean from a histogram” — that’s not categorical. “Is this a find the mean from a histogram?” Doesn’t make sense.

So we can take 60 seconds tops to give that a go. Here are a bunch of different things we might want to teach in maths — for each one, just write down yes or no: is it a categorical concept?

[30 seconds]

If you have access to the chat channel, could you just type Y or N for each one?

[Responses come in]

OK, we’ve got everything through there. It looks like that was 100%. Brilliant.

OK, so now we can identify categoricals. How do we communicate them? First and foremost, we don’t just show an example and then talk about that example a little bit. We don’t even show many examples. We’re going to use the concept of triangle here. We have to use non-examples.

But as I hinted at earlier, it used to be that you could say “use non-examples” and that was a revelation. Actually, you can use them well or not so well. One way of not using them so well — when I just went to Google and typed in “examples and non-examples” and went to Google Images, the very first image that came up was one that takes a whole bunch of examples and non-examples plus other stuff totally at random and just throws them at the kids all at once. It’s a shame because there are some good ones in there, but it would be very overloading for most people.

So what we’re going to do instead — what Engelmann and Carnine propose — is use what they call an initial instructional sequence. To communicate the concept “triangle,” this would mean a sequence of examples and non-examples, one at a time. And as a rule of thumb they would follow this pattern: a negative example first, followed by a positive example, followed by another positive example, another positive example, then a negative example. This structure is now often abbreviated to NPPPN.

I’ll show an example of what an initial instructional sequence could look like for communicating the concept of triangle.

“I’m going to teach you what we mean by the word triangle. To start, this is not a triangle. But this is a triangle. This is a triangle. And this is a triangle. And this is not a triangle.”

Negative, positive, positive, positive, negative — NPPPN.

Now I’m going to share a second instructional sequence. Like this one, it will have five examples in total, structured NPPPN. But my second example is going to be a bad initial instructional sequence — not a good one — because there are four principles of instruction that should be used each time in the design of these sequences. My good example adhered to all four. This next one does not adhere to any of them. Afterwards we’ll play a game of spot the difference.

“OK, so I’m going to teach you what we mean by the word triangle. To start off, this isn’t a triangle — obviously not a triangle, because we’ve seen this before. We all know that this one’s a rectangle. It’s got four sides. It’s obviously a rectangle. But this — this is a triangle. It’s got three sides, unlike the rectangle. The red thing doesn’t matter. Like here, we’ve got one that’s black and not red. That doesn’t really matter either. It’s still a triangle. Same as this one. This one is also a triangle — kind of. I know it’s sort of filled in now, it doesn’t have an outline, it’s sort of all orange. Just ignore that. That doesn’t matter. That doesn’t change the fact that it’s a triangle. And then finally, this — this isn’t a triangle. Well, this is actually a shape called a pentagon. So pentagons have five sides, not like the rectangle and the triangle. Actually, ignore — look, we’re going to look at pentagons better next week. Ignore what I was just saying. What really matters at this point is just that that is not a triangle. So yeah — not a triangle, and then not a triangle. But these ones here, these were triangles.”

OK. There were four meaningful distinctions between those two sequences. What were they?

Participant: You were far too wordy in your second one.

Kris: Yeah, everyone always likes that one.

Participant: All your three triangles are incredibly similar in shape and orientation.

Kris: Correct, that’s the second one.

Participant: You were teaching a new concept with the pentagon as well, whereas actually your focus was on the triangles.

Kris: That is very true, absolutely — I’ve moved on too far there.

Participant: You didn’t just stick with the word triangle, you used other language as well.

Kris: We’ve still got two out of the four. Any others?

Participant: There were some non-important features in the second set — features that weren’t germane to the definition of a triangle.

Kris: That is very true. So we’ve got things that are irrelevant that have been thrown in and varied all over. That gives us a third principle. There’s one last one.

Participant: Your negative examples are totally different — not triangle-like.

Kris: Yeah, exactly. My negatives are quite close to the positives in the good sequence, but not in the bad one. So we’ve got the four — we’ll go through them one by one.

First: don’t change irrelevant features — in this case, like colour. Here, all the irrelevant features are kept quite constant. That helps to reduce informational noise and direct students’ attention to what matters. I don’t know how familiar everybody already is with the cognitive science model of working memory and long-term memory, but in a nutshell: working memory is very constrained for everybody. Some of us have a little bit more available, some a little bit less, but it’s constrained for everybody. The more we’re asking people to attend to and think about and process all in one go, the more likely they’re just going to miss everything. So here, we reduce the informational noise and focus attention on the things that actually matter. That first principle — don’t vary irrelevant features — is called the setup principle.

Second — the one that everybody enjoys and picks up on pretty quickly — use the same words to present each example. For each of the good examples I presented, very simple, very consistent wording every time. For the bad example it was all over the place. That’s called the wording principle.

Those first two principles apply to the whole sequence, top to bottom. The next two apply to the specific examples we choose and how we sequence them.

Third: for positive examples, show maximal variation. Within the constraints of the first principle, vary them as much as possible. In the bad example, we just had three right-angled triangles all oriented the same way — you’d leave thinking triangles always look like that. In the good example, we’ve tried to show differently shaped triangles in different orientations. This is called the sameness principle — because what we’re doing is varying our positives a lot, but treating them the same each time. Even though I’ve made this look totally different, it is still a triangle. If I have positive examples that vary massively but I’m still saying “yes, this is an example,” there must be something that is the same about them all — and that’s what we’re communicating.

Fourth: between the positives and negatives, show very minimal variation. In the bad sequence, the negatives look very different from the positives. In the good sequence, small change — and suddenly it’s no longer a triangle. By varying the positives a lot, we’re showing the range of what the concept could be. And by keeping the variation between positive and negative minimal, we’re putting a limit on that range — this tiny change suddenly means it’s no longer a triangle. This is called the difference principle.

And then there is a fifth principle as well. However, that brings us to the end of what we have time for today — so it will have to remain a mystery.

Phil, would you like to take it from there? We have a couple of options. Option one: I can give an exercise for everyone to try now to construct one of these sequences. Option two: we can take questions.

Philip Bell: Let’s start with questions. If anyone has a question please raise your hand or put it in the chat. But to start with, I had one question.

This approach is clearly applicable to different subjects and phases. But there’s a spectrum of applicability — the mindset of getting 100% for every student is clearly applicable, and the idea of logically faultless communication is clearly applicable. But do you think there are some aspects of the approach that differ depending on the subject? Specifically, do you think some subjects are more easily broken down into atoms than others?

Kris: So they are. I need to explain several concepts to make sense of this — I’ll give it a go.

The first distinction that’s really important is between substantive and disciplinary knowledge. The easiest example to understand comes from science. In science, things that are scientific facts, scientific truths, or scientific laws — things we’ve figured out about science, like Newton’s three laws of motion or the ways that chemicals react — that is all substantive knowledge or substantive content. Disciplinary knowledge is how scientists think and how they behave in order to discover all of those ideas about science.

And that breakdown exists for a lot of different subjects. With mathematics it’s the same — pretty much all the content we teach in school is substantive. But professional mathematicians are obviously not having a teacher hand them work to do. They’re doing work to discover new ideas about mathematics — conjecture, inquiry, problem solving, et cetera — and that’s disciplinary. History: what we know or think we know about the past is substantive; how historians have come to know that is disciplinary.

So different subjects in our examinations prioritise the substantive and the disciplinary differently. Speaking about GCSE exams: maths is fully, 100% substantive. Science is 90, 95% substantive — it tries to creep in a little bit of the disciplinary as well. Languages are 100% substantive. English almost goes the other direction — very little of it is substantive content, and very often it is trying to promote thinking skills, writing skills, reading skills that are baked into the way we assess the subject. And history and geography are a bit of a 50/50 mix.

The next distinction is between a quality-based and a difficulty-based assessment paradigm. Mathematics is the purest example of difficulty-based — it’s like doing the high jump. The way you get people to compete in the high jump and see who’s better: you just have a bar, get them to jump over it, and then raise the bar again. And you keep going until one person can no longer clear it. Maths exams are set up like that — we ask 100 questions and they get harder the further you get, until people basically tap out and can’t get any further.

Whereas the alternative is something more like Strictly Come Dancing. You go out there, do your one big performance, and then multiple people assess the quality of your performance. And that, again, is how English exams tend to operate — you get about four questions, 30 marks each, and you’re being assessed on the quality of your writing. Usually multiple people have to moderate that assessment as well.

And then the final thing I need to explain is the distinction between convergent and divergent learning intentions. In a maths classroom, if you gave all the kids an exit ticket before they leave and every single one of them wrote down the exact same answer and the exact same working, you would call that a massive win — I’ve completely succeeded in teaching this, assuming you’re able to prevent cheating and copying. On the other hand, if you asked all your kids to write a sentence or a paragraph about something and they all handed in exactly the same words, word for word — something weird has happened here. That doesn’t seem right.

So in maths we have very convergent learning intentions. We want everyone to learn and do the same thing. In English we intuitively want them to be more creative and expressive and do different things, within constraints.

All those different things — having a substantive focus, a difficulty-based assessment, and convergent learning intentions — they tend to line up. And the more substantive, difficulty-based, and importantly convergent a subject is in its school manifestation, the much more amenable it is to this kind of analysis. So the sciences are very amenable to it. Languages are very amenable to it. When you get to the humanities, some of the content is very amenable to it, but I don’t think the more disciplinary stuff is.

Now, back to English — it sort of depends, because some parts of English do actually fall on the convergent spectrum: things like spelling, grammar, punctuation, writing a technically correct sentence rather than one that is expressive. And Tom Needham in particular is very good at this — he’s a secondary English teacher who’s drawing from the same body of research and thinking about how to apply it specifically to English in the secondary context.

And then finally, for those subjects that do prioritise creativity, inquiry, problem solving, divergent thinking — all of those outcomes are predicated on having a load of the substantive knowledge. You kind of still need that before you can get to that bit.

So the short answer is yes, it applies across the board, and there’s a bit more detail there about how it might apply differently across different subjects across the spectrum.

Philip Bell: That’s pretty interesting. That was an interesting answer — as a kind of example of how to very clearly articulate something complex by substantiating some of the background thinking. I’d never actually thought about the comparison between assessment in English and the high jump, versus maths being more gradual. That’s a really interesting point.

That was awesome. I have one final question to round off the session: do you use this approach in your own learning? I think I use some of these approaches in my own learning — I was just intrigued to know whether you do yourself.

Kris: Do you mean if I’m trying to learn something myself?

Philip Bell: Yeah, I’m kind of thinking about your knowledge diet in general and how you think about this in relation to how you try to learn yourself.

Kris: I think it informs some of how I try to learn. I think it looks different in its application for a variety of reasons. Sometimes if I’m reading a description of something, I might ask myself: OK, well, what if we change this — would that not be that then? Or I’m very aware of my own uncertainty because I realise I’m asking about edge cases. Well, what if we change that little bit — does that still apply? Yes or no? GPT has actually been pretty good for that because I can sometimes push it and, depending on whether you can trust it, it will at least give you an answer. And I think there are ways of pushing it to give true answers.

I’m also hyper aware of when I don’t understand something. And I think a lot of people in society more broadly blame themselves when they don’t understand something — there’s probably a lot of insecurity going around about how smart people feel, possibly hangups from our experiences in school.

And I think also, just as a casual observation of society, quite often people will answer questions when they don’t really know what the answer is — someone’s asked them and they’ll give an answer, and they have no idea what the answer is. And then if their answer was unsatisfactory, or if I start probing — well, what about this, what about this? — I can almost see them starting to feel flustered and almost intimidated, because what’s happening is they’re getting caught out. They’re getting caught out in the lie of having pretended they knew. And it’s interesting that they feel like they have to pretend that they know.

Whereas I’ve kind of reached a point where I go very much the other way. If I don’t know or understand something, I don’t ever blame myself for that. I always think that must be a function of the experience or what is being communicated to me — and now maybe I’ll see what I can do to try and improve upon that or figure out what’s missing, if I care enough to invest in it. So it doesn’t manifest in me constructing NPPPN sequences for myself, but I guess it gets reinterpreted and filtered in that way.

Philip Bell: Amazing. Thank you so much. I’ve been learning Chinese lately and I’ve been trying to use the Michel Thomas method, which I know is sort of partly related to this in terms of thinking about — I did the first two lessons of Mandarin Chinese on Michel Thomas, but I didn’t get further than that because of time. It was 10 years ago.

Kris: Fair enough. I have done a good bit further with Urdu — my wife’s Pakistani heritage — with Pimsleur, which uses the same method. So if you ever go through a course like the European ones — in particular French, Spanish, German, Italian — Michel Thomas and Pimsleur, it’s absolutely fascinating to see the same kind of methodology, which is all the same as what I’m talking about here. They use the same ideas applied in two different ways, and they’re actually both slightly good at different things.

And with the Urdu one, there isn’t a Michel Thomas version. The deficiencies in the Pimsleur method and the Pimsleur courses I was able to make up for a lot with ChatGPT — by knowing what questions to ask ChatGPT, thanks to Michel Thomas. So that helped a lot.

Philip Bell: Interesting. I’d actually be really interested to get into the differences between Pimsleur and Michel Thomas. I might look that up after this.

Kris: It’d be a different conversation for a different group.