Evidence receipt / belief
Published · transcript-backedTim Scarfe: belief
23 Nov 2025 Machine Learning Street Talk He Co-Invented the Transformer. Now: Continuous Thought Machines - Llion Jones and Luke Darlow [Sakana AI]
“Very cool. And I think I didn't pick up on this. So you're doing a fixed number of steps.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 23 Nov 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Yes. I think the boundary in terms of computation from a Turing machine perspective, if you wish, is really interesting because the notion of being able to write your tape, read from that tape, then write again to be in a Turing compute system, Turing complete system, is obviously an incredible idea that has completely changed the world. And I think the primary difference with let's talk about transformers versus what we're trying to do with the CTM is that the process that the CTM thinks in, we can apply that process, that internal process to breaking down a problem. So the problem itself can be a single there is a single solution to this problem. And you could do that in 1 shot. You could, as I explained with the maze, you could just process that in 1 shot. But there are certain phrasings of problems that are real problems that doing so becomes exponentially more challenging. So in the maze task, a really good example is that if you try to predict a 100 200 steps down the path in 1 shot, no models that we could train, not even our model could do that. And we needed to actually build an auto curriculum system where the model first predicted the first step, and then when it could predict the first step, then it we started training it on the second and third and fourth step. And the sort of result in behavior of this is where it gets interesting. 1 of the of the ways that I like to do research and that I encourage people who work with me to do research is understand the, if you wish, the behavior of a model. We are getting to a point now where the models that we build are demonstrably intelligent in ways that keep surprising us. And breaking that down into a single set of metrics or even a finite single metric about performance seems maybe not to be the right way to do it for me. And understanding the behavior actions that those models take when you put them in a system and train them in a certain way seems to reveal more about what's actually going on under the hood. Very cool. And I think I didn't pick up on this. So you're doing a fixed number of steps. So you have like a context window. And did you say that you've set that around a 100 steps? So for the for the maze task, the model always observes the full image. At every step, the CTM will absorb observe the full image. For argument's sake, those images could be tokens from a language the output of a language model. Those inputs could be numbers that that model has to sort, whatever the case may be. It should be agnostic to data. That's how we've tried to build it. But in the maze task, the model can continuously just observe the data. No matter where, it can look at the whole image simultaneously but it uses attention to retrieve information from the data. And it has, let's call it, 100 steps that it can think through. And what we do is we pick up at some point, the model solves 3 steps through the maze. So it says I'm going to go up, up, and right. And then it's correct. But then it makes the wrong tone. And at that point, we stop supervision. We only train it to solve the fourth step. So 1 more than what it could. In practice, we do it 5 but the principle holds. And when you do that, it's a self bootstrapping mechanism. And I think the intuitive listener will understand how that extends to other domains, other sequential domains, for instance, like language prediction, many tokens ahead, that sort of thing.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.