Evidence receipt / belief
Published · transcript-backedLuke Darlow: belief
23 Nov 2025 Machine Learning Street Talk He Co-Invented the Transformer. Now: Continuous Thought Machines - Llion Jones and Luke Darlow [Sakana AI]
“I I did a lot of the work, but we also had a lot of people in different areas and doing different parts of it that I think an 8 month life cycle for a paper seems a bit long for AI research at the moment.”
Source trail
Everything needed to verify it.
- Speaker
- Luke Darlow
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 23 Nov 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…I think now is a good segue to talk about this paper in a little bit more detail. So that this continuous thought machine, you were just pointing to it before. Look, first first of all first of all, mate, introduce yourself and set this thing up for us. Name's Luke. I am a research scientist at Sakana AI. And my primary sector of research is this continuous thought machines. It took us somewhere in the region of about 8 months working on this project with the whole team. I I did a lot of the work, but we also had a lot of people in different areas and doing different parts of it that I think an 8 month life cycle for a paper seems a bit long for AI research at the moment. But, yes, to the to the actual technical points of the paper. So we call it continuous thought machines. It originally had a different name. We called it asynchronous thought machines before. But every single time people asked us what's the asynchronous part, it became a bit confusing. So continuous thought machines basically depends on 3 novelties. The first 1 is having what we call a internal thoughts dimension. And this is not necessarily something new. It's related conceptually to the ideas of latent reasoning. And it's essentially applying compute in a sequential dimension. And when you start thinking about ideas and problems in this domain and in this framework, you start understanding that many problems that look like or solutions to problems that look intelligent are often solutions that have a sequential nature. So for instance, 1 of the primary tasks that we tested in the continuous thought machines was this maze solving task. And solving mazes for deep learning is is quite trivial. It's really easy to do if you make the task easy for machines. And 1 of the ways to do this is you give an image of a maze to a neural network, like a convolutional neural network, And it outputs a image. Same size of the maze and it's zeros where there isn't a path and there's ones where there is a path. There's some really brilliant work showing how you can train these in a careful way and scale them up essentially indefinitely. And this is fascinating and really interesting idea of how to solve this. However, when you take that approach out of the picture and you ask what is a more human way to solve this problem, it becomes a sequential problem. You have to say, well, go up, go right, go up, go left, whatever the case may be to trace a route from start to finish. And when you constrain that simple problem space and you ask a machine learning system to solve it like that, turns out to actually get much, much more challenging. So this became our hello world problem for the CTM. And applying an internal sequential thought dimension to this is how we went about solving this. 2 other novelties that we can touch on and talk about. We we sort of rethought the idea of what neurons should be. There's a lot of excellent research in this world in cognitive neuroscience particularly exploring how neurons work in biological systems. And then we get on the other side of the scale how deep learning neurons work which the quintessential example is a ReLU. It's off or on in a sense. y exploring how neurons work in biological systems. And then we get on the other side of the scale how deep learning neurons work which the quintessential example is a ReLU. It's off or on in a sense. And this very very high level abstraction of neurons in the brains feels a little bit myopic. So we approached this problem and said well, let's on a neuron by neuron basis, let this neuron be a little model itself. And this ended up doing a lot of interesting work on how to build dynamics in the system. The third novelty here is, as I said before, we have this internal dimension over which thinking happens. We ask the question, well, what is the representation? What is the representation for a biological system when it's thinking? Is it just the state of the neurons at any given time? Does that capture a thought, if you wish? If I can be controversial and use the term thinking and thought. And my philosophy with this is no, it doesn't. That the concept of a thought is something that exists over time. So how do we capture that in in engineering speak? We instead of measuring the states of the model that is recurrent, we measure how it synchronizes, how neurons synchronize in pairs along with other neurons. And this opens up the door to a huge array of things that we can do with this type of representation.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.