Evidence receipt / belief
Published · transcript-backedLuke Darlow: belief
23 Nov 2025 Machine Learning Street Talk He Co-Invented the Transformer. Now: Continuous Thought Machines - Llion Jones and Luke Darlow [Sakana AI]
“I think a lot of the really fascinating work in the last few years that I found fascinating in the literature of language models has been related to what 1 can actually call a new scaling dimension.”
Source trail
Everything needed to verify it.
- Speaker
- Luke Darlow
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 23 Nov 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…So we were speaking about this yesterday, Luke, that when folks apply transformers to things like the ARC challenge or things that need reasoning, we need to do lots of domain specific hacks. So the architects who were the winners of last year's challenge, they did depth first search sampling. And some folks have been experimenting with using language representations or, you know, using DSLs. And some part of this is to do with the reachability of language, right? And language is quite dense, which means you can kind of monotonically increase. But if I understand correctly, your system might have some interesting properties for reasoning and for discrete and sparse domains and also for sample efficiency. Because we want we wanna build a system that can actually do well on things like the ARC challenge. But can you kind of explain in simple terms why you think this architecture could be significantly better than transformers for doing those things? I think a lot of the really fascinating work in the last few years that I found fascinating in the literature of language models has been related to what 1 can actually call a new scaling dimension. I, in some sense, see continue a chain of thought reasoning as a way of adding more compute to a system. That's obviously just 1 small part of what that really is and what that really means, but I think it's quite a profound breakthrough in some sense. Now, what we're trying to do is is have that reasoning component be entirely internal yet still running in some sort of sequential manner. And I think that that's rather important. And you spoke earlier about Gemini's diffusion language modeling, and I think that there are a lot of different directions that are exploring this right now. I do think that the continuous thought machine with the ideas of synchronization and multi hierarchical temporal representations gives a certain flexibility on that space that other people are not yet exploring. And that richness of that space being able to project the next step to solve the challenge and the next 100, the next 200 steps to be able to break that down into a process that a model can then very quickly search that process in its high dimensional latent space becomes something that feels like a good approach to take. Do you see any relationship between this architecture and, you know, Alex Graves' neural Turing machine?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.