Evidence receipt / evaluation
Published · transcript-backedTim Scarfe: evaluation
23 Nov 2025 Machine Learning Street Talk He Co-Invented the Transformer. Now: Continuous Thought Machines - Llion Jones and Luke Darlow [Sakana AI]
“Yes, and on that point, I think maybe the most exciting thing about your paper is, you know, we were talking about path dependence and having this understanding which is built step by step, this process of complexification.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 23 Nov 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…The flavor of this kind of research is such that we didn't actually go out and actually try to create a very well calibrated model. Right? And we didn't even try to create a model that was necessarily going to be able to do some kind of adaptive computation time. Right? I was a very big fan of the the paper yeah. Adaptive computation time. It was Alex Graves, was it? But that paper, it had a massive amount of hyperparameter sweeps in it because in that paper, he needed to have a loss on the amount of computation that was being done. Yeah. Because anytime you try to do some sort of adaptive computation time research, what you're fighting is the fact that neural networks are greedy. Right? Because, obviously, the way to get the lowest loss is to use all the computation that you have access to. So unless you had, like, an extra loss that had a penalty that said, okay. Actually, you're not allowed to use all the computation. That's and and very, very carefully balanced loss. That's when you actually got the interesting dynamic computation time behavior falling out of the the model in that paper. But what was really gratifying to see with the the continuous thought machine is that because of the way that we set up the loss that Luke described earlier, adaptive commutation times seem to just fall out naturally. So that's more the way that I think research should go. Okay? Because we don't actually have a specific goal, or a specific problem we're trying to fix like that, or something we're trying to invent. It's more that we have this interesting architecture, and that we're just following the gradients of interestingness. Yes, and on that point, I think maybe the most exciting thing about your paper is, you know, we were talking about path dependence and having this understanding which is built step by step, this process of complexification. And I mean, maybe this is this is apropos in in the theme of world models in general. And also active inference. And I say active inference in big quotes because it's not Carl Friston's active, you know, maybe adaptive inference or something like that. We we want to build agents that can continue to learn, that can update their parameters, and most importantly, can construct path dependent understanding. And because it that's completely different to just understanding what the thing is. It's how you got there is very important. And this architecture potentially allows these agents using this algorithm to explore trajectories in spaces, find the best trajectories, and actually construct an understanding which carves the world up by the joints. Yeah. That's a that's a really neat perspective. I haven't actually thought about it like that. But yes, I think that particular stance becomes really interesting when you think about ambiguous problems. Because carving the world up in 1 way is as performant as carving it up in another way. Yeah. You know, perhaps the hallucination in language models is carving the world up in some fine way but it's just not performance in our measure of this is hallucination and actually that's not true. But in some other trace down the path of wanting to carve the world up through a auto regressive generation of tokens, you end up in a different carve up of that world. And being able to train a model that can be implicitly aware of the fact that it is actually carving up the world in a different way and can explore those manners, those descends down the carve up is something that we're often. I think it's quite an exciting approach to be trying to take a stance of let's break up this problem into small solvable parts and learn to do it like that. And how can we do this in a natural way without too many hacks?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.