Evidence receipt / evaluation
Published · transcript-backedLlion Jones: evaluation
23 Nov 2025 Machine Learning Street Talk He Co-Invented the Transformer. Now: Continuous Thought Machines - Llion Jones and Luke Darlow [Sakana AI]
“It's unfortunate that they work so well. It's unfortunate that scaling works so well, because it's too easy for people to just sweep these problems under the carpet.”
Source trail
Everything needed to verify it.
- Speaker
- Llion Jones
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 23 Nov 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…You're touching on something fascinating there, because, you know, we were talking about the need for adaptivity and adaptive computation. I'm really inspired by Randall Bellis Dreiro's spline theory of of neural networks. And we we've had them on many times. And you can look on the TensorFlow playground. You can look what happens when you have a ReLU network on on this, you know, spiral manifold. And, you know, you'd you'd be forgiven for thinking that these things are basically a locality sensitive hashing table. Right? Because they they do. They they they they partition the space and and they they can predict the spiral manifold. Right? But we wanna do something a little bit more different than that. And it also comes into this imposters thing because just tracing the spiral manifold, but not continuing the pattern, there's a big difference between that. And so from an impostor perspective, just just tracing the pattern is not learning it abstractly or constructively. Right? If we learned it constructively, so we, you know, you speak about this in your paper, this complexification, the abstract building blocks, and you can do adaptive computation. You understand the spiral. That means that with adaptive computation, can continue the spiral. And then you can update the model's weight so it has adaptivity because that's so important for intelligence. So we know that we need models that can do these things. But for some reason, they're they're so sycophantic. They're they're almost better than an adaptive intelligence system because they tell us exactly what we want to hear. They seem so intelligent. But we know that they're missing these fundamental properties. I'm still fairly sceptical when I see video generation models. You know, we went through a phase where you could detect them because of the number of fingers on somebody's hand. Right? And, yes, with more data, with more compute, with better training tricks, okay, they submit, and now they usually do have 5 fingers. But did we fix the problem, or did we just use more brute force to just, you know, force the the neural network to know it's 5 fingers, where something that actually had a much better kind of representation space it's almost mad that it's controversial to say that we should represent a spiral like a spiral. But, you know, something that could do that generally, that if it if it represented a human hand the way that, you know, maybe I represent a human hand, then maybe it would be much easier to count how many fingers are on a on a hand. It's unfortunate that they work so well. It's unfortunate that scaling works so well, because it's too easy for people to just sweep these problems under the carpet. You guys have possibly created what I think might be the best paper of of the year. This could actually be the innovation, which takes us to the next step. And you did did you get the spotlight in Europe's as well? Yeah. This year. And congratulations on that. So I think that's testament to how amazing this paper is.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.