High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Llion Jones: evaluation

23 Nov 2025 Machine Learning Street Talk He Co-Invented the Transformer. Now: Continuous Thought Machines - Llion Jones and Luke Darlow [Sakana AI]

“Transformers where you just applied it to a new problem, and it just was so so much faster to train, and you just got such higher accuracy that you just had to move. And I think the deep the deep learning revolution was also another example of that.”

— Llion Jones

Source trail

Everything needed to verify it.

Speaker
Llion Jones
Attribution
Verified speaker
Claim type
evaluation
Recorded
23 Nov 2025
Publisher
Machine Learning Street Talk

Transcript context

…I mean, there's also this interesting thing that even in spite of great new research, I mean, was speaking to Seth Hochreiter, and he's got all of these new architectural ideas, and OpenAI aren't implementing them. I mean, Google are doing this diffusion language model, which is quite cool. And I'd like to know your opinion on why that is. So there's a few philosophies floating around like this concept of a universal representation, that there are universal patterns. And the transformer representations resemble those in the brain. And it's rather led to this idea of, well, we don't need to use different architectures, because if we just have more scale and more compute, then all roads lead to Rome, so why would we bother doing it any differently? There's actually better. Right? There is actually already architectures that have been shown in the research to work better than Transformers. Okay? But not better enough in order to move the entire industry away from such an established architecture where you're familiar with it, you know how to train it, you know how it works, you know how the internals work. Right? You know how to fine tune them. You have all these softwares already set up for training transformers, fine tuning transformers, inference. So if you wanna move the industry away from that, being better is not good enough. It has to be obviously, crushingly better. Transformers were that much better over RNNs. Okay? Transformers where you just applied it to a new problem, and it just was so so much faster to train, and you just got such higher accuracy that you just had to move. And I think the deep the deep learning revolution was also another example of that. Right? Where you had plenty of skeptics, and people were pushing neural networks even back then, and people are going, no. We think symbolic stuff will work better. But then they demonstrated it as being so much better that you couldn't ignore it. This fact makes finding the next thing even harder. Right? That's the gravitational pull of always pull pulling you back to, oh, okay. But a transformer's good enough. And, yeah, you made a cool little architecture over here that, yeah, it looks like it's it's got better accuracy, but OpenAI over here just made it 10 times bigger and it beats that. So let's just keep going. May I also submit that there could be an additional reason, which is, you know, I love that fractured and tangled representations paper. There's there's this shortcut learning problem. Mhmm. And I think that there's a little bit of a mirage going on here. And there there might be problems with these language models that we don't, you know, that we're not fully aware of. And there's also this thing that we're seeing that we are starting to bastardize the architecture. So we know we need to have adaptive computation for reasoning. We know we want things like uncertainty quantification. And what we're doing is is we're bolting these things on top, rather than having an architecture which intrinsically does all of these things that we know we need.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence