High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / preference

Published · transcript-backed

Tim Scarfe: preference

23 Nov 2025 Machine Learning Street Talk He Co-Invented the Transformer. Now: Continuous Thought Machines - Llion Jones and Luke Darlow [Sakana AI]

“May I also submit that there could be an additional reason, which is, you know, I love that fractured and tangled representations paper.”

— Tim Scarfe

Source trail

Everything needed to verify it.

Speaker
Tim Scarfe
Attribution
Verified speaker
Claim type
preference
Recorded
23 Nov 2025
Publisher
Machine Learning Street Talk

Transcript context

…There's actually better. Right? There is actually already architectures that have been shown in the research to work better than Transformers. Okay? But not better enough in order to move the entire industry away from such an established architecture where you're familiar with it, you know how to train it, you know how it works, you know how the internals work. Right? You know how to fine tune them. You have all these softwares already set up for training transformers, fine tuning transformers, inference. So if you wanna move the industry away from that, being better is not good enough. It has to be obviously, crushingly better. Transformers were that much better over RNNs. Okay? Transformers where you just applied it to a new problem, and it just was so so much faster to train, and you just got such higher accuracy that you just had to move. And I think the deep the deep learning revolution was also another example of that. Right? Where you had plenty of skeptics, and people were pushing neural networks even back then, and people are going, no. We think symbolic stuff will work better. But then they demonstrated it as being so much better that you couldn't ignore it. This fact makes finding the next thing even harder. Right? That's the gravitational pull of always pull pulling you back to, oh, okay. But a transformer's good enough. And, yeah, you made a cool little architecture over here that, yeah, it looks like it's it's got better accuracy, but OpenAI over here just made it 10 times bigger and it beats that. So let's just keep going. May I also submit that there could be an additional reason, which is, you know, I love that fractured and tangled representations paper. There's there's this shortcut learning problem. Mhmm. And I think that there's a little bit of a mirage going on here. And there there might be problems with these language models that we don't, you know, that we're not fully aware of. And there's also this thing that we're seeing that we are starting to bastardize the architecture. So we know we need to have adaptive computation for reasoning. We know we want things like uncertainty quantification. And what we're doing is is we're bolting these things on top, rather than having an architecture which intrinsically does all of these things that we know we need. Yeah. And I and I think the the our continuous thought machine is is an attempt at addressing those more directly. Right? Which which Luke will be able to tell you more about later. There's something still not quite right with the current technology. Right? I I think the the phrase that's becoming popular is jagged intelligence. Right? That the fact that you can ask an LLM something and they can solve literally, like, a PhD level problem. And then, you know, in the next sentence, it can say something just so clearly, obviously wrong that it it's it's jarring. Right? So I think this is actually a reflection of something probably quite fundamentally wrong with the current architecture, as amazing as they are. The current technology is actually too good. Okay? Another reason why it's it's difficult to move away from them. Right? So they're they're too good in in in the following sense. And you you spoke about the fact that we have these foundation models. That's okay. So that we have the foundation, we can do anything with them. Yes. I think current neural networks are so powerful that if you have enough patience and enough compute and enough data, you can make them do anything. But I don't necessarily think that they want to. Right? We're sort of forcing them. Like, they're universal pro approximators. But I think there are probably a space of, you know, function approximators that will more want to represent things in the way that a human represents them. So there's actually quite an obscure paper that is my poster child for this. It's called intelligence matrix exponentiation. Mhmm. And I think that it was actually rejected. So, you know, you can probably project the image of a figure 1, but there's an image of it solving, you know, the the classal spiral dataset of needing to separate the 2 classes in the spiral. Yes. And it has the decision boundary for a for both a classic RNN multilayer perceptron and a a tan h multilayer perceptron. And you can see they both solve it. Right? Technically, they both solve the problem because they they they classify all the points correctly and get a very good good test score on this on this very simple dataset. And then they show you the decision boundary for the for the m layer that they built in this paper, and it's a spiral. The layer represented the spiral as a spiral. Should shouldn't we? Should you know, if the data is a spiral, shouldn't we represent it as a spiral? And then if you look back at the decision boundaries for for the spiral and the classic ReLU multilayer perceptron, it's clear that you just have these tiny little piecewise linear separations. And and that's what I mean. Yes. If, you know, if you train these things enough and you push these little piecewise linear boundaries around enough, it can it can fit the spiral and get a high accuracy, but there's no feeling when I look at those that that image that the ReLU version actually understands that it is a spiral. Right?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence