High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Tim Scarfe: belief

23 Nov 2025 Machine Learning Street Talk He Co-Invented the Transformer. Now: Continuous Thought Machines - Llion Jones and Luke Darlow [Sakana AI]

“If we learned it constructively, so we, you know, you speak about this in your paper, this complexification, the abstract building blocks, and you can do adaptive computation.”

— Tim Scarfe

Source trail

Everything needed to verify it.

Speaker
Tim Scarfe
Attribution
Verified speaker
Claim type
belief
Recorded
23 Nov 2025
Publisher
Machine Learning Street Talk

Transcript context

…around enough, it can it can fit the spiral and get a high accuracy, but there's no feeling when I look at those that that image that the ReLU version actually understands that it is a spiral. Right? And when you represent it as a spiral, it actually extrapolates correctly, because the spiral just keeps going out. You're touching on something fascinating there, because, you know, we were talking about the need for adaptivity and adaptive computation. I'm really inspired by Randall Bellis Dreiro's spline theory of of neural networks. And we we've had them on many times. And you can look on the TensorFlow playground. You can look what happens when you have a ReLU network on on this, you know, spiral manifold. And, you know, you'd you'd be forgiven for thinking that these things are basically a locality sensitive hashing table. Right? Because they they do. They they they they partition the space and and they they can predict the spiral manifold. Right? But we wanna do something a little bit more different than that. And it also comes into this imposters thing because just tracing the spiral manifold, but not continuing the pattern, there's a big difference between that. And so from an impostor perspective, just just tracing the pattern is not learning it abstractly or constructively. Right? If we learned it constructively, so we, you know, you speak about this in your paper, this complexification, the abstract building blocks, and you can do adaptive computation. You understand the spiral. That means that with adaptive computation, can continue the spiral. And then you can update the model's weight so it has adaptivity because that's so important for intelligence. So we know that we need models that can do these things. But for some reason, they're they're so sycophantic. They're they're almost better than an adaptive intelligence system because they tell us exactly what we want to hear. They seem so intelligent. But we know that they're missing these fundamental properties. I'm still fairly sceptical when I see video generation models. You know, we went through a phase where you could detect them because of the number of fingers on somebody's hand. Right? And, yes, with more data, with more compute, with better training tricks, okay, they submit, and now they usually do have 5 fingers. But did we fix the problem, or did we just use more brute force to just, you know, force the the neural network to know it's 5 fingers, where something that actually had a much better kind of representation space it's almost mad that it's controversial to say that we should represent a spiral like a spiral. But, you know, something that could do that generally, that if it if it represented a human hand the way that, you know, maybe I represent a human hand, then maybe it would be much easier to count how many fingers are on a on a hand. It's unfortunate that they work so well. It's unfortunate that scaling works so well, because it's too easy for people to just sweep these problems under the carpet.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence