High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Tim Scarfe: evaluation

22 Jun 2026 Machine Learning Street Talk He won a Nobel here for AlphaFold. Then he left. - John Jumper

“Even that, I think you can argue, maybe not entirely the story, but it's definitely not the story for proteins because that's the hardest problem is the large scale structure.”

— Tim Scarfe

Source trail

Everything needed to verify it.

Speaker
Tim Scarfe
Attribution
Verified speaker
Claim type
evaluation
Recorded
22 Jun 2026
Publisher
Machine Learning Street Talk

Transcript context

…re module was a geometrization engine that took a set of really quite good constraints that had very clear notion of the structure within those constraints and solved up the details. I think AlphaFold 3 diffusion is similar, and it's especially similar because, in fact, in images, okay, you start generating an image and you see especially these early trained diffusion models generate kind of colored blobs, and they start to decide what those colored blobs mean. And they pretty clearly kind of decide what those colored blobs will mean later because you could stop them in the middle of the process and run them again and get a somewhat different interpretation of those colored blobs. In AlphaFold 3, you actually have an interesting thing that if you look at AlphaFold 2, we can kind of, through this process of projecting out intermediate layers, see what it solves 1st. And it basically solves local details, local pieces. It starts to put local pieces together. It's agglomerative, in how it solves a structure as is kind of natural. The easiest thing to predict is your local structure. The hardest thing to predict is your largest scale structure. That's how alpha 2 works. If you look at alpha fold 3 and you take coordinates, which you've added a very large amount of noise to, well, the very 1st thing you have to solve is how, say, you have 2 proteins, how do they associate it? Where are their 2 blobs relative to each other? What are their Gaussians? So the the problem that AlphaFold 2 is solving last is the problem that AlphaFold three's diffusion has to realize 1st. And how does it do it? The answer is not that it comes up with an orientation and builds the protein around it because, of course, it's going for 1 correct answer or at least a very narrow distribution. The answer is really the big network before it plus the 1st pass through the diffusion network is solving the overall structure. And then the diffusion is realizing in any details it couldn't solve before, it's basically sampling among. So it is diffusion technically, but it's much closer to AlphaFold 2. I think there's no reason that it was kind of very specific technical reasons around kind of laziness and geometry that made diffusion a really good choice for Alpha Fold 3. It made it easier to handle ligands and handled some bond distances and local things. But it's not like diffusion in the same way as, oh, it's drawing the blobs and deciding what they mean at the end. So I think all of these are people like to think of these like to say, this works because it's a transformer. And this works because it's a transformer doesn't explain why chat models have gotten vastly better in the last 3, 4 years. It doesn't explain all the research. It doesn't explain what researchers do every day. All of these details are far more important than this high level bit of, is it a transformer, is it a diffusion model, that we wanna talk about. And then also even these diffusion mechanisms don't work in the way of kind of progressive refinement that makes sense for images. Right? Maybe you'll make colored blobs, and you'll decide what those colored blobs mean. Even that, I think you can argue, maybe not entirely the story, but it's definitely not the story for proteins because that's the hardest problem is the large scale structure. I mean, a sense, this is leaning towards this idea of constructive complexity that we were talking about before. And I'd love to get your your general take on on what this means for artificial general intelligence. Because with language models, for example, we train them basically with behavior cloning. So, you know, we have this this rich adaptive generative process and we we generate all of this language and we train language models on them. And for me, intelligence is the adaptive acquisition of coarse grained representations. Culture and language is changing all of the time. So we invent the word unalive to get around the filters on social media platforms and that's an example of ling linguistic agency. Language models, we noticed that when we do this iterative adaptive refining with active active fine tuning and adaptation, they become a bit intelligent. They they learn new representations and and they adapt. And in a way, what they're doing is even though they're ungrounded from the path, they can they can take a code solution like AlphaEvolve and they can refine it and they can refine it. And it seems to work really really well. But are we in this regime, do you think, that we're not necessarily building artifacts that have the same type of generality. I mean, what what do you think about intelligence in general? So this question of representations is very, very important and far less important than people believe 5 years ago in the explicit way. So just like we were talking about the things that AlphaFold does and the things that AlphaFold is forced to do by its code or or, you know, obviously, everything that's forced to do by its code, it does. But many things it does, it does without being forced. Because it had to learn it to make a good predictive model of the data. It had to find good intermediate representations. So in a certain sense, I think the most seductive idea in in machine learning is always there's this thing I know will have to be in the end there in the end, so I'm gonna have to have a u I'm gonna have to have a place in my code that is named that and then forces the mechanism to high level concept builder thingamajigger. Right? And that was a very popular kind of I'll make the concepts units. I shall force disentangled representations via this law sometimes on the intermediate layer. This is where it will store those. And that's reasonable to go test. But what we've seen a lot of is that a lot of the things that you would imagine needed or needed for intelligence are developed by desperately trying to predict the next token really, really well. And they're not they're not developed because you predict next tokens at all. They develop because you do a really, really good job at it. And so these kind of generalized spaces, representations, understanding of concepts is forced very slowly with data. Right? Pretty much all the kind of you know, there's a lot of log linears or my you know, everyone's least favorite functional is right. The things go up as the things go up linearly with the exponent of effort that we see all the time in our scaling laws. But we do get these concepts and representations, and what we don't really, I think, have an answer for is how do we get them cheaper? Now sometimes we can get them via programming. Right? We get memory like things. Now we have language models writing notes for itself and then retrieving those notes. So we find out it's better to keep reminding agents what they're doing so they don't forget over long trajectories. So we we build weights. We build artifacts. We find deficiencies. We can often paper over those deficiencies in some kind of software harnesses, but then we don't yet know we don't that doesn't immediately drive back…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence