Evidence receipt / belief
Published · transcript-backedSpeaker unverified: belief
22 Jun 2026 Machine Learning Street Talk He won a Nobel here for AlphaFold. Then he left. - John Jumper
“I think AlphaFold 3 diffusion is similar, and it's especially similar because, in fact, in images, okay, you start generating an image and you see especially these early trained diffusion models generate kind of colored blobs, and they start to decide what those colored blobs mean.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- belief
- Recorded
- 22 Jun 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…ration. And then we put in a kind of code idea, architectural idea to help this process that it was gonna learn from the data. Going back to the earlier thing about exactly how far residues are apart, we didn't tell AlphaFold that. We knew that the data would scream at it, that I and I plus 1 were 1.3 angstroms apart. So I think when we think about our human understanding, I think 1 of the you know, I don't really love the bitter lesson as people try and apply it. In fact, AlphaFold 2 is the opposite of that. We did a whole bunch of specialty stuff because our data is not finite. And in fact, now that we've gone to language models, we found our data is still finite. The Internet is finite. So I think, you know, don't do architectural research is the wrong thing to draw from it. But have some humility about which things go into your code and which things will be derived from your data. Look at what's missing. Understand the algorithm that deep learning that the deep learning is trying to learn, how can you accelerate it, how can you add hypotheses, and where you especially you add kind of communication. The most important thing we would do within the architecture is modify which units communicated and how. I think all of these have been kind of how we drive understanding to ultimately make an iterative process. And it should shock no 1 that if you're trying to make an intricate geometric object that you are going to iterate. Or similarly, if you think about generating text. Right? And 1 kind of naive assumption that people will make is that these are next word generators, so they have no idea what's gonna happen in 2 words ahead or 3 words ahead. But, of course, you can't think of the you can't write down the next word without I don't I don't start a sentence not knowing how it's gonna end most of the time. Right? At some points, I change, but I think ahead a little bit in order to accomplish my task. And so the understanding that we see built into these models are kind of the structures that we want sometimes emerge and sometimes don't. And I think we valorize the high level ideas that impose, for example, an AlphaFold 3 coming back to AlphaFold 3. Right? You said it is a diffusion model. But I would argue it's a different diffusion model than an image model. Maybe well, there's some different for 1 thing, there's a huge trunk that is not in any way a diffusion model that's only run once. That trunk is probably where the structure is actually determined, and the diffusion is just like the structure module was a geometrization engine that took a set of really quite good constraints that had very clear notion of the structure within those constraints and solved up the details. I think AlphaFold re module was a geometrization engine that took a set of really quite good constraints that had very clear notion of the structure within those constraints and solved up the details. I think AlphaFold 3 diffusion is similar, and it's especially similar because, in fact, in images, okay, you start generating an image and you see especially these early trained diffusion models generate kind of colored blobs, and they start to decide what those colored blobs mean. And they pretty clearly kind of decide what those colored blobs will mean later because you could stop them in the middle of the process and run them again and get a somewhat different interpretation of those colored blobs. In AlphaFold 3, you actually have an interesting thing that if you look at AlphaFold 2, we can kind of, through this process of projecting out intermediate layers, see what it solves 1st. And it basically solves local details, local pieces. It starts to put local pieces together. It's agglomerative, in how it solves a structure as is kind of natural. The easiest thing to predict is your local structure. The hardest thing to predict is your largest scale structure. That's how alpha 2 works. If you look at alpha fold 3 and you take coordinates, which you've added a very large amount of noise to, well, the very 1st thing you have to solve is how, say, you have 2 proteins, how do they associate it? Where are their 2 blobs relative to each other? What are their Gaussians? So the the problem that AlphaFold 2 is solving last is the problem that AlphaFold three's diffusion has to realize 1st. And how does it do it? The answer is not that it comes up with an orientation and builds the protein around it because, of course, it's going for 1 correct answer or at least a very narrow distribution. The answer is really the big network before it plus the 1st pass through the diffusion network is solving the overall structure. And then the diffusion is realizing in any details it couldn't solve before, it's basically sampling among. So it is diffusion technically, but it's much closer to AlphaFold 2. I think there's no reason that it was kind of very specific technical reasons around kind of laziness and geometry that made diffusion a really good choice for Alpha Fold 3. It made it easier to handle ligands and handled some bond distances and local things. But it's not like diffusion in the same way as, oh, it's drawing the blobs and deciding what they mean at the end. So I think all of these are people like to think of these like to say, this works because it's a transformer. And this works because it's a transformer doesn't explain why chat models have gotten vastly better in the last 3, 4 years. It doesn't explain all the research. It doesn't explain what researchers do every day. All of these details are far more important than this high level bit of, is it a transformer, is it a diffusion model, that we wanna talk about. And then also even these diffusion mechanisms don't work in the way of kind of progressive refinement that makes sense for images. Right? Maybe you'll make colored blobs, and you'll decide what those colored blobs mean. Even that, I think you can argue, maybe not entirely the story, but it's definitely not the story for proteins because that's the hardest problem is the large scale structure. I mean, a sense, this is leaning towards this idea of constructive complexity that we were talking about before. And I'd love to get your your general take on on what this means for artificial general intelligence. Because with language models, for example, we train them basically with behavior cloning. So, you know, we have this this rich adaptive generative process and we we generate all of this language and we train language models on them. And for me, intelligence is the adaptive acquisition of coarse grained representations. Culture and language is changing all of the time. So we invent the word unalive to get around the filters on social media platforms and that's an example of ling linguistic agency. Language models, we noticed that when we do this iterative adaptive refining with active active fine tuning and adaptation, they become a bit intelligent. They they learn new representations and and they adapt. And in a way, what they're doing is even though they're ungrounded from the path, they can they can take a code solution like AlphaEvolve and they can refine it and they can refine it. And it seems to work really really well. But are we in this regime, do you think, that we're not necessarily building artifacts that have the same type of generality. I mean, what what do you think about intelligence in general? So this question of representations is very, very important…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.