High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / preference

Published · transcript-backed

Speaker unverified: preference

22 Jun 2026 Machine Learning Street Talk He won a Nobel here for AlphaFold. Then he left. - John Jumper

“But after AlphaFold 1 and then kind of all the protein specific bits were kind of wrapped around the machine learning. And so the I would say alpha fold 2 was, let's build the science instead of building the science of image recognition and then applying it to proteins because, you know, human visual the human visual system is exactly what we needed to fold proteins is not something true.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
preference
Recorded
22 Jun 2026
Publisher
Machine Learning Street Talk

Transcript context

…kind of a narrow predictor. But we are doing something truly useful. Can we talk through the predictive architectures of of the different versions of AlphaFold? So, you know, the 1st version was was a CNN. The last version is a diffusion model. The 2nd version, we spoke about this last night, it had a structure component. And, you know, obviously, geometric deep learning is spoken about a lot. And I think people misattributed the the benefit of having these, you know, and kind of symmetries. It did these SC 3 symmetries. And just just talk me through that process because you were kind of saying at the beginning, you were really trying to imbue your human understanding of this and then kind of experience told you differently. I think there's 2 or 3 things. I mean, 1st, I would almost object not technically, but kind of thematically to AlphaFold 3 as a diffusion model. We love to stick things in boxes. We love to have the highest level bit be the answer for why these things work. Yes. Right? Oh, they switched from CNN to I think the answer is really okay. AlphaFold 1 really was, as a network, it predicted a subpart of the problem. It started from kind of biological data, evolutionary correlations. It ended in kind of geometric ish data, distance between atoms. In between was a CNN. Right? It was actually an off the shelf CNN from a computer vision that someone else had done. Okay. That was a CNN. But after AlphaFold 1 and then kind of all the protein specific bits were kind of wrapped around the machine learning. And so the I would say alpha fold 2 was, let's build the science instead of building the science of image recognition and then applying it to proteins because, you know, human visual the human visual system is exactly what we needed to fold proteins is not something true. Right? Humans were bad, are bad at at predicting protein structures. How are we going to actually build all the pieces? Now there was an s e 3 piece. In fact, AlphaFold 2 was built iteratively. There were many stages. Actually, the s e 3 piece was the 1st part of AlphaFold built. But AlphaFold 2 at the end was really this giant trunk of an architecture we called Evoformer, which is axial attention plus a bunch of other stuff. And that is 90 plus percent of the compute and the accuracy. And then but it produces this kind of n by n. So you start off with 2 pieces of data. You start off with the protein sequence, and then you find the sequence of every protein evolutionarily related. And protein structure changes slowly. The the structures of my proteins are in most cases similar to the structure of proteins in yeast, sometimes even out in E. Coli. So you grab many related structures. You often have hundreds or thousands. You provide this information, and we have this this specialty architecture called Evoformer, which had 2 forms of axial attention that were kind of having a conversation between what we believed about geometry and what we believed about evolution. We had these 2 representations. And we end up we take the geometric bit, the n by n, which we have actually as an intermediate loss said, what are the distances between these atoms and made categorical predictions. And then we hand it to what we call the structure module, and it's best thought of as a geometrization engine. Right? If you you have n squared predictions about n positions, you're somebody's gonna have to harmonize this thing. And this used a s e 3 I guess it was invariant in the sense we collapsed it on every layer. S e 3 invariant attention. This was actually 1. This was 1 of mine ody's gonna have to harmonize this thing. And this used a s e 3 I guess it was invariant in the sense we collapsed it on every layer. S e 3 invariant attention. This was actually 1. This was 1 of mine that was kind of even starting at DeepMine. I'm like, oh, we should probably put points in. Or I was thinking about pro even then, protein residues. Right? So you have this backbone, which has 3 atoms. You can align a frame to it. It's extraordinarily rigid. And I know the business in the places where all these at where all these residues differ is kind of just off that. So if you align them to reference frames then and you operate in points in those reference frames, then it's natural. And you can just take attention, and you can let it project points in its local frame. You can transform it. Then you can use the distance of those points as a way in order to bias your attention. And this is, invariant point attention is what we called it. At the end, it's kind of fun. More important than that probably was this defining of frames. Actually, almost certainly. Defining of frames let us write down a really interesting loss function. So we call it frame align, point error, or FAPE. And this was kind of saying, in the reference frame of the I th residue, where is everyone else? And it's kind of locally registered, and then you have n squared errors, and then you average them together. That, I think, was really, really important. I think that was 1 of the breakthroughs early on was this loss. But, of course, the really fun part is an SE 3 invariance. And but remember, we didn't start with any geometric data. We started only with nongeometric data. So our geometry emerged in the middle. We started with what we would call black hole initialization, where we just stick all the residues on top of each other in the world's least physical structure, it was important also that we disrespected known symmetries of a protein. For example, the known symmetries of a protein are these residues are separated. The atoms this atom and that atom in a residue are separated by 1.3 angstroms plus or minus 0.015. And in fact, even in AlphaFold 1, when we would do the optimization, we would actually use turning kind of like a jointed robot arm to optimize these, say, typically 300 residues. And so you would do all this twisting. You would actually have a very ugly geometry. The geometry of a 300 joint or actually, no. Sorry. Wouldn't be 300. It would be a 900 joint robot arm is really bad, and so that means that your optimizer has to take many steps. So 1 of the important things is let's just break it up. Let's just treat them as a residue gas, we called it, so that this can proceed in, say, 4 steps, 8 steps instead of the number of steps of this twisty geometry. And then we used equivariance, and it helped. But 1 of the things that was really surprising, think, is maybe because of the early talk or maybe people were working on equivariance. Geometric deep learning has been very popular, and people said, ah, they mentioned my keyword. That must be the reason it worked. And I remember being a little bit confused, and I thought, okay. But we'll we'll…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence