Evidence receipt / belief
Published · transcript-backedJacob Kimmel: belief
21 Aug 2025 Dwarkesh Podcast Evolution designed us to die fast; we can change that — Jacob Kimmel
“We don't actually use RL at the moment, so I don't want to overstate the level of sophistication we've got. But I think the general problem reduces down in a similar way.”
Source trail
Everything needed to verify it.
- Speaker
- Jacob Kimmel
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 21 Aug 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…This is so similar to, in LLMs, you first have imitation learning with pre-training that builds a general-purpose representation of the world. Then you do RL about a particular objective in math or coding or whatever that you care about. You are describing an extremely similar procedure where first you just learn to predict perturbations in genes to broad effects on the cell. That's the pre-training, just learning how cells work. Then there's another afterward layer of these value judgments of, “How would we have to perturb it to have effect X?” That actually seems very similar to “How do we get the base model to answer this math problem or answer this coding problem?” I don't know if people usually put it this way, but it actually just seems extremely similar. That makes me more optimistic on this. LLMs work and RL works. Yeah, they do. I think the conceptual analogy is very apt. We don't actually use RL at the moment, so I don't want to overstate the level of sophistication we've got. But I think the general problem reduces down in a similar way. You can think about your earlier question of what does the general model look like that enables you to actually have compounding returns in drug discovery. You might have something like this base model, which as you said, just predicts this object function of, “How are these perturbations hitting these targets going to change which genes are turned on and off in this cell?” Then there's an entirely other task, which is, well, which genes do you want to turn on and off? What state do I want the cell to adopt? Our lens on that is that across many different diseases people have, age is one of the strongest predictors of how they're going to progress, whether that disease arises. In many, many circumstances you have evidence in humans where you can say, “Ah, if I could make the cell younger, maybe that's not a perfect fix, but that's going to dramatically benefit not only patients who have a diagnosed disease, but it might actually help most of us stay healthier longer, even subclinically before anyone would formally say that we're sick.” Now that's another more general function. The same way that in LLMs, you might have to create these particular RLHF environments, you need to have places where you can state a value function of the particular task that you're trying to optimize for. In drug discovery, you would then need to know, “Well, what are the cell states I want to engineer for?” That's kind of the next generation of what a target might be. Beyond just which genes do I want to move up and down, and which gene perturbations do I put in, you then need to know what cell state am I engineering for? What do I want this T cell to do? You’ll have a bunch of labelers in Nigeria clicking different pictures of cells. Like, “Oh, this one looks young. This one looks old.”…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.