Evidence receipt / preference
Published · transcript-backedJacob Kimmel: preference
21 Aug 2025 Dwarkesh Podcast Evolution designed us to die fast; we can change that — Jacob Kimmel
“” We only work on a few cell types at NewLimit because those are some of the only cell types today with which we believe we can get really effective delivery of medicines.”
Source trail
Everything needed to verify it.
- Speaker
- Jacob Kimmel
- Attribution
- Verified speaker
- Claim type
- preference
- Recorded
- 21 Aug 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…If it actually is the case that this is very similar to the way LLM dynamics work, then once this technology is mature and you get the GPT-3 equivalent of the virtual cell, what you would expect to happen is you get many different companies, at least a couple, that are doing these cheap, Perturb-seq-like experiments and building their own virtual cells. Then they're leasing this out to other people who then have their own ideas about, "We want to see if we can come up with the labels for this particular thing we care about and test for that." What it seems like is happening right now is, at least at NewLimit, you are like, "We know the end use case we're going after." It would be as if Cursor in 2018 was like, "We're going to build our own LLM from scratch so that we can enable our application," rather than some foundation model company being like, "We don't care what you use it for, we're going to build this." Does that make sense? It seems like you're combining two different layers of the stack. Because nobody else is doing the other layer, you're just doing both of them. I don't know to what extent this analogy maps on. To play with the analogy a bit, imagine that you think about NewLimit as an LLM company. If I'm going to put us in the shoes of Cursor, which oh I so wish, imagine we're trying to, in 2018, create Cursor Tab, but we're not trying to create a full LLM. I don't know enough about the underlying mechanics to know if that would have been feasible, but it's a much more feasible problem than trying to create the most recent Cursor agent or compete with modern Claude Code. That's roughly the equivalent. The problem we're breaking off is a subset of the more general virtual cell problem. We're trying to predict, “What do groups of transcription factors do to the age of very specific types of cells? ” We only work on a few cell types at NewLimit because those are some of the only cell types today with which we believe we can get really effective delivery of medicines. We think they're just more important because we can act on them today. If we solve the problem of what TFs to use, we can make a medicine pretty quickly. In a way, we're carving out a region of this massive parameter space and saying, "If we can learn the distribution of effects even just in this small region, it's going to be really effective for us, and we can make really amazing products, unlike the world has ever seen." Over time, we can expand to the corpus of predicting every possible gene perturbation in every possible cell type. I think that's maybe the way the analogy maps on, but it is true that we are vertically integrating here. We're generating our own data in a way that's proprietary. We think we have a much, much larger data set for this particular regime than the rest of the world combined. That enables us to build what we think are the best models. In many cases, what we found is that unlike with LLMs, where a lot of the data that was necessary to build these was a common good—it was produced as a function of the internet and shared across everyone, it's pretty common across all the domains everyone wants to use it for—this biological data is still in its infancy. Imagine we're in the early 1980s and we are just now thinking about trying to create some of the first web pages. That's the era we're in. We're going after and generating some of our own data in this very niche circumstance, building the very high-quality corpus, the Wikipedia that you might train your overly analogized- LLM on, and then building the first products based on that and then expanding from there. We think that's necessary because of where we are today. There isn't this Internet-like equivalent of data that everyone can go out and reap rewards from. Interesting. This is more a question about the broader pharma industry rather than just NewLimit. In the future, how are people going to make money? With the GLPs, we've got peptides from China that are just a gray market that people can easily consume. Presumably, with these future AI models, even if you have a patent on a molecule, finding an isomorphic molecule or an isomorphic treatment is relatively easy. If you do come up with these crazy treatments and if pharma in general is able to come up with these crazy treatments, will they be able to make money?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.