Evidence receipt / belief
Published · transcript-backedDwarkesh Patel: belief
22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken
“I think there's a paper from Tsinghua University, where they showed that if you give a base model enough tries to answer a question, it can still answer the question as well as the reasoning model.”
Source trail
Everything needed to verify it.
- Speaker
- Dwarkesh Patel
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 22 May 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…This was just over a conversation. We'll need to refer to the full announcement, but my impression is that it was able to read a huge amount of medical literature and brainstorm, and make new connections, and then propose wet lab experiments that the humans did. Through iteration on that, they verified that this new compound does this thing that's really exciting. Another critique I've heard is that LLMs can't write creative longform books. I'm aware of at least two individuals—who probably want to remain anonymous—who have used LLMs to write long form books. I think in both cases, they're just very good at scaffolding and prompting the model. Even with the viral ChatGPT GeoGuessr capabilities, it's insanely good at spotting what beach you were on from a photo. Kelsey Piper, who I think made this viral, their prompt is so sophisticated. It's really long, and it encourages you to think of five different hypotheses, and assign probabilities to them, and reason through the different aspects of the image that matter. I haven't A/B tested it, but I think unless you really encourage the model to be this thoughtful, you wouldn't get the level of performance that you see with that ability. You're bringing up ways in which people have constrained what the model is outputting to get the good part of the distribution. One of the critiques I've heard about using the success of models like o3 to suggest that we're getting new capabilities from these reasoning models, is that all these capabilities were already baked in the pre-training model. I think there's a paper from Tsinghua University, where they showed that if you give a base model enough tries to answer a question, it can still answer the question as well as the reasoning model. It basically just has a lower probability of answering correctly. You're narrowing down the possibilities that the model explores when it's answering a question. Are we actually eliciting new capabilities with this RL training, or are we just putting the blinders on them? Right, like carving away the marbles on this. I think it's worth noting that that paper was, I'm pretty sure, on the Llama and Qwen models. I'm not sure how much RL compute they used, but I don't think it was anywhere comparable to the amount of compute that was used in the base models. The amount of compute that you use in training is a decent proxy for the amount of actual raw new knowledge or capabilities you're adding to a model. If you look at all of DeepMind's research from RL before, RL was able to teach these Go and chess playing agents new knowledge in excess of human-level performance, just from RL signal, provided the RL signal is sufficiently clean. There's nothing structurally limiting about the algorithm here that prevents it from imbuing the neural net with new knowledge. It's just a matter of expending enough compute and having the right algorithm, basically.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.