Evidence receipt / belief
Published · transcript-backedSholto Douglas: belief
22 May 2025 Dwarkesh Podcast Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken
“I think at the beginning, the hill to climb. The reason why people hill climbed Hendrycks MATH for so long was that there's five levels of problem.”
Source trail
Everything needed to verify it.
- Speaker
- Sholto Douglas
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 22 May 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…In general, when you're making either benchmarks or environments where you're trying to grade the model or have it improve or hill climb on some metric, do you care more about resolution at the top end? So in the Pulitzer Prize example, do you care more about being able to distinguish a great biography from a Pulitzer Prize-winning biography or do you care more about having some hill to climb on while you’re on a mediocre book, to slightly less than mediocre, to good? Which one is more important? I think at the beginning, the hill to climb. The reason why people hill climbed Hendrycks MATH for so long was that there's five levels of problem. It starts off reasonably easy. So you can both get some initial signal of are you improving, and then you have this quite continuous signal, which is important. Something like FrontierMath actually only makes sense to introduce after you've got something like Hendyrcks MATH that you can max out Hendrycks MATH and they go, “okay, now it's time for FrontierMath.” How does one get models to output less slop? What is the benchmark or the metric? Why do you think they will be outputting less slop in a year?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.