High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / recommendation

Published · transcript-backed

Dylan Patel: recommendation

3 Feb 2025 Lex Fridman Podcast #459 – DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters

“Maybe we use some sort of reward model outside of this to select even the best one to preference, as well.”

— Dylan Patel

Source trail

Everything needed to verify it.

Speaker
Dylan Patel
Attribution
Verified speaker
Claim type
recommendation
Recorded
3 Feb 2025
Publisher
Lex Fridman Podcast

Transcript context

…Scientific discovery, like when you use this sort of reasoning problem in it? Just something we fully don’t expect. I think it’s actually probably simpler than that. It’s probably something related to computer use or robotics rather than science discovery. Because the important aspect here is models take so much data to learn. They’re not sample efficient. Trillions. They take the entire web, over 10 trillion tokens to train on. This would take a human thousands of years to read. A human does not… And humans know most of the stuff, a lot of the stuff models know better than it, right? Humans are way, way, way more sample efficient. That is because of the self-play, right? How does a baby learn what its body is as it sticks its foot in its mouth and it says, “Oh, this is my body, right?” It sticks its hand in its mouth and it calibrates its touch on its fingers with the most sensitive touch thing on its tongue is how babies learn and it’s just self-play over and over and over and over again. And now we have something that is similar to that with these verifiable proofs, whether it’s a unit testing code or a mathematical verifiable task, generate many traces of reasoning and keep branching them out, keep branching them out, and then check at the end, hey, which one actually has the right answer? Most of them are wrong. Great. These are the few that are right. Maybe we use some sort of reward model outside of this to select even the best one to preference, as well. But now you’ve started to get better and better at these benchmarks. And so you’ve seen over the last six months a skyrocketing in a lot of different benchmarks. All math and code benchmarks were pretty much solved except for frontier math, which is designed to be almost questions that aren’t practical to most people. They’re exam-level, open math problem-type things. So it’s like on the math problems that are somewhat reasonable, which is somewhat complicated word problems or coding problems, is just what Dylan is saying.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence