High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / commitment

Published · transcript-backed

Eric Jang: commitment

15 May 2026 Dwarkesh Podcast Eric Jang – Building AlphaGo from scratch

“For a long time, we made the assumption that if we can’t model the dynamics with a world model, we’re going to instead just collect trajectories and then plan with respect to the only number that really matters, which is reward.”

— Eric Jang

Source trail

Everything needed to verify it.

Speaker
Eric Jang
Attribution
Verified speaker
Claim type
commitment
Recorded
15 May 2026
Publisher
Dwarkesh Podcast

Transcript context

…Earlier I was like, “Hey, why are we training policy? Why don’t we just train the value alone?” That is what this is. This is an algorithm for recovering value estimates of intermediate steps when you don’t have the ability to do forward search. You must collect a trajectory first of n steps before you’re able to do this trick. But the intuition is kind of the same. Knowing something about the Q value here can tell you something about the Q value here. And indeed, you can recover a policy from a Q value. You don’t need to explicitly model the policy distribution. You can recover the policy distribution by doing argmax over your Q values. Q-learning, or approximate dynamic programming, propagates what you know about the future Qs backward like this. You can see that there’s a similar structure going on here, where in one case you’re planning over trajectories your agent hasn’t been to yet, whereas in the other you’re planning over trajectories your agent has visited. Importantly, why was Q-learning a big deal? It’s because historically we just haven’t had the ability to do search on fairly high-dimensional problems like robotics or whatever. For a long time, we made the assumption that if we can’t model the dynamics with a world model, we’re going to instead just collect trajectories and then plan with respect to the only number that really matters, which is reward. This is very interesting. To unify this with our discussion of LLMs, you don’t have Q values, but you’re doing this backwards learning where you find the trajectories that pass some unit test in some coding environment and then reinforce those trajectories. There’s a huge difference between that and this forward approach with MCTS. The reason it’s much more preferable to do MCTS is because you can do it per move and make each move better, rather than having to learn per trajectory and hope, as Karpathy said, to learn this……

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence