High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Eric Jang: belief

15 May 2026 Dwarkesh Podcast Eric Jang – Building AlphaGo from scratch

“I would say the jury is probably still out on what the final instantiation of reasoning for LLMs will look like.”

— Eric Jang

Source trail

Everything needed to verify it.

Speaker
Eric Jang
Attribution
Verified speaker
Claim type
belief
Recorded
15 May 2026
Publisher
Dwarkesh Podcast

Transcript context

…Right, you get supervision through a straw. You basically just upgrade all the tokens in a trajectory that might or might not have been relevant to getting the answer right. The reason you can do this much more sample-efficient, favorable thing with Go is that because MCTS works in Go, you basically know that if you just do the search locally here—and this search is truncated at the end by a value function that works even if you haven’t unfolded your whole trajectory—you can just say, “This is my new policy,” and improve in a more iterative, local way, rather than having to unfold all these trajectories. There was some research from Google in 2023 or 2024 where they tried to apply tree structures to reasoning. The jury is still out as to whether this can ever work. We probably will see a revisiting of this idea of forward search in the future. But there are two things that make MCTS very simple for Go. Value estimation is concrete. You can determine it for real, and then you can use it to truncate depth, as you said. The breadth is also determined. What’s critical is that the action selection algorithm, where you iteratively visit and grow the tree, is well suited for the size and depth of problem that Go is. But for something like LLM reasoning, PUCT might not be a good enough heuristic. It might be too greedy with local tokens, and it might only give you obvious thoughts that are correct but don’t really solve your final problem. I would say the jury is probably still out on what the final instantiation of reasoning for LLMs will look like. I wouldn’t rule out that this stuff could come back, but it’s a bit hard. Don’t LLMs natively learn to do MCTS, where they’ll try an approach and be like, “Oh, that doesn’t work. Let’s back up. Let’s try this other thing,” and then go in the direction that proves to be more fruitful?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence