High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Demis Hassabis: belief

28 Feb 2024 Dwarkesh Podcast Demis Hassabis — Scaling, superhuman AIs, AlphaZero atop LLMs, AlphaFold

“If you improve the models, then I think your search can be more efficient and therefore you can get further with your search.”

— Demis Hassabis

Source trail

Everything needed to verify it.

Speaker
Demis Hassabis
Attribution
Verified speaker
Claim type
belief
Recorded
28 Feb 2024
Publisher
Dwarkesh Podcast

Transcript context

…How do you get past the immense amount of compute that these approaches tend to require? Even the AlphaGo system was a pretty expensive system because you sort of had to run an LLM on each node of the tree. How do you anticipate that’ll get made more efficient? One thing is Moore’s law tends to help. Over every year more computation comes in. But we focus a lot on sample-efficient methods and reusing existing data, things like experience replay and also just looking at more efficient ways. The better your world model is, the more efficient your search can be. One example I always give is AlphaZero, our system to play Go and chess and any game. It’s stronger than human world champion level in all these games and it uses a lot less search than a brute force method like Deep Blue to play chess. One of these traditional Stockfish or Deep Blue systems would maybe look at millions of possible moves for every decision it’s going to make. AlphaZero and AlphaGo may look at around tens of thousands of possible positions in order to make a decision about what to move next. A human grandmaster or world champion probably only looks at a few hundred moves, even the top ones, in order to make their very good decision about what to play next. So that suggests that the brute force systems don’t have any real model other than the heuristics about the game. AlphaGo has quite a decent model but the top human players have a much richer, much more accurate model of Go or chess. That allows them to make world-class decisions on a very small amount of search. So I think there’s a sort of trade-off there. If you improve the models, then I think your search can be more efficient and therefore you can get further with your search. I have two questions based on that. With AlphaGo, you had a very concrete win condition: at the end of the day, do I win this game of Go or not? You can reinforce on that. When you’re thinking of an LLM putting out thought, do you think there will be this ability to discriminate in the end, whether that was a good thing to reward or not?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence