Speakers in the public record
Claim mix
evaluation 9belief 9prediction 3recommendation 2commitment 2observation 1uncertainty 1
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
27 published records
“I highly recommend John Schulman’s “Generalized Advantage Estimation” paper as a good treatment on how to think about various ways to compute it.”
- Publisher
- Dwarkesh Podcast
“There were many algorithmic ideas applied, and then you can see that with modern Blackwell GPUs and Ada-class GPUs—which are much better than the V100-grade GPUs that that paper used—some of these algorithmic tricks to speed up convergence just don’t matter so much compared to something else.”
- Publisher
- Dwarkesh Podcast
“However, the problem is you spend most of training in this regime, in the low pass rate regime.”
- Publisher
- Dwarkesh Podcast
“I will say for the audience, for previous episodes when I was prepping and it seemed relevant to understand how AlphaGo works, I would find it very confusing.”
- Publisher
- Dwarkesh Podcast
“Always pick something that works and then get it to do something better, rather than start from something that doesn’t work at all and try to make it work.”
- Publisher
- Dwarkesh Podcast
“We care about this broader ability to do economically useful work, which is not super easy to measure, at least until you automate everything.”
- Publisher
- Dwarkesh Podcast
“I think there are a couple of interesting follow-on questions. There are questions on the inner loop and the outer loop.”
- Publisher
- Dwarkesh Podcast
“I think the most profound thing here is that a 10-layer neural network pass, basically 10 steps of reasoning… Of course, the reasoning is not just one trail of thought.”
- Publisher
- Dwarkesh Podcast
“I think there’s something quite profound there, and probably underexplored just because Go has been relatively underexplored compared to the boom in LLMs.”
- Publisher
- Dwarkesh Podcast
“I think the question we should be asking ourselves is about how we’ve been formulating solutions to NP-hard problems in worst-case complexity.”
- Publisher
- Dwarkesh Podcast
“I think that’s where a lot of the lateral thinking ability of humans comes from: knowing when the track you’re pursuing or the objective you’re pursuing is not right, and you should be asking a different question.”
- Publisher
- Dwarkesh Podcast
“Don’t LLMs natively learn to do MCTS, where they’ll try an approach and be like, “Oh, that doesn’t work.”
- Publisher
- Dwarkesh Podcast
“I think there’s a rich library of subtasks and sub-environments that you can train an automated scientist to work on, with Go as a sort of outer verification loop.”
- Publisher
- Dwarkesh Podcast
“Certainly, LLMs manage to do something that looks like real human reasoning without having to do an explicit tree structure. That being said, I think the idea of doing forward search and simulation to get a better sense of what is valuable might make a comeback, even if not in exactly the same instantiation as AlphaGo.”
- Publisher
- Dwarkesh Podcast
“While you’re erasing, another thing that was important for me to understand was about the MCTS data structure with nodes and children of nodes.”
- Publisher
- Dwarkesh Podcast
“You want to make sure that, similar to how in this case we were training on a lot of neutral labels, you’re sort of penalizing the labels that don’t help and only rewarding the ones that actually make you better.”
- Publisher
- Dwarkesh Podcast
“I would say the jury is probably still out on what the final instantiation of reasoning for LLMs will look like.”
- Publisher
- Dwarkesh Podcast
“I think automated scientific research is one of the most exciting skills that the frontier labs are developing right now.”
- Publisher
- Dwarkesh Podcast
“Because ln(n) grows slower than n, over time you will move from the argmax being dominated by the exploration term, which is the second term here, to the argmax being dominated by the Q term, which is when you’ve done enough simulations and are confident that this is the branch to go down.”
- Publisher
- Dwarkesh Podcast
“For a long time, we made the assumption that if we can’t model the dynamics with a world model, we’re going to instead just collect trajectories and then plan with respect to the only number that really matters, which is reward.”
- Publisher
- Dwarkesh Podcast
“For small data regimes, my experience is that ResNets still outperform transformers and give you more bang for the buck at lower budgets.”
- Publisher
- Dwarkesh Podcast
“In current LLM RL, they treat this entire sequence as a single action, at, and big T is just one. It’s true that because of how transformers are formulated through the product of conditional probabilities, the log probability of the whole sequence is equal to the sum of the probabilities of individual tokens.”
- Publisher
- Dwarkesh Podcast
“I was thinking, “Let’s see if the Bitter Lesson had happened, where a lot of these tricks just go away because Nvidia made faster GPUs.”
- Publisher
- Dwarkesh Podcast
“So Πθ (a | s) is great. But somehow, maybe because we’re playing a lot of games where the bots just resign instead of playing all the way to the Tromp-Taylor resolution, they forget how to evaluate those late-stage plans.”
- Publisher
- Dwarkesh Podcast
“I also don’t know if this one is proved to have logarithmically or square-root-bounded regret, but I think the algorithm was derived to look something like this.”
- Publisher
- Dwarkesh Podcast
“If you just take this policy recommendation and take the argmax over these probabilities—if you take the argmax and just take this action as your Go play—it’ll be a very fast Go player that doesn’t think in terms of reasoning steps.”
- Publisher
- Dwarkesh Podcast
“Experts might decide to end the game well before that, but under Tromp-Taylor scoring, you actually have to play things all the way to the end.”
- Publisher
- Dwarkesh Podcast