paper / recommends
Generalized Advantage Estimation
“I highly recommend John Schulman’s “Generalized Advantage Estimation” paper as a good treatment on how to think about various ways to compute it.”
Public evidence record
Published podcast speaker
Books, apps, and tools
paper / recommends
“I highly recommend John Schulman’s “Generalized Advantage Estimation” paper as a good treatment on how to think about various ways to compute it.”
Claim ledger
20 transcript-backed records
01 / recommendation
“I highly recommend John Schulman’s “Generalized Advantage Estimation” paper as a good treatment on how to think about various ways to compute it.”
02 / evaluation
“There were many algorithmic ideas applied, and then you can see that with modern Blackwell GPUs and Ada-class GPUs—which are much better than the V100-grade GPUs that that paper used—some of these algorithmic tricks to speed up convergence just don’t matter so much compared to something else.”
03 / evaluation
“Always pick something that works and then get it to do something better, rather than start from something that doesn’t work at all and try to make it work.”
04 / belief
“I think the most profound thing here is that a 10-layer neural network pass, basically 10 steps of reasoning… Of course, the reasoning is not just one trail of thought.”
05 / belief
“I think there’s something quite profound there, and probably underexplored just because Go has been relatively underexplored compared to the boom in LLMs.”
06 / belief
“I think the question we should be asking ourselves is about how we’ve been formulating solutions to NP-hard problems in worst-case complexity.”
07 / belief
“I think that’s where a lot of the lateral thinking ability of humans comes from: knowing when the track you’re pursuing or the objective you’re pursuing is not right, and you should be asking a different question.”
08 / belief
“I think there’s a rich library of subtasks and sub-environments that you can train an automated scientist to work on, with Go as a sort of outer verification loop.”
09 / belief
“Certainly, LLMs manage to do something that looks like real human reasoning without having to do an explicit tree structure. That being said, I think the idea of doing forward search and simulation to get a better sense of what is valuable might make a comeback, even if not in exactly the same instantiation as AlphaGo.”
10 / evaluation
“You want to make sure that, similar to how in this case we were training on a lot of neutral labels, you’re sort of penalizing the labels that don’t help and only rewarding the ones that actually make you better.”
11 / belief
“I would say the jury is probably still out on what the final instantiation of reasoning for LLMs will look like.”
12 / prediction
“I think automated scientific research is one of the most exciting skills that the frontier labs are developing right now.”
13 / commitment
“For a long time, we made the assumption that if we can’t model the dynamics with a world model, we’re going to instead just collect trajectories and then plan with respect to the only number that really matters, which is reward.”
14 / evaluation
“For small data regimes, my experience is that ResNets still outperform transformers and give you more bang for the buck at lower budgets.”
15 / evaluation
“In current LLM RL, they treat this entire sequence as a single action, at, and big T is just one. It’s true that because of how transformers are formulated through the product of conditional probabilities, the log probability of the whole sequence is equal to the sum of the probabilities of individual tokens.”
16 / prediction
“I was thinking, “Let’s see if the Bitter Lesson had happened, where a lot of these tricks just go away because Nvidia made faster GPUs.”
17 / evaluation
“So Πθ (a | s) is great. But somehow, maybe because we’re playing a lot of games where the bots just resign instead of playing all the way to the Tromp-Taylor resolution, they forget how to evaluate those late-stage plans.”
18 / uncertainty
“I also don’t know if this one is proved to have logarithmically or square-root-bounded regret, but I think the algorithm was derived to look something like this.”
19 / recommendation
“If you just take this policy recommendation and take the argmax over these probabilities—if you take the argmax and just take this action as your Go play—it’ll be a very fast Go player that doesn’t think in terms of reasoning steps.”
20 / evaluation
“Experts might decide to end the game well before that, but under Tromp-Taylor scoring, you actually have to play things all the way to the end.”