paper / recommends
Generalized Advantage Estimation
“I highly recommend John Schulman’s “Generalized Advantage Estimation” paper as a good treatment on how to think about various ways to compute it.”
Public evidence record
Published podcast speaker
Books, apps, and tools
paper / recommends
“I highly recommend John Schulman’s “Generalized Advantage Estimation” paper as a good treatment on how to think about various ways to compute it.”
Claim ledger
7 transcript-backed records
01 / belief
“I think the most profound thing here is that a 10-layer neural network pass, basically 10 steps of reasoning… Of course, the reasoning is not just one trail of thought.”
02 / belief
“I think there’s something quite profound there, and probably underexplored just because Go has been relatively underexplored compared to the boom in LLMs.”
03 / belief
“I think the question we should be asking ourselves is about how we’ve been formulating solutions to NP-hard problems in worst-case complexity.”
04 / belief
“I think that’s where a lot of the lateral thinking ability of humans comes from: knowing when the track you’re pursuing or the objective you’re pursuing is not right, and you should be asking a different question.”
05 / belief
“I think there’s a rich library of subtasks and sub-environments that you can train an automated scientist to work on, with Go as a sort of outer verification loop.”
06 / belief
“Certainly, LLMs manage to do something that looks like real human reasoning without having to do an explicit tree structure. That being said, I think the idea of doing forward search and simulation to get a better sense of what is valuable might make a comeback, even if not in exactly the same instantiation as AlphaGo.”
07 / belief
“I would say the jury is probably still out on what the final instantiation of reasoning for LLMs will look like.”