app / uses
Claude
“We have some very early agents that are extremely impressive and that I use daily—Claude and Codex and so on—but I still feel there’s so much work to be done.”
Public evidence record
Researcher · Eureka Labs
Books, apps, and tools
app / uses
“We have some very early agents that are extremely impressive and that I use daily—Claude and Codex and so on—but I still feel there’s so much work to be done.”
app / uses
“We have some very early agents that are extremely impressive and that I use daily—Claude and Codex and so on—but I still feel there’s so much work to be done.”
course / built
“Earlier on, I built CS231n at Stanford, which I think was the first deep learning class at Stanford, which became very popular.”
Claim ledger
7 transcript-backed records
01 / evaluation
“I guess they just don’t work as well empirically because right now the models are collapsed.”
02 / evaluation
“The paper that blew my mind was InstructGPT, because it pointed out that you can take the pretrained model, which is autocomplete, and if you just fine-tune it on text that looks like conversations, the model will very rapidly adapt to become very conversational, and it keeps all the knowledge from pre-training.”
03 / evaluation
“Just a lot of it. So maybe one way to think about it, I don’t know if this is the best way, but I almost feel like — again, making these analogies imperfect as they are — we’ve stumbled by with the transformer neural network, which is extremely powerful, very general.”
04 / evaluation
“” The Atari deep reinforcement learning shift in 2013 or so was part of that early effort of agents, in my mind, because it was an attempt to try to get agents that not just perceive the world, but also take actions and interact and get rewards from environments.”
05 / evaluation
“In fact, I think the paper was even stronger because they hardcoded the weights of a neural network to do gradient descent through attention and all the internals of the neural network.”
06 / evaluation
“We have some very early agents that are extremely impressive and that I use daily—Claude and Codex and so on—but I still feel there’s so much work to be done.”
07 / evaluation
“I would say nanochat is not an example of those because it’s a fairly unique repository.”