High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Andrej Karpathy

Researcher · Eureka Labs

Claims
38
Episodes
1
Shows
1
Named items
3

Books, apps, and tools

The evidenced stack.

Browse the grouped index →

app / uses

Claude

“We have some very early agents that are extremely impressive and that I use daily—Claude and Codex and so on—but I still feel there’s so much work to be done.”

Dwarkesh Podcast · 17 Oct 2025

Evidence receipt · Source ↗

app / uses

Codex

“We have some very early agents that are extremely impressive and that I use daily—Claude and Codex and so on—but I still feel there’s so much work to be done.”

Dwarkesh Podcast · 17 Oct 2025

Evidence receipt · Source ↗

course / built

CS231n

“Earlier on, I built CS231n at Stanford, which I think was the first deep learning class at Stanford, which became very popular.”

Dwarkesh Podcast · 17 Oct 2025

Evidence receipt · Source ↗

Claim ledger

What Andrej said.

38 transcript-backed records

05 / belief

I am betting a bit implicitly on some of the timelessness of human nature. It will be desirable to do all these things, and I think people will look up to it as they have for millennia.

“I am betting a bit implicitly on some of the timelessness of human nature. It will be desirable to do all these things, and I think people will look up to it as they have for millennia.”
Speaker
Andrej Karpathy
Publisher
Dwarkesh Podcast

06 / belief

I still think you’re presupposing some discrete jump that has no historical precedent that I can’t find in any of the statistics and that I think probably won’t happen.

“I still think you’re presupposing some discrete jump that has no historical precedent that I can’t find in any of the statistics and that I think probably won’t happen.”
Speaker
Andrej Karpathy
Publisher
Dwarkesh Podcast

07 / belief

You just take all the course materials and then I think you could serve a very good automated TA for the student when they have more basic questions or something like that.

“You just take all the course materials and then I think you could serve a very good automated TA for the student when they have more basic questions or something like that.”
Speaker
Andrej Karpathy
Publisher
Dwarkesh Podcast

09 / observation

The way to synchronize gradients between them is to use a Distributed Data Parallel container of PyTorch, which automatically as you’re doing the backward, it will start communicating and synchronizing gradients.

“The way to synchronize gradients between them is to use a Distributed Data Parallel container of PyTorch, which automatically as you’re doing the backward, it will start communicating and synchronizing gradients.”
Speaker
Andrej Karpathy
Publisher
Dwarkesh Podcast

13 / belief

One example that’s prominently in my mind, this was probably public, if you’re using an LLM judge for a reward, you just give it a solution from a student and ask it if the student did well or not.

“One example that’s prominently in my mind, this was probably public, if you’re using an LLM judge for a reward, you just give it a solution from a student and ask it if the student did well or not.”
Speaker
Andrej Karpathy
Publisher
Dwarkesh Podcast

19 / uncertainty

We’re really far into a territory where I don’t know what this looks like, but if I were to write sci-fi novels, they would look along the lines of not even a single entity that takes over everything, but multiple competing entities that gradually become more and more autonomous.

“We’re really far into a territory where I don’t know what this looks like, but if I were to write sci-fi novels, they would look along the lines of not even a single entity that takes over everything, but multiple competing entities that gradually become more and more autonomous.”
Speaker
Andrej Karpathy
Publisher
Dwarkesh Podcast

24 / belief

Even the early iPhone didn’t have the App Store, and it didn’t have a lot of the bells and whistles that the modern iPhone has. So even though we think of 2008, when the iPhone came out, as this major seismic change, it’s actually not.

“Even the early iPhone didn’t have the App Store, and it didn’t have a lot of the bells and whistles that the modern iPhone has. So even though we think of 2008, when the iPhone came out, as this major seismic change, it’s actually not.”
Speaker
Andrej Karpathy
Publisher
Dwarkesh Podcast

29 / evaluation

The paper that blew my mind was InstructGPT, because it pointed out that you can take the pretrained model, which is autocomplete, and if you just fine-tune it on text that looks like conversations, the model will very rapidly adapt to become very conversational, and it keeps all the knowledge from pre-training.

“The paper that blew my mind was InstructGPT, because it pointed out that you can take the pretrained model, which is autocomplete, and if you just fine-tune it on text that looks like conversations, the model will very rapidly adapt to become very conversational, and it keeps all the knowledge from pre-training.”
Speaker
Andrej Karpathy
Publisher
Dwarkesh Podcast

31 / evaluation

Just a lot of it. So maybe one way to think about it, I don’t know if this is the best way, but I almost feel like — again, making these analogies imperfect as they are — we’ve stumbled by with the transformer neural network, which is extremely powerful, very general.

“Just a lot of it. So maybe one way to think about it, I don’t know if this is the best way, but I almost feel like — again, making these analogies imperfect as they are — we’ve stumbled by with the transformer neural network, which is extremely powerful, very general.”
Speaker
Andrej Karpathy
Publisher
Dwarkesh Podcast

32 / evaluation

” The Atari deep reinforcement learning shift in 2013 or so was part of that early effort of agents, in my mind, because it was an attempt to try to get agents that not just perceive the world, but also take actions and interact and get rewards from environments.

“” The Atari deep reinforcement learning shift in 2013 or so was part of that early effort of agents, in my mind, because it was an attempt to try to get agents that not just perceive the world, but also take actions and interact and get rewards from environments.”
Speaker
Andrej Karpathy
Publisher
Dwarkesh Podcast

33 / evaluation

In fact, I think the paper was even stronger because they hardcoded the weights of a neural network to do gradient descent through attention and all the internals of the neural network.

“In fact, I think the paper was even stronger because they hardcoded the weights of a neural network to do gradient descent through attention and all the internals of the neural network.”
Speaker
Andrej Karpathy
Publisher
Dwarkesh Podcast

37 / recommendation

I have a whole rant on how everyone should learn physics in early school education because early school education is not about accumulating knowledge or memory for tasks later in the industry.

“I have a whole rant on how everyone should learn physics in early school education because early school education is not about accumulating knowledge or memory for tasks later in the industry.”
Speaker
Andrej Karpathy
Publisher
Dwarkesh Podcast
Search evidence