High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / observation

Published · transcript-backed

Andrej Karpathy: observation

17 Oct 2025 Dwarkesh Podcast Andrej Karpathy — AGI is still a decade away

“The way to synchronize gradients between them is to use a Distributed Data Parallel container of PyTorch, which automatically as you’re doing the backward, it will start communicating and synchronizing gradients.”

— Andrej Karpathy

Source trail

Everything needed to verify it.

Speaker
Andrej Karpathy
Attribution
Verified speaker
Claim type
observation
Recorded
17 Oct 2025
Publisher
Dwarkesh Podcast

Transcript context

…Maybe one example? You have eight GPUs that are all doing forward, backwards. The way to synchronize gradients between them is to use a Distributed Data Parallel container of PyTorch, which automatically as you’re doing the backward, it will start communicating and synchronizing gradients. I didn’t use DDP because I didn’t want to use it, because it’s not necessary. I threw it out and wrote my own synchronization routine that’s inside the step of the optimizer. The models were trying to get me to use the DDP container. They were very concerned. This gets way too technical, but I wasn’t using that container because I don’t need it and I have a custom implementation of something like it. They just couldn’t internalize that you had your own.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence