High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Dylan Patel: belief

3 Feb 2025 Lex Fridman Podcast #459 – DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters

“I think there is one aspect to note though is that there is the general ability for that to transfer across different types of runs.”

— Dylan Patel

Source trail

Everything needed to verify it.

Speaker
Dylan Patel
Attribution
Verified speaker
Claim type
belief
Recorded
3 Feb 2025
Publisher
Lex Fridman Podcast

Transcript context

…So, high-quality, readable code. Yeah. I think there is one aspect to note though is that there is the general ability for that to transfer across different types of runs. You may make really, really high-quality code for one specific model architecture at one size, and then that is not transferable to, ” Hey, when I make this architecture tweak, everything’s broken again,” right? That’s something that could be with their specific low-level coding of scheduling SMs is specific to this model architecture and size. Whereas, Nvidia’s Collectives Library is more like, “Hey, it’ll work for anything,” right? “You want to do an allreduce? Great, I don’t care what your model architecture is, it’ll work,” and you’re giving up a lot of performance when you do that in many cases, but it’s worthwhile for them to do the specific optimization for the specific run given the constraints that they have regarding compute. I wonder how stressful it is to these frontier models, like initiate training to have the code-…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence