Evidence receipt / prediction
Published · transcript-backedJeff Dean: prediction
12 Feb 2025 Dwarkesh Podcast Jeff Dean & Noam Shazeer — 25 years at Google: from PageRank to AGI
“" Or even below that, some people are quantizing models to two bits or one bit, and I think that's a trend that definitely –”
Source trail
Everything needed to verify it.
- Speaker
- Jeff Dean
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 12 Feb 2025
- Publisher
- Dwarkesh Podcast
Transcript context
…What are the trade-offs that you're considering changing for future versions of TPU to integrate how you're thinking about algorithms? I think one general trend is we're getting better at quantizing or having much more reduced precision models. We started with TPUv1, and we weren't even quite sure we could quantize and model for serving with eight-bit integers. But we sort of had some early evidence that seemed like it might be possible. So we're like, "Great, let's build the whole chip around that." And then over time, I think you've seen people able to use much lower precision for training as well. But also the inference precision has gone. People are now using INT4 or FP4, which sounded like, if you said to someone like we're going to use FP4, like a supercomputing floating point person 20 years ago, they'd be like, "What? That's crazy. We like 64 bits in our floats. " Or even below that, some people are quantizing models to two bits or one bit, and I think that's a trend that definitely – One bit? Just like a zero-or-one?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.