Evidence receipt / belief
Published · transcript-backedShawn Wang: belief
22 Aug 2023 Latent Space Cursor.so: The AI-first Code Editor — with Aman Sanger of Anysphere
“You also had an interesting observation because I think we had talked a lot about quantization in the podcast just for running locally or more efficient running.”
Source trail
Everything needed to verify it.
- Speaker
- Shawn Wang
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 22 Aug 2023
- Publisher
- Latent Space
Transcript context
…We don't have it in production, but we've experimented with it for a few things. You also had an interesting observation because I think we had talked a lot about quantization in the podcast just for running locally or more efficient running. You said quantization and imperfect utilization cancel each other out. Yes. That's a cool observation. Yeah. So this is like a little bit hand wavy, but the core thing is, yeah, we expect that when you don't have like complete utilization, right, you're never going to like saturate all your GPUs. There's going to be some idle time. Like from things that we've experimented with in the past, it ends up being, you know, 50% is a reasonable amount as a more liberal estimate of how much you can get. So the interesting thing about quantization is that there's a bunch of these kind of new quantization libraries that have cropped up and they're all very good at reducing costs for low batch inference when you're memory bound. But the key thing is when you increase the batch size, they actually end up resulting in no real speed ups over FP16. The reason why is because they only quantize the model weights, right? So that operation of kind of reading the model weights when they're now, you know, 4x smaller instead of FP16, they're, you know, 4 bits or something.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.