High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Quentin Anthony

Published podcast speaker

Claims
16
Episodes
1
Shows
1
Named items
0

Claim ledger

What Quentin said.

16 transcript-backed records

01 / recommendation

You might have synchronization issues. And that's why I'm saying practically that you should take the minimum number of GPUs that you have because those are the easier cases to debug.

“You might have synchronization issues. And that's why I'm saying practically that you should take the minimum number of GPUs that you have because those are the easier cases to debug.”
Speaker
Quentin Anthony
Publisher
Latent Space

02 / observation

The problem is that in practice, full utilization, that's the key word, right? Because in practice, there are a lot of cases where like you're spending time waiting on data movement from like the GPU to CPU.

“The problem is that in practice, full utilization, that's the key word, right? Because in practice, there are a lot of cases where like you're spending time waiting on data movement from like the GPU to CPU.”
Speaker
Quentin Anthony
Publisher
Latent Space

04 / belief

I would say flash attention. Guys, just talk to Tree. And flash attention is just sort of a really great set of kernels that I thought would take a while to get to us.

“I would say flash attention. Guys, just talk to Tree. And flash attention is just sort of a really great set of kernels that I thought would take a while to get to us.”
Speaker
Quentin Anthony
Publisher
Latent Space

06 / belief

Next up would be something like selective recomputation. So in selective recomputation, which Megatron has a good paper on, and I believe the figure that we have in our blog post is from, in that case, you sort of do a weighted decision for each activation.

“Next up would be something like selective recomputation. So in selective recomputation, which Megatron has a good paper on, and I believe the figure that we have in our blog post is from, in that case, you sort of do a weighted decision for each activation.”
Speaker
Quentin Anthony
Publisher
Latent Space

10 / evaluation

And we did a lot of experiments there. It's a challenging system to scale to because your interconnect across nodes is kind of slow in comparison to within a node, which I think we'll get to later.

“And we did a lot of experiments there. It's a challenging system to scale to because your interconnect across nodes is kind of slow in comparison to within a node, which I think we'll get to later.”
Speaker
Quentin Anthony
Publisher
Latent Space

12 / observation

So once the hardware moved to V100, the software moved to like mixed precision and APEX and AMP and such. And one counterintuitive part of mixed precision is that you actually require more memory when you're trained because you need an FP16 copy of the weights and an FP32 copy of the weights.

“So once the hardware moved to V100, the software moved to like mixed precision and APEX and AMP and such. And one counterintuitive part of mixed precision is that you actually require more memory when you're trained because you need an FP16 copy of the weights and an FP32 copy of the weights.”
Speaker
Quentin Anthony
Publisher
Latent Space

13 / recommendation

I always feel like you should go with the minimum number of GPUs because the more number of GPUs that you have, the more likely it is for things to break.

“I always feel like you should go with the minimum number of GPUs because the more number of GPUs that you have, the more likely it is for things to break.”
Speaker
Quentin Anthony
Publisher
Latent Space

16 / preference

I would say even if you had infinite VRAM, you would still want a quantized model, just a bigger model that's quantized is what I would say. And that's because like I was mentioning there at the end, how like deep learning is very stochastic and a lot, you could have all the precision in the world, but ultimately it's meaningless when you still depend so much like on what the input is.

“I would say even if you had infinite VRAM, you would still want a quantized model, just a bigger model that's quantized is what I would say. And that's because like I was mentioning there at the end, how like deep learning is very stochastic and a lot, you could have all the precision in the world, but ultimately it's meaningless when you still depend so much like on what the input is.”
Speaker
Quentin Anthony
Publisher
Latent Space
Search evidence