Speakers in the public record
Claim mix
belief 11evaluation 4recommendation 3observation 3uncertainty 3preference 2prediction 1
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
27 published records
“Honestly, that's one of the fundamental things I love about the math of deep learning so far that as I've explored it, which is, it's surprisingly efficient as compared to other, I guess, numerical methods you might be exposed to and, you know, college calculus.”
- Publisher
- Latent Space
“You might have synchronization issues. And that's why I'm saying practically that you should take the minimum number of GPUs that you have because those are the easier cases to debug.”
- Publisher
- Latent Space
“The problem is that in practice, full utilization, that's the key word, right? Because in practice, there are a lot of cases where like you're spending time waiting on data movement from like the GPU to CPU.”
- Publisher
- Latent Space
“I think Alessio had more questions before we move too far into high level, more questions on just the equation itself.”
- Publisher
- Latent Space
“The naive way to implement fine-tuning is just to do more training. But I don't know if you've developed any intuitions over fine-tuning that's worth inserting here.”
- Publisher
- Latent Space
“I think in my area of research, there are two things that I think people should really care about.”
- Publisher
- Latent Space
“I would say flash attention. Guys, just talk to Tree. And flash attention is just sort of a really great set of kernels that I thought would take a while to get to us.”
- Publisher
- Latent Space
“One, it's an extra effort to write things down, of course, but two, it is secret sauce and so that not many people do it. And therefore, oftentimes the only way to learn this stuff is to actually work in one of the large model labs.”
- Publisher
- Latent Space
“You essentially have model memory, optimizer memory, gradient memory, and activation memory. I think that's one of the last discussed things.”
- Publisher
- Latent Space
“I would say the Atom optimizer is good at what it does. It's sort of a broad question.”
- Publisher
- Latent Space
“Next up would be something like selective recomputation. So in selective recomputation, which Megatron has a good paper on, and I believe the figure that we have in our blog post is from, in that case, you sort of do a weighted decision for each activation.”
- Publisher
- Latent Space
“I know we're at the hour mark, and I think we put our listeners through a very intense class today.”
- Publisher
- Latent Space
“I think we probably want to spend the most of our time on optimizer states and the Atom optimizer.”
- Publisher
- Latent Space
“I would say that a lot of major GPT-based models use this scheme. A lot of them now are sort of going with just a pure zero scheme.”
- Publisher
- Latent Space
“I think a lot of times the focus is on tokens parameter ratio in the training dataset and people don't think as much about the actual flops per GPU, which you're going to mention later in the blog post too, in terms of how much you can get out.”
- Publisher
- Latent Space
“Other than that, I think there was like a doc from Microsoft on like their Turing NLG.”
- Publisher
- Latent Space
“I don't know if I have a smell test myself, to be honest, like maybe I'll look at like what sort of flops, what you would expect on like an A100.”
- Publisher
- Latent Space
“And we did a lot of experiments there. It's a challenging system to scale to because your interconnect across nodes is kind of slow in comparison to within a node, which I think we'll get to later.”
- Publisher
- Latent Space
“If you have a eight bits model quantized down, you need one byte per parameter. So for example, in an H100, which is 80 gigabyte of memory, you could fit a 70 billion parameters in eight, you cannot fit a FP32 because you will need like 280 gigabytes of memory.”
- Publisher
- Latent Space
“I think there's a lot of questions that are unanswered for fine-tuning. For example, we know scaling laws for training.”
- Publisher
- Latent Space
“I think after all of this, you can quickly do the math and see that training needs to be distributed to actually work because we just don't have hardware that can easily run this.”
- Publisher
- Latent Space
“So once the hardware moved to V100, the software moved to like mixed precision and APEX and AMP and such. And one counterintuitive part of mixed precision is that you actually require more memory when you're trained because you need an FP16 copy of the weights and an FP32 copy of the weights.”
- Publisher
- Latent Space
“I always feel like you should go with the minimum number of GPUs because the more number of GPUs that you have, the more likely it is for things to break.”
- Publisher
- Latent Space
“It matters that you are going for the sort of good enough rules of thumb, because I think a lot of people try to go for precision and being overly precise actually is not helpful.”
- Publisher
- Latent Space
“I would say most things run there, not particularly efficiently, but you're going to have weird bugs that no one has encountered before.”
- Publisher
- Latent Space
“I would say I do look at ways to improve upon it using things like second order optimizers.”
- Publisher
- Latent Space
“I would say even if you had infinite VRAM, you would still want a quantized model, just a bigger model that's quantized is what I would say. And that's because like I was mentioning there at the end, how like deep learning is very stochastic and a lot, you could have all the precision in the world, but ultimately it's meaningless when you still depend so much like on what the input is.”
- Publisher
- Latent Space