Evidence receipt / belief
Published · transcript-backedSpeaker unverified: belief
6 Oct 2024 Lex Fridman Podcast #447 – Cursor Team: Future of Programming with AI
“Actually, I would say you swoop in and you get all the information, all the little heuristics, all the little parameters, all the parameters that define how the thing is trained.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- belief
- Recorded
- 6 Oct 2024
- Publisher
- Lex Fridman Podcast
Transcript context
…it’s easier to train for that. I think we might get Field’s Medal before AGI. I mean, I’d be very happy. I’d be very happy. But I don’t know if I… I think 2028, 2030. For Field’s Medal? Field’s Medal. All right. It feels like forever from now, given how fast things have been going. Speaking of how fast things have been going, let’s talk about scaling laws. So for people who don’t know, maybe it’s good to talk about this whole idea of scaling laws. What are they, where’d you think stand, and where do you think things are going? I think it was interesting. The original scaling laws paper by open AI was slightly wrong. Because I think of some issues they did with learning right schedules. And then Chinchilla showed a more correct version. And then from then people have again deviated from doing the compute optimal thing. Because people start now optimizing more so for making the thing work really well given an inference budget. And I think there are a lot more dimensions to these curves than what we originally used, of just compute number of parameters and data. Like inference compute is the obvious one. I think context length is another obvious one. So let’s say you care about the two things of inference compute and then context window, maybe the thing you want to train, is some kind of SSM. Because they’re much, much cheaper and faster at super, super long context. And even if, maybe it was 10 X more scaling properties during training, meaning you spend 10 X more compute to train the thing to get the same level of capabilities, it’s worth it. Because you care most about that inference budget for really long context windows. So it’ll be interesting to see how people play with all these dimensions. So yeah, I mean you speak to the multiple dimensions, obviously. The original conception was just looking at the variables of the size of the model as measured by parameters, and the size of the data as measured by the number of tokens, and looking at the ratio of the two. Yeah. And it’s kind of a compelling notion that there is a number, or at least a minimum. And it seems like one was emerging. Do you still believe that there is a kind of bigger is better? I mean I think bigger is certainly better for just raw performance. And raw intelligence. And raw intelligence. I think the path that people might take, is… I’m particularly bullish on distillation. And how many knobs can you turn to, if we spend a ton, ton of money on training, get the most capable cheap model. Really, really caring as much as you can. Because the naive version of caring as much as you can about inference time compute, is what people have already done with the Llama models. Or just over-training the shit out of 7B models on way, way, way more tokens than is essential optimal. But if you really care about it, maybe the thing to do is what Gamma did, which is let’s not just train on tokens, let’s literally train on minimizing the KL divergence with the distribution of gemma 27B, right? So knowledge distillation there. And you’re spending the compute of literally training this 27 billion parameter model on all these tokens, just to get out this, I don’t know, smaller model. B, right? So knowledge distillation there. And you’re spending the compute of literally training this 27 billion parameter model on all these tokens, just to get out this, I don’t know, smaller model. And the distillation gives you just a faster model, smaller means faster. Yeah. Distillation in theory is, I think, getting out more signal from the data that you’re training on. And it’s perhaps another way of getting over, not completely over, but partially helping with the data wall. Where you only have so much data to train on, let’s train this really, really big model on all these tokens and we’ll distill it into this smaller one. And maybe we can get more signal per token for this much smaller model than we would’ve originally if we trained it. So if I gave you $10 trillion, how would you spend it? I mean you can’t buy an island or whatever. How would you allocate it in terms of improving the big model versus maybe paying for HF in the RLHF? Or- Yeah, yeah. I think there’s a lot of these secrets and details about training these large models that I just don’t know, and are only privy to the large labs. And the issue is, I would waste a lot of that money if I even attempted this, because I wouldn’t know those things. Suspending a lot of disbelief and assuming you had the know- how, or if you’re saying you have to operate with the limited information you have now- No, no, no. Actually, I would say you swoop in and you get all the information, all the little heuristics, all the little parameters, all the parameters that define how the thing is trained. Mm-hmm. If we look in how to invest money for the next five years in terms of maximizing what you called raw intelligence- I mean, isn’t the answer really simple? You just try to get as much compute as possible. At the end of the day all you need to buy, is the GPUs. And then the researchers can find all… You can tune whether you want to pre-train a big model or a small model. Well this gets into the question of are you really limited by compute and money, or are you limited by these other things? I’m more privy to Arvid’s belief that we’re sort of idea-limited, but there’s always that like- But if you have a lot of compute, you can run a lot of experiments. So you would run a lot of experiments versus use that compute to trend a gigantic model? I would, but I do believe that we are limited in terms of ideas that we have. you can run a lot of experiments. So you would run a lot of experiments versus use that compute to trend a gigantic model? I would, but I do believe that we are limited in terms of ideas that we have. I think yeah, because even with all this compute and all the data you could collect in the world, I think you really are ultimately limited by not even ideas, but just really good engineering. Even with all the capital in the world, would you really be able to assemble… There aren’t that many people in the world who really can make the difference here. And there’s so much work that goes into research that is just pure, really, really hard engineering work. As a very hand-wavy example, if you look at the original Transformer paper, how much work was joining together a lot of these really interesting concepts embedded in the literature, versus then going in and writing all the codes, maybe the CUDA kernels, maybe whatever else. I don’t know if it ran them GPUs or TPUs. Originally such that it actually saturated the GPU performance. Getting GNOME Azure to go in and do all this code. And GNOME is probably one of the best engineers in the world. Or maybe going a step further, like the next generation of models, having these things… Like getting model parallelism to work, and scaling it on thousands of, or maybe tens of thousands of V100s, which I think GBDE-III may have been. There’s just so much engineering effort that has to go into all of these things to make it work. If you really brought that cost down to maybe not zero, but just made it 10 X easier, made it super easy for someone with really fantastic ideas, to immediately get to the version of the new architecture they dreamed up, that is getting 50, 40% utilization on their GPUs, I think that would just speed up research by a ton. I mean I think if you see a clear path to improvement, you should always take the low-hanging fruit first, right? I think probably OpenAI and all the other labs that did the right thing to pick off the low-hanging fruit. Where the low-hanging fruit is like, you could scale up to a GPT-4.25 scale and you just keep scaling, and things keep getting better. And as long as… There’s no point of experimenting with new ideas when everything is working. And you should sort of bang on and to try to get as much as much juice out of the possible. And then maybe when you really need new ideas for… I think if you’re spending $10 trillion, you probably want to spend some… Then actually reevaluate probably your idea a little bit at that point. I think all of us believe new ideas are probably needed to get all the way there to AGI. And all of us also probably believe there exist ways of testing out those ideas at smaller scales, and being fairly confident that they’ll play out. It’s just quite difficult for the labs in their current position to dedicate their very limited research and engineering talent to exploring all these other ideas, when there’s this core thing that will probably improve performance for some decent amount of time.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.