Evidence receipt / belief
Published · transcript-backedSpeaker unverified: belief
6 Oct 2024 Lex Fridman Podcast #447 – Cursor Team: Future of Programming with AI
“I think yeah, because even with all this compute and all the data you could collect in the world, I think you really are ultimately limited by not even ideas, but just really good engineering.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- belief
- Recorded
- 6 Oct 2024
- Publisher
- Lex Fridman Podcast
Transcript context
…B, right? So knowledge distillation there. And you’re spending the compute of literally training this 27 billion parameter model on all these tokens, just to get out this, I don’t know, smaller model. And the distillation gives you just a faster model, smaller means faster. Yeah. Distillation in theory is, I think, getting out more signal from the data that you’re training on. And it’s perhaps another way of getting over, not completely over, but partially helping with the data wall. Where you only have so much data to train on, let’s train this really, really big model on all these tokens and we’ll distill it into this smaller one. And maybe we can get more signal per token for this much smaller model than we would’ve originally if we trained it. So if I gave you $10 trillion, how would you spend it? I mean you can’t buy an island or whatever. How would you allocate it in terms of improving the big model versus maybe paying for HF in the RLHF? Or- Yeah, yeah. I think there’s a lot of these secrets and details about training these large models that I just don’t know, and are only privy to the large labs. And the issue is, I would waste a lot of that money if I even attempted this, because I wouldn’t know those things. Suspending a lot of disbelief and assuming you had the know- how, or if you’re saying you have to operate with the limited information you have now- No, no, no. Actually, I would say you swoop in and you get all the information, all the little heuristics, all the little parameters, all the parameters that define how the thing is trained. Mm-hmm. If we look in how to invest money for the next five years in terms of maximizing what you called raw intelligence- I mean, isn’t the answer really simple? You just try to get as much compute as possible. At the end of the day all you need to buy, is the GPUs. And then the researchers can find all… You can tune whether you want to pre-train a big model or a small model. Well this gets into the question of are you really limited by compute and money, or are you limited by these other things? I’m more privy to Arvid’s belief that we’re sort of idea-limited, but there’s always that like- But if you have a lot of compute, you can run a lot of experiments. So you would run a lot of experiments versus use that compute to trend a gigantic model? I would, but I do believe that we are limited in terms of ideas that we have. you can run a lot of experiments. So you would run a lot of experiments versus use that compute to trend a gigantic model? I would, but I do believe that we are limited in terms of ideas that we have. I think yeah, because even with all this compute and all the data you could collect in the world, I think you really are ultimately limited by not even ideas, but just really good engineering. Even with all the capital in the world, would you really be able to assemble… There aren’t that many people in the world who really can make the difference here. And there’s so much work that goes into research that is just pure, really, really hard engineering work. As a very hand-wavy example, if you look at the original Transformer paper, how much work was joining together a lot of these really interesting concepts embedded in the literature, versus then going in and writing all the codes, maybe the CUDA kernels, maybe whatever else. I don’t know if it ran them GPUs or TPUs. Originally such that it actually saturated the GPU performance. Getting GNOME Azure to go in and do all this code. And GNOME is probably one of the best engineers in the world. Or maybe going a step further, like the next generation of models, having these things… Like getting model parallelism to work, and scaling it on thousands of, or maybe tens of thousands of V100s, which I think GBDE-III may have been. There’s just so much engineering effort that has to go into all of these things to make it work. If you really brought that cost down to maybe not zero, but just made it 10 X easier, made it super easy for someone with really fantastic ideas, to immediately get to the version of the new architecture they dreamed up, that is getting 50, 40% utilization on their GPUs, I think that would just speed up research by a ton. I mean I think if you see a clear path to improvement, you should always take the low-hanging fruit first, right? I think probably OpenAI and all the other labs that did the right thing to pick off the low-hanging fruit. Where the low-hanging fruit is like, you could scale up to a GPT-4.25 scale and you just keep scaling, and things keep getting better. And as long as… There’s no point of experimenting with new ideas when everything is working. And you should sort of bang on and to try to get as much as much juice out of the possible. And then maybe when you really need new ideas for… I think if you’re spending $10 trillion, you probably want to spend some… Then actually reevaluate probably your idea a little bit at that point. I think all of us believe new ideas are probably needed to get all the way there to AGI. And all of us also probably believe there exist ways of testing out those ideas at smaller scales, and being fairly confident that they’ll play out. It’s just quite difficult for the labs in their current position to dedicate their very limited research and engineering talent to exploring all these other ideas, when there’s this core thing that will probably improve performance for some decent amount of time. n to dedicate their very limited research and engineering talent to exploring all these other ideas, when there’s this core thing that will probably improve performance for some decent amount of time. But also, these big labs like winning. So they’re just going wild. Okay, so big question, looking out into the future: You’re now at the center of the programming world. How do you think programming, the nature of programming changes in the next few months, in the next year, in the next two years and the next five years, 10 years? I think we’re really excited about a future where the programmer is in the driver’s seat for a long time. And you’ve heard us talk about this a little bit, but one that emphasizes speed and agency for the programmer and control. The ability to modify anything you want to modify, the ability to iterate really fast on what you’re building. And this is a little different, I think, than where some people are jumping to in the space, where I think one idea that’s captivated people, is can you talk to your computer? Can you have it build software for you? As if you’re talking to an engineering department or an engineer over Slack. And can it just be this sort of isolated text box? And part of the reason we’re not excited about that, is some of the stuff we’ve talked about with latency, but then a big piece, a reason we’re not excited about that, is because that comes with giving up a lot of control. It’s much harder to be really specific when you’re talking in the text box. And if you’re necessarily just going to communicate with a thing like you would be communicating with an engineering department, you’re actually advocating tons of really important decisions to this bot. And this kind of gets at, fundamentally, what engineering is. I think that some people who are a little bit more removed from engineering might think of it as the spec is completely written out and then the engineers just come and they just implement. And it’s just about making the thing happen in code and making the thing exist. But I think a lot of the best engineering, the engineering we enjoy, involves tons of tiny micro decisions about what exactly you’re building, and about really hard trade-offs between speed and cost and just all the other things involved in a system. As long as humans are actually the ones designing the software and the ones specifying what they want to be built, and it’s not just like company run by all AIs, we think you’ll really want the human in a driver’s seat dictating these decisions.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.