Evidence receipt / belief
Published · transcript-backedSpeaker unverified: belief
6 Oct 2024 Lex Fridman Podcast #447 – Cursor Team: Future of Programming with AI
“The original scaling laws paper by open AI was slightly wrong. Because I think of some issues they did with learning right schedules.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- belief
- Recorded
- 6 Oct 2024
- Publisher
- Lex Fridman Podcast
Transcript context
…odel has a much easier time verifying some solution than it does generating it. Then you actually could perhaps get this kind of recursive loop. But I don’t think it’s going to look exactly like that. The other thing you could do, that we kind of do, is a little bit of a mix of RLAIF and RLHF, where usually the model is actually quite correct and this is the case of precursor tap picking between two possible generations of what is the better one. And then it just needs a little bit of human nudging with only on the order 50, 100 examples to align that prior the model has with exactly with what you want. It looks different than I think normal RLHF where you’re usually training these reward models in tons of examples. What’s your intuition when you compare generation and verification or generation and ranking? Is ranking way easier than generation? My intuition would just say, yeah, it should be. This is going back to… Like, if you believe P does not equal NP, then there’s this massive class of problems that are much, much easier to verify given proof, than actually proving it. I wonder if the same thing will prove P not equal to NP or P equal to NP. That would be really cool. That’d be a whatever Field’s Medal by AI. Who gets the credit? Another the open philosophical question. Whoever prompted it. I’m actually surprisingly curious what a good bet for one AI will get the Field’s Medal will be. I actually don’t have- Isn’t this Aman’s specialty? I don’t know what Aman’s bet here is. Oh, sorry, Nobel Prize or Field’s Medal first? Field’s Medal- Oh, Field’s Medal level? Field’s Medal comes first, I think. [inaudible 02:06:41]. Field’s Medal comes first. Well, you would say that, of course. But it’s also this isolated system you verify and… Sure. I don’t even know if I- You don’t need to do [inaudible 02:06:50]. I feel like I have much more to do there. It felt like the path to get to IMO was a little bit more clear. Because it already could get a few IMO problems and there was a bunch of low-hanging fruit, given the literature at the time, of what tactics people could take. I think I’m, one, much less versed in the space of theorem proving now. And two, less intuition about how close we are to solving these really, really hard open problems. So you think you’ll be Field’s Medal first? It won’t be in physics or in- Oh, 100%. I think that’s probably more likely. It is probably much more likely that it’ll get in. Yeah, yeah, yeah. Well I think it both to… I don’t know, BSD, which is a Birch and Swinnerton-Dyer conjecture, or [inaudible 02:07:33] iPods, or any one of these hard math problems are just actually really hard. It’s sort of unclear what the path to get even a solution looks like. We don’t even know what a path looks like, let alone [inaudible 02:07:47]. And you don’t buy the idea this is just like an isolated system and you can actually have a good reward system, and it feels like it’s easier to train for that. I think we might get Field’s Medal before AGI. I mean, I’d be very happy. I’d be very happy. But I don’t know if I… I think 2028, 2030. For Field’s Medal? Field’s Medal. it’s easier to train for that. I think we might get Field’s Medal before AGI. I mean, I’d be very happy. I’d be very happy. But I don’t know if I… I think 2028, 2030. For Field’s Medal? Field’s Medal. All right. It feels like forever from now, given how fast things have been going. Speaking of how fast things have been going, let’s talk about scaling laws. So for people who don’t know, maybe it’s good to talk about this whole idea of scaling laws. What are they, where’d you think stand, and where do you think things are going? I think it was interesting. The original scaling laws paper by open AI was slightly wrong. Because I think of some issues they did with learning right schedules. And then Chinchilla showed a more correct version. And then from then people have again deviated from doing the compute optimal thing. Because people start now optimizing more so for making the thing work really well given an inference budget. And I think there are a lot more dimensions to these curves than what we originally used, of just compute number of parameters and data. Like inference compute is the obvious one. I think context length is another obvious one. So let’s say you care about the two things of inference compute and then context window, maybe the thing you want to train, is some kind of SSM. Because they’re much, much cheaper and faster at super, super long context. And even if, maybe it was 10 X more scaling properties during training, meaning you spend 10 X more compute to train the thing to get the same level of capabilities, it’s worth it. Because you care most about that inference budget for really long context windows. So it’ll be interesting to see how people play with all these dimensions. So yeah, I mean you speak to the multiple dimensions, obviously. The original conception was just looking at the variables of the size of the model as measured by parameters, and the size of the data as measured by the number of tokens, and looking at the ratio of the two. Yeah. And it’s kind of a compelling notion that there is a number, or at least a minimum. And it seems like one was emerging. Do you still believe that there is a kind of bigger is better? I mean I think bigger is certainly better for just raw performance. And raw intelligence. And raw intelligence. I think the path that people might take, is… I’m particularly bullish on distillation. And how many knobs can you turn to, if we spend a ton, ton of money on training, get the most capable cheap model. Really, really caring as much as you can. Because the naive version of caring as much as you can about inference time compute, is what people have already done with the Llama models. Or just over-training the shit out of 7B models on way, way, way more tokens than is essential optimal. But if you really care about it, maybe the thing to do is what Gamma did, which is let’s not just train on tokens, let’s literally train on minimizing the KL divergence with the distribution of gemma 27B, right? So knowledge distillation there. And you’re spending the compute of literally training this 27 billion parameter model on all these tokens, just to get out this, I don’t know, smaller model. B, right? So knowledge distillation there. And you’re spending the compute of literally training this 27 billion parameter model on all these tokens, just to get out this, I don’t know, smaller model. And the distillation gives you just a faster model, smaller means faster. Yeah. Distillation in theory is, I think, getting out more signal from the data that you’re training on. And it’s perhaps another way of getting over, not completely over, but partially helping with the data wall. Where you only have so much data to train on, let’s train this really, really big model on all these tokens and we’ll distill it into this smaller one. And maybe we can get more signal per token for this much smaller model than we would’ve originally if we trained it. So if I gave you $10 trillion, how would you spend it? I mean you can’t buy an island or whatever. How would you allocate it in terms of improving the big model versus maybe paying for HF in the RLHF? Or- Yeah, yeah. I think there’s a lot of these secrets and details about training these large models that I just don’t know, and are only privy to the large labs. And the issue is, I would waste a lot of that money if I even attempted this, because I wouldn’t know those things. Suspending a lot of disbelief and assuming you had the know- how, or if you’re saying you have to operate with the limited information you have now- No, no, no. Actually, I would say you swoop in and you get all the information, all the little heuristics, all the little parameters, all the parameters that define how the thing is trained. Mm-hmm. If we look in how to invest money for the next five years in terms of maximizing what you called raw intelligence- I mean, isn’t the answer really simple? You just try to get as much compute as possible. At the end of the day all you need to buy, is the GPUs. And then the researchers can find all… You can tune whether you want to pre-train a big model or a small model. Well this gets into the question of are you really limited by compute and money, or are you limited by these other things? I’m more privy to Arvid’s belief that we’re sort of idea-limited, but there’s always that like- But if you have a lot of compute, you can run a lot of experiments. So you would run a lot of experiments versus use that compute to trend a gigantic model? I would, but I do believe that we are limited in terms of ideas that we have.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.