High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Speaker unverified: belief

6 Oct 2024 Lex Fridman Podcast #447 – Cursor Team: Future of Programming with AI

“The last category I think is, I guess the main one that it feels like the big labs are doing for synthetic data, which is producing text with language models that can then be verified easily.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
belief
Recorded
6 Oct 2024
Publisher
Lex Fridman Podcast

Transcript context

…es this mean Cursor is done? I think I saw one comment saying that. It’s a time to shut down Cursor. Yeah. Time to shut down Cursor. [inaudible 01:58:38]. Thank you. So is it time to shut down Cursor? I think this space is a little bit different from past software spaces over the 2010s, where I think that the ceiling here is really, really, really incredibly high. So I think that the best product in three to four years will just be soon much more useful than the best product today. You can wax poetic about moats this and brand that and this is our advantage, but I think in the end, just if you stop innovating on the product, you will lose. That’s also great for startups, that’s great for people trying to enter this market because it means you have an opportunity to win against people who have lots of users already by just building something better. So I think over the next few years, it’s just about building the best product, building the best system. That both comes down to the modeling engine side of things, and it also comes down to the editing experience. Yeah, I think most of the additional value from Cursor versus everything else out there is not just integrating the new model fast like o1. It comes from all of the depth that goes into these custom models that you don’t realize are working for you in every facet of the product, as well as the really thoughtful UX with every single feature. All right. From that profound answer- All right, from that profound answer, let’s descend back down to the technical. You mentioned you have a taxonomy of synthetic data. Oh yeah. Can you please explain? Yeah, I think there are three main kinds of synthetic data. So what is synthetic data, first? So there’s normal data, like non-synthetic data, which is just data that’s naturally created, i.e. usually it’ll be from humans having done things. So from some human process you get this data. Synthetic data, the first one would be distillation. So having a language model, output tokens or probability distributions over tokens, and then you can train some less capable model on this. This approach is not going to get you a more capable model than the original one that has produced the tokens, but it’s really useful for if there’s some capability you want to elicit from some really expensive high-latency model. You can then distill that down into some smaller task-specific model. The second kind is when one direction of the problem is easier than the reverse. So a great example of this is bug detection, like we mentioned earlier, where it’s a lot easier to introduce reasonable-looking bugs than it is to actually detect them. And this is probably the case for humans too. And so what you can do, is you can get a model that’s not trained in that much data, that’s not that smart, to introduce a bunch of bugs and code. And then you can use that to then train… Use the synthetic data to train a model that can be really good at detecting bugs. hat much data, that’s not that smart, to introduce a bunch of bugs and code. And then you can use that to then train… Use the synthetic data to train a model that can be really good at detecting bugs. The last category I think is, I guess the main one that it feels like the big labs are doing for synthetic data, which is producing text with language models that can then be verified easily. So extreme example of this is if you have a verification system that can detect if language is Shakespeare level, and then you have a bunch of monkeys typing and typewriters. You can eventually get enough training data to train a Shakespeare-level language model. And I mean this is very much the case for math where verification is actually really, really easy for formal languages. And then what you can do, is you can have an okay model, generate a ton of rollouts, and then choose the ones that you know have actually proved the ground truth theorems, and train that further. There’s similar things you can do for code with lead code like problems, where if you have some set of tests that you know correspond to if something passes these tests, it actually solved problem. You could do the same thing where you verify that it’s passed the test and then train the model in the outputs that have passed the tests. I think it’s going to be a little tricky getting this to work in all domains, or just in general. Having the perfect verifier feels really, really hard to do with just open-ended miscellaneous tasks. You give the model or more long horizon tasks, even in coding. That’s because you’re not as optimistic as Arvid. But yeah, so yeah, that third category requires having a verifier. Verification, it feels like it’s best when you know for a fact that it’s correct. And then it wouldn’t be like using a language model to verify. It would be using tests or formal systems. Or running the thing too. Doing the human form of verification, where you just do manual quality control. Yeah. But the language model version of that, where it’s running the thing and it actually understands the output. Yeah. No, that’s- I’m sure it’s somewhere in between. Yeah. I think that’s the category that is most likely to result in massive gains. What about RL with feedback side RLHF versus RLAIF? What’s the role of that in getting better performance on the models? Yeah. So RLHF is when the reward model you use is trained from some labels you’ve collected from humans giving feedback. I think this works if you have the ability to get a ton of human feedback for this kind of task that you care about. RLAIF is interesting because you’re depending on… This is actually, it’s depending on the constraint that verification is actually a decent bit easier than generation. Because it feels like, okay, what are you doing? Are you using this language model to look at the language model outputs and then prove the language model? But no, it actually may work if the language model has a much easier time verifying some solution than it does generating it. Then you actually could perhaps get this kind of recursive loop. But I don’t think it’s going to look exactly like that. odel has a much easier time verifying some solution than it does generating it. Then you actually could perhaps get this kind of recursive loop. But I don’t think it’s going to look exactly like that. The other thing you could do, that we kind of do, is a little bit of a mix of RLAIF and RLHF, where usually the model is actually quite correct and this is the case of precursor tap picking between two possible generations of what is the better one. And then it just needs a little bit of human nudging with only on the order 50, 100 examples to align that prior the model has with exactly with what you want. It looks different than I think normal RLHF where you’re usually training these reward models in tons of examples. What’s your intuition when you compare generation and verification or generation and ranking? Is ranking way easier than generation? My intuition would just say, yeah, it should be. This is going back to… Like, if you believe P does not equal NP, then there’s this massive class of problems that are much, much easier to verify given proof, than actually proving it. I wonder if the same thing will prove P not equal to NP or P equal to NP. That would be really cool. That’d be a whatever Field’s Medal by AI. Who gets the credit? Another the open philosophical question. Whoever prompted it. I’m actually surprisingly curious what a good bet for one AI will get the Field’s Medal will be. I actually don’t have- Isn’t this Aman’s specialty? I don’t know what Aman’s bet here is. Oh, sorry, Nobel Prize or Field’s Medal first? Field’s Medal- Oh, Field’s Medal level? Field’s Medal comes first, I think. [inaudible 02:06:41]. Field’s Medal comes first. Well, you would say that, of course. But it’s also this isolated system you verify and… Sure. I don’t even know if I- You don’t need to do [inaudible 02:06:50]. I feel like I have much more to do there. It felt like the path to get to IMO was a little bit more clear. Because it already could get a few IMO problems and there was a bunch of low-hanging fruit, given the literature at the time, of what tactics people could take. I think I’m, one, much less versed in the space of theorem proving now. And two, less intuition about how close we are to solving these really, really hard open problems. So you think you’ll be Field’s Medal first? It won’t be in physics or in- Oh, 100%. I think that’s probably more likely. It is probably much more likely that it’ll get in. Yeah, yeah, yeah. Well I think it both to… I don’t know, BSD, which is a Birch and Swinnerton-Dyer conjecture, or [inaudible 02:07:33] iPods, or any one of these hard math problems are just actually really hard. It’s sort of unclear what the path to get even a solution looks like. We don’t even know what a path looks like, let alone [inaudible 02:07:47]. And you don’t buy the idea this is just like an isolated system and you can actually have a good reward system, and it feels like it’s easier to train for that. I think we might get Field’s Medal before AGI. I mean, I’d be very happy. I’d be very happy. But I don’t know if I… I think 2028, 2030. For Field’s Medal? Field’s Medal.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence