High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Speaker unverified: evaluation

6 Oct 2024 Lex Fridman Podcast #447 – Cursor Team: Future of Programming with AI

“I think there’s no model that Pareto dominates others, meaning it is better in all categories that we think matter, the categories being speed, ability to edit code, ability to process lots of code, long context, a couple of other things and coding capabilities.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
evaluation
Recorded
6 Oct 2024
Publisher
Lex Fridman Podcast

Transcript context

…ggestion of what new things to do. And the seemingly for humans trivial step of combining the two, you’re saying is not so trivial. Contrary to popular perception, it is not a deterministic algorithm. Yeah, I think you see shallow copies of apply elsewhere and it just breaks most of the time because you think you can try to do some deterministic matching and then it fails at least 40% of the time and that just results in a terrible product experience. I think in general, this regime of you are going to get smarter and smarter models. So one other thing that Apply lets you do is it lets you use fewer tokens with the most intelligent models. This is both expensive in terms of latency for generating all these tokens and cost. So you can give this very, very rough sketch and then have your model models go and implement it because it’s a much easier task to implement this very, very sketched out code. And I think that this regime will continue where you can use smarter and smarter models to do the planning and then maybe the implementation details can be handled by the less intelligent ones. Perhaps you’ll have maybe o1, maybe it’ll be even more capable models given an even higher level plan that is recursively applied by sauna and then the apply model. Maybe we should talk about how to make it fast if you like. Fast is always an interesting detail. Fast is good. Yeah, how do you make it fast? Yeah, so one big component of making it fast is speculative edits. So speculative edits are a variant of speculative decoding, and maybe it’d be helpful to briefly describe speculative decoding. With speculative decoding, what you do is you can take advantage of the fact that most of the time, and I’ll add the caveat that it would be when you’re memory bound in language model generation, if you process multiple tokens at once, it is faster than generating one token at a time. So this is the same reason why if you look at tokens per second with prompt tokens versus generated tokens, it’s much much faster for prompt tokens. So what we do is instead of using what speculative decoding normally does, which is using a really small model to predict these draft tokens that your larger model will then go in and verify, with code edits, we have a very strong prior of what the existing code will look like and that prior is literally the same exact code. So you can do is you can just feed chunks of the original code back into the model, and then the model will just pretty much agree most of the time that, “Okay, I’m just going to spit this code back out.” And so you can process all of those lines in parallel and you just do this with sufficiently many chunks. And then eventually you’ll reach a point of disagreement where the model will now predict text that is different from the ground truth original code. It’ll generate those tokens and then we will decide after enough tokens match the original code to re- start speculating in chunks of code. text that is different from the ground truth original code. It’ll generate those tokens and then we will decide after enough tokens match the original code to re- start speculating in chunks of code. What this actually ends up looking like is just a much faster version of normal editing code. So it looks like a much faster version of the model rewriting all the code. So we can use the same exact interface that we use for diffs, but it will just stream down a lot faster. And then the advantage is that while it’s streaming, you can just also start reviewing the code before it’s done so there’s no big loading screen. Maybe that is part of the advantage. So the human can start reading before the thing is done. I think the interesting riff here is something like… I feel like speculation is a fairly common idea nowadays. It’s not only in language models. There’s obviously speculation in CPUs and there’s speculation for databases and there’s speculation all over the place. Well, let me ask the ridiculous question of which LLM is better at coding? GPT, Claude, who wins in the context of programming? And I’m sure the answer is much more nuanced because it sounds like every single part of this involves a different model. I think there’s no model that Pareto dominates others, meaning it is better in all categories that we think matter, the categories being speed, ability to edit code, ability to process lots of code, long context, a couple of other things and coding capabilities. The one that I’d say right now is just net best is Sonnet. I think this is a consensus opinion. o1’s really interesting and it’s really good at reasoning. So if you give it really hard programming interview style problems or lead code problems, it can do quite well on them, but it doesn’t feel like it understands your rough intent as well as Sonnet does. If you look at a lot of the other frontier models, one qualm I have is it feels like they’re not necessarily over… I’m not saying they train on benchmarks, but they perform really well in benchmarks relative to everything that’s in the middle. So if you tried on all these benchmarks and things that are in the distribution of the benchmarks they’re evaluated on, they’ll do really well. But when you push them a little bit outside of that, Sonnet is I think the one that does best at maintaining that same capability. You have the same capability in the benchmark as when you try to instruct it to do anything with coding. Another ridiculous question is the difference between the normal programming experience versus what benchmarks represent? Where do benchmarks fall short, do you think, when we’re evaluating these models? ther ridiculous question is the difference between the normal programming experience versus what benchmarks represent? Where do benchmarks fall short, do you think, when we’re evaluating these models? By the way, that’s a really, really hard, critically important detail of how different benchmarks are versus real coding, where real coding, it’s not interview style coding. Humans are saying half-broken English sometimes and sometimes you’re saying, “Oh, do what I did before.” Sometimes you’re saying, “Go add this thing and then do this other thing for me and then make this UI element.” And then it’s just a lot of things are context dependent. You really want to understand the human and then do what the human wants, as opposed to this… Maybe the way to put it abstractly is the interview problems are very well specified. They lean a lot on specification while the human stuff is less specified. I think that this benchmark question is both complicated by what Sualeh just mentioned, and then also what Aman was getting into is that even if you… There’s this problem of the skew between what can you actually model in a benchmark versus real programming, and that can be sometimes hard to encapsulate because it’s real programming’s very messy and sometimes things aren’t super well specified what’s correct or what isn’t. But then it’s also doubly hard because of this public benchmark problem. And that’s both because public benchmarks are sometimes hill climbed on, then it’s really, really hard to also get the data from the public benchmarks out of the models. And so for instance, one of the most popular agent benchmarks, SWE-Bench, is really, really contaminated in the training data of these foundation models. And so if you ask these foundation models to do a SWE-Bench problem, but you actually don’t give them the context of a code base, they can hallucinate the right file pass, they can hallucinate the right function names. And so it’s also just the public aspect of these things is tricky. In that case, it could be trained on the literal issues or pull requests themselves, and maybe the labs will start to do a better job or they’ve already done a good job at decontaminating those things, but they’re not going to omit the actual training data of the repository itself. These are all some of the most popular Python repositories. SimPy is one example. I don’t think they’re going to handicap their models on SimPy and all these popular Python repositories in order to get true evaluation scores in these benchmarks. I think that given the dirts in benchmarks, there have been a few interesting crutches that places that build systems with these models or build these models actually use to get a sense of are they going the right direction or not. And in a lot of places, people will actually just have humans play with the things and give qualitative feedback on these. One or two of the foundation model companies, they have people who that’s a big part of their role. And internally, we also qualitatively assess these models and actually lean on that a lot in addition to private emails that we have. It’s like the vibe. The vibe, yeah, the vibe. It’s like the vibe.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence