High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Speaker unverified: belief

6 Oct 2024 Lex Fridman Podcast #447 – Cursor Team: Future of Programming with AI

“Then after we had a good model, I think there’ve been a lot of effort to make the inference fast for having a good experience, and we’ve been starting to incorporate… I mean, Michael sort of mentioned this ability to jump to different places and that jump to different places I think came from a feeling of once you accept an edit, it’s like man, it should be just really obvious where to go next.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
belief
Recorded
6 Oct 2024
Publisher
Lex Fridman Podcast

Transcript context

…t build the most useful stuff. Okay. Well then the natural question is, VS Code is kind of with Copilot a competitor, so how do you win? Is it basically just the speed and the quality of the features? Yeah, I mean I think this is a space that is quite interesting, perhaps quite unique where if you look at previous tech waves, maybe there’s kind of one major thing that happened and it unlocked a new wave of companies, but every single year, every single model capability or jump you get in model capabilities, you now unlock this new wave of features, things that are possible, especially in programming. And so I think in AI programming, being even just a few months ahead, let alone a year ahead makes your product much, much, much more useful. I think the Cursor a year from now will need to make the Cursor of today look obsolete. And I think Microsoft has done a number of fantastic things, but I don’t think they’re in a great place to really keep innovating and pushing on this in the way that a startup can. Just rapidly implementing features. Yeah. And kind of doing the research experimentation necessary to really push the ceiling. I don’t know if I think of it in terms of features as I think of it in terms of capabilities for programmers. As the new O1 model came out, and I’m sure there are going to be more models of different types, like longer context and maybe faster, there’s all these crazy ideas that you can try and hopefully 10% of the crazy ideas will make it into something kind of cool and useful and we want people to have that sooner. To rephrase, an underrated fact is we’re making it for ourself. When we started Cursor, you really felt this frustration that models… You could see models getting better, but the Copilot experience had not changed. It was like, man, these guys, the ceiling is getting higher, why are they not making new things? They should be making new things. Where’s all the alpha features? There were no alpha features. I’m sure it was selling well. I’m sure it was a great business, but it didn’t feel… I’m one of these people that really want to try and use new things and there was no new thing for a very long while. Yeah, it’s interesting. I don’t know how you put that into words, but when you compare a Cursor with Copilot, Copilot pretty quickly started to feel stale for some reason. Yeah, I think one thing that I think helps us is that we’re sort of doing it all in one where we’re developing the UX and the way you interact with the model at the same time as we’re developing how we actually make the model give better answers. So how you build up the prompt or how do you find the context and for a Cursor Tab, how do you train the model? So I think that helps us to have all of it the same people working on the entire experience [inaudible 00:15:17] . Yeah, it’s like the person making the UI and the person training the model sit like 18 feet away- Often the same person even. Yeah, often even the same person. You can create things that are sort of not possible if you’re not talking, you’re not experimenting. And you’re using, like you said, Cursor to write Cursor? Of course. Oh yeah. even the same person. You can create things that are sort of not possible if you’re not talking, you’re not experimenting. And you’re using, like you said, Cursor to write Cursor? Of course. Oh yeah. Well let’s talk about some of these features. Let’s talk about the all-knowing the all-powerful praise be to the Tab, auto complete on steroids basically. So how does Tab work? What is Tab? To highlight and summarize at a high level, I’d say that there are two things that Cursor is pretty good at right now. There are other things that it does, but two things that it helps programmers with. One is this idea of looking over your shoulder and being a really fast colleague who can kind of jump ahead of you and type and figure out what you’re going to do next. And that was the original idea behind… That was kind of the kernel of the idea behind a good auto complete was predicting what you’re going to do next, but you can make that concept even more ambitious by not just predicting the characters after your Cursor but actually predicting the next entire change you’re going to make, the next diff, next place you’re going to jump to. And the second thing Cursor is pretty good at right now too is helping you sometimes jump ahead of the AI and tell it what to do and go from instructions to code. And on both of those we’ve done a lot of work on making the editing experience for those things ergonomic and also making those things smart and fast. One of the things we really wanted was we wanted the model to be able to edit code for us. That was kind of a wish and we had multiple attempts at it before we had a good model that could edit code for you. Then after we had a good model, I think there’ve been a lot of effort to make the inference fast for having a good experience, and we’ve been starting to incorporate… I mean, Michael sort of mentioned this ability to jump to different places and that jump to different places I think came from a feeling of once you accept an edit, it’s like man, it should be just really obvious where to go next. It’s like I’d made this change, the model should just know that the next place to go to is 18 lines down. If you’re a WIM user, you could press 18JJ or whatever, but why am I doing this? The model should just know it. So the idea was you just press Tab, it would go 18 lines down and then show you the next edit and you would press Tab, so as long as you could keep pressing Tab. And so the internal competition was, how many Tabs can we make someone press? Once you have the idea, more abstractly, the thing to think about is how are the edits zero entropy? So once you’ve expressed your intent and the edit is… There’s no new bits of information to finish your thought, but you still have to type some characters to make the computer understand what you’re actually thinking, then maybe the model should just sort of read your mind and all the zero entropy bits should just be like tabbed away. That was sort of the abstract version. understand what you’re actually thinking, then maybe the model should just sort of read your mind and all the zero entropy bits should just be like tabbed away. That was sort of the abstract version. There’s this interesting thing where if you look at language model loss on different domains, I believe the bits per byte, which is a kind of character normalize loss for code is lower than language, which means in general there are a lot of tokens in code that are super predictable, a lot of characters that are super predictable. And this is I think even magnified when you’re not just trying to auto complete code, but predicting what the user’s going to do next in their editing of existing code. And so the goal of Cursor Tab is let’s eliminate all the low entropy actions you take inside of the editor. When the intent is effectively determined, let’s just jump you forward in time, skip you forward. Well, what’s the intuition and what’s the technical details of how to do next Cursor prediction? That jump, that’s not so intuitive I think to people. Yeah. I think I can speak to a few of the details on how to make these things work. They’re incredibly low latency, so you need to train small models on this task. In particular, they’re incredibly pre-fill token hungry. What that means is they have these really, really long prompts where they see a lot of your code and they’re not actually generating that many tokens. And so the perfect fit for that is using a sparse model, meaning an MOE model. So that was one breakthrough we made that substantially improved its performance at longer context. The other being a variant of speculative decoding that we built out called speculative edits. These are two, I think, important pieces of what make it quite high quality and very fast. Okay, so MOE [inaudible 00:20:22], the input is huge, the output is small. Yeah. Okay. So what else can you say about how to make… Does caching play a role- Oh, caching plays a huge role. Because you’re dealing with this many input tokens, if every single keystroke that you’re typing in a given line you had to rerun the model on all of those tokens passed in, you’re just going to one, significantly degrade latency, two, you’re going to kill your GPUs with load. So you need to design the actual prompts you use for the model such that they’re caching aware. And then yeah, you need to reuse the KV cache across requests just so that you’re spending less work, less compute. Again, what are the things that Tab is supposed to be able to do in the near term, just to linger on that? Generate code, fill empty space, also edit code across multiple lines and then jump to different locations inside the same file and then- Hopefully jump to different files also. So if you make an edit in one file and maybe you have to go to another file to finish your thought, it should go to the second file also.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence