Evidence receipt / evaluation
Published · transcript-backedSpeaker unverified: evaluation
6 Oct 2024 Lex Fridman Podcast #447 – Cursor Team: Future of Programming with AI
“There’s this interesting thing where if you look at language model loss on different domains, I believe the bits per byte, which is a kind of character normalize loss for code is lower than language, which means in general there are a lot of tokens in code that are super predictable, a lot of characters that are super predictable.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- evaluation
- Recorded
- 6 Oct 2024
- Publisher
- Lex Fridman Podcast
Transcript context
…even the same person. You can create things that are sort of not possible if you’re not talking, you’re not experimenting. And you’re using, like you said, Cursor to write Cursor? Of course. Oh yeah. Well let’s talk about some of these features. Let’s talk about the all-knowing the all-powerful praise be to the Tab, auto complete on steroids basically. So how does Tab work? What is Tab? To highlight and summarize at a high level, I’d say that there are two things that Cursor is pretty good at right now. There are other things that it does, but two things that it helps programmers with. One is this idea of looking over your shoulder and being a really fast colleague who can kind of jump ahead of you and type and figure out what you’re going to do next. And that was the original idea behind… That was kind of the kernel of the idea behind a good auto complete was predicting what you’re going to do next, but you can make that concept even more ambitious by not just predicting the characters after your Cursor but actually predicting the next entire change you’re going to make, the next diff, next place you’re going to jump to. And the second thing Cursor is pretty good at right now too is helping you sometimes jump ahead of the AI and tell it what to do and go from instructions to code. And on both of those we’ve done a lot of work on making the editing experience for those things ergonomic and also making those things smart and fast. One of the things we really wanted was we wanted the model to be able to edit code for us. That was kind of a wish and we had multiple attempts at it before we had a good model that could edit code for you. Then after we had a good model, I think there’ve been a lot of effort to make the inference fast for having a good experience, and we’ve been starting to incorporate… I mean, Michael sort of mentioned this ability to jump to different places and that jump to different places I think came from a feeling of once you accept an edit, it’s like man, it should be just really obvious where to go next. It’s like I’d made this change, the model should just know that the next place to go to is 18 lines down. If you’re a WIM user, you could press 18JJ or whatever, but why am I doing this? The model should just know it. So the idea was you just press Tab, it would go 18 lines down and then show you the next edit and you would press Tab, so as long as you could keep pressing Tab. And so the internal competition was, how many Tabs can we make someone press? Once you have the idea, more abstractly, the thing to think about is how are the edits zero entropy? So once you’ve expressed your intent and the edit is… There’s no new bits of information to finish your thought, but you still have to type some characters to make the computer understand what you’re actually thinking, then maybe the model should just sort of read your mind and all the zero entropy bits should just be like tabbed away. That was sort of the abstract version. understand what you’re actually thinking, then maybe the model should just sort of read your mind and all the zero entropy bits should just be like tabbed away. That was sort of the abstract version. There’s this interesting thing where if you look at language model loss on different domains, I believe the bits per byte, which is a kind of character normalize loss for code is lower than language, which means in general there are a lot of tokens in code that are super predictable, a lot of characters that are super predictable. And this is I think even magnified when you’re not just trying to auto complete code, but predicting what the user’s going to do next in their editing of existing code. And so the goal of Cursor Tab is let’s eliminate all the low entropy actions you take inside of the editor. When the intent is effectively determined, let’s just jump you forward in time, skip you forward. Well, what’s the intuition and what’s the technical details of how to do next Cursor prediction? That jump, that’s not so intuitive I think to people. Yeah. I think I can speak to a few of the details on how to make these things work. They’re incredibly low latency, so you need to train small models on this task. In particular, they’re incredibly pre-fill token hungry. What that means is they have these really, really long prompts where they see a lot of your code and they’re not actually generating that many tokens. And so the perfect fit for that is using a sparse model, meaning an MOE model. So that was one breakthrough we made that substantially improved its performance at longer context. The other being a variant of speculative decoding that we built out called speculative edits. These are two, I think, important pieces of what make it quite high quality and very fast. Okay, so MOE [inaudible 00:20:22], the input is huge, the output is small. Yeah. Okay. So what else can you say about how to make… Does caching play a role- Oh, caching plays a huge role. Because you’re dealing with this many input tokens, if every single keystroke that you’re typing in a given line you had to rerun the model on all of those tokens passed in, you’re just going to one, significantly degrade latency, two, you’re going to kill your GPUs with load. So you need to design the actual prompts you use for the model such that they’re caching aware. And then yeah, you need to reuse the KV cache across requests just so that you’re spending less work, less compute. Again, what are the things that Tab is supposed to be able to do in the near term, just to linger on that? Generate code, fill empty space, also edit code across multiple lines and then jump to different locations inside the same file and then- Hopefully jump to different files also. So if you make an edit in one file and maybe you have to go to another file to finish your thought, it should go to the second file also. he same file and then- Hopefully jump to different files also. So if you make an edit in one file and maybe you have to go to another file to finish your thought, it should go to the second file also. The full generalization is next action prediction. Sometimes you need to run a command in the terminal and it should be able to suggest the command based on the code that you wrote too, or sometimes you actually need to… It suggests something, but it’s hard for you to know if it’s correct because you actually need some more information to learn. You need to know the type to be able to verify that it’s correct. And so maybe it should actually take you to a place that’s the definition of something and then take you back so that you have all the requisite knowledge to be able to accept the next completion. So providing the human the knowledge. Yes. Right. Yeah. I just gotten to know a guy named Primeagen who I believe has an… You can order coffee via SSH. Oh yeah. We did that. We did that. So can also the model do that and provide you with caffeine? Okay. So that’s the general framework. Yeah. And the magic moment would be if… Programming is this weird discipline where sometimes the next five minutes, not always, but sometimes the next five minutes of what you’re going to do is actually predictable from the stuff you’ve done recently. And so can you get to a world where that next five minutes either happens by you disengaging and it taking you through? Or maybe a little bit more of just you seeing next step what it’s going to do and you’re like, okay, that’s good, that’s good, that’s good, that’s good, and you can just sort of tap, tap through these big changes. As we’re talking about this, I should mention one of the really cool and noticeable things about Cursor is that there’s this whole diff interface situation going on. So the model suggests with the red and the green of here’s how we’re going to modify the code, and in the chat window you can apply and it shows you the diff and you can accept the diff. So maybe can you speak to whatever direction of that? We’ll probably have four or five different kinds of diffs. So we have optimized the diff for the auto complete, so that has a different diff interface than when you’re reviewing larger blocks of code. And then we’re trying to optimize another diff thing for when you’re doing multiple different files. And at a high level, the difference is for when you’re doing auto- complete, it should be really, really fast to read. Actually it should be really fast to read in all situations, but in auto-complete your eyes are focused in one area, you can’t be in too many… The humans can’t look in too many different places. So you’re talking about on the interface side? On the interface side. So it currently has this box on this side. So we have the current box, and it you tries to delete code in some place and tries to add other code, it tries to show you a box on the side. You can maybe show it if we pull it up in Cursor.com. This is what we’re talking. So that box- Exactly here.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.