High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / commitment

Published · transcript-backed

Speaker unverified: commitment

6 Oct 2024 Lex Fridman Podcast #447 – Cursor Team: Future of Programming with AI

“I will mark this with a little red squiggly and say, you should probably review this part of the diff.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
commitment
Recorded
6 Oct 2024
Publisher
Lex Fridman Podcast

Transcript context

…he same file and then- Hopefully jump to different files also. So if you make an edit in one file and maybe you have to go to another file to finish your thought, it should go to the second file also. The full generalization is next action prediction. Sometimes you need to run a command in the terminal and it should be able to suggest the command based on the code that you wrote too, or sometimes you actually need to… It suggests something, but it’s hard for you to know if it’s correct because you actually need some more information to learn. You need to know the type to be able to verify that it’s correct. And so maybe it should actually take you to a place that’s the definition of something and then take you back so that you have all the requisite knowledge to be able to accept the next completion. So providing the human the knowledge. Yes. Right. Yeah. I just gotten to know a guy named Primeagen who I believe has an… You can order coffee via SSH. Oh yeah. We did that. We did that. So can also the model do that and provide you with caffeine? Okay. So that’s the general framework. Yeah. And the magic moment would be if… Programming is this weird discipline where sometimes the next five minutes, not always, but sometimes the next five minutes of what you’re going to do is actually predictable from the stuff you’ve done recently. And so can you get to a world where that next five minutes either happens by you disengaging and it taking you through? Or maybe a little bit more of just you seeing next step what it’s going to do and you’re like, okay, that’s good, that’s good, that’s good, that’s good, and you can just sort of tap, tap through these big changes. As we’re talking about this, I should mention one of the really cool and noticeable things about Cursor is that there’s this whole diff interface situation going on. So the model suggests with the red and the green of here’s how we’re going to modify the code, and in the chat window you can apply and it shows you the diff and you can accept the diff. So maybe can you speak to whatever direction of that? We’ll probably have four or five different kinds of diffs. So we have optimized the diff for the auto complete, so that has a different diff interface than when you’re reviewing larger blocks of code. And then we’re trying to optimize another diff thing for when you’re doing multiple different files. And at a high level, the difference is for when you’re doing auto- complete, it should be really, really fast to read. Actually it should be really fast to read in all situations, but in auto-complete your eyes are focused in one area, you can’t be in too many… The humans can’t look in too many different places. So you’re talking about on the interface side? On the interface side. So it currently has this box on this side. So we have the current box, and it you tries to delete code in some place and tries to add other code, it tries to show you a box on the side. You can maybe show it if we pull it up in Cursor.com. This is what we’re talking. So that box- Exactly here. ete code in some place and tries to add other code, it tries to show you a box on the side. You can maybe show it if we pull it up in Cursor.com. This is what we’re talking. So that box- Exactly here. It was like three or four different attempts at trying to make this thing work where first the attempt was this blue crossed out line. So before it was a box on the side, it used to show you the code to delete by showing you Google Docs style, you would see a line through it and then you would see the new code. That was super distracting. And then we tried many different… There was deletions, there was trying the red highlight. Then the next iteration of it, which is sort of funny, you would hold the, on Mac, the option button. So it would sort of highlight a region of code to show you that there might be something coming. So maybe in this example, the input and the value would all get blue. And the blue was to highlight that the AI had a suggestion for you. So instead of directly showing you the thing, it would just hint that the AI had a suggestion and if you really wanted to see it, you would hold the option button and then you would see the new suggestion. And if you release the option button, you would then see your original code. So by the way, that’s pretty nice, but you have to know to hold the option button. Yeah. And by the way, I’m not a Mac user, but I got it. Option. It’s a button I guess you people have. Again, it’s just not intuitive. I think that’s the key thing. And there’s a chance this is also not the final version of it. I am personally very excited for making a lot of improvements in this area. We often talk about it as the verification problem where these diffs are great for small edits. For large edits or when it’s multiple files or something, it’s actually a little bit prohibitive to review these diffs. So there are a couple of different ideas here. One idea that we have is, okay, parts of the diffs are important. They have a lot of information. And then parts of the diff are just very low entropy. They’re the same thing over and over again. And so maybe you can highlight the important pieces and then gray out the not so important pieces. Or maybe you can have a model that looks at the diff and sees, oh, there’s a likely bug here. I will mark this with a little red squiggly and say, you should probably review this part of the diff. Ideas in that vein I think are exciting. Yeah, that’s a really fascinating space of UX design engineering. So you’re basically trying to guide the human programmer through all the things they need to read and nothing more, optimally. And you want an intelligent model to do it. Currently, diff algorithms, they’re just like normal algorithms. There’s no intelligence. There’s intelligence that went into designing the algorithm, but then you don’t care if it’s about this thing or this thing as you want the model to do this. ormal algorithms. There’s no intelligence. There’s intelligence that went into designing the algorithm, but then you don’t care if it’s about this thing or this thing as you want the model to do this. So I think the general question is like, man, these models are going to get much smarter. As the models get much smarter, changes they will be able to propose are much bigger. So as the changes gets bigger and bigger and bigger, the humans have to do more and more and more verification work. It gets more and more and more… You need to help them out. I don’t want to spend all my time reviewing code. Can you say a little more across multiple files [inaudible 00:28:19]? Yeah. I mean, so GitHub tries to solve this with code review. When you’re doing code review, you’re reviewing multiple diffs across multiple files. But like Arvid said earlier, I think you can do much better than code review. Code review kind of sucks. You spend a lot of time trying to grok this code that’s often quite unfamiliar to you and it often doesn’t even actually catch that many bugs. And I think you can significantly improve that review experience using language models, for example, using the kinds of tricks that Arvid had described of maybe pointing you towards the regions that actually matter. I think also if the code is produced by these language models and it’s not produced by someone else… The code review experience is design for both the reviewer and the person that produced the code. In the case where the person that produced the code is a language model, you don’t have to care that much about their experience and you can design the entire thing around the reviewer such that the reviewer’s job is as fun, as easy, as productive as possible. I think that feels like the issue with just naively trying to make these things look like code review. I think you can be a lot more creative and push the boundary on what’s possible. And just one idea there is, I think ordering matters. Generally, when you review a PR, you have this list of files and you’re reviewing them from top to bottom, but actually, you actually want to understand this part first because that came logically first, and then you want to understand the next part and you don’t want to have to figure out that yourself, you want a model to. And you don’t want to have to figure out that yourself. You want a model to guide you through the thing. And is the step of creation going to be more and more natural language, is the goal versus with actual writing the book?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence