Evidence receipt / recommendation
Published · transcript-backedSpeaker unverified: recommendation
6 Oct 2024 Lex Fridman Podcast #447 – Cursor Team: Future of Programming with AI
“I think for the more aggressive things where you’re making larger changes that take longer periods of time, you’ll probably want to do this in some sandbox remote environment and that’s another incredibly tricky problem of how do you exactly reproduce or mostly reproduce to the point of it being effectively equivalent for running code the user’s environment with this remote sandbox.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- recommendation
- Recorded
- 6 Oct 2024
- Publisher
- Lex Fridman Podcast
Transcript context
…ct, you kind of need that. If you don’t have that, it’s really hard to code in a big project. Can you say, again, how that’s being used inside Cursor, the language server protocol communication thing? So it’s being used in Cursor to show to the programmer just like in VS Code, but then the idea is you want to show that same information to the models, the IM models, and you want to do that in a way that doesn’t affect the user because you want to do it in background. And so the idea behind the Shadow Workspace was, okay, one way we can do this is we spawn a separate window of Cursor that’s hidden, and so you can set this flag in it and like turn it’s hidden. There is a window but you don’t actually see it. And inside of this window, the AI agents can modify code however they want as long as they don’t save it because it’s still the same folder and then can get feedback from the linters and go to definition and iterate on their code. So literally run everything in the background as if… Right, maybe even run the code. So that’s the eventual version and that’s what you want. And a lot of the blog post is actually about how do you make that happen because it’s a little bit tricky. You want it to be on the user’s machine so that it exactly mirrors the user’s environment. And then on Linux, you can do this cool thing where you can actually mirror the file system and have the AI make changes to the files, and it thinks that it’s operating on the file level, but actually, that’s stored in memory and you can create this kernel-like extension to make it work. Whereas on Mac and Windows, it’s a little bit more difficult, but it’s a fun technical problem, so that’s why. One may be hacky but interesting idea that I like is holding a lock on saving. And so basically, you can then have the language model kind of hold the lock on saving to disk and then instead of you operating in the ground truth version of the files that are saved to disk, you actually are operating what was the Shadow Workspace before and these unsaved things that only exist in memory that you still get linter errors for, and you can code in. And then when you try to maybe run code, it’s just like there’s a small warning that there’s a lock, and then you kind of will take back the lock from the language server if you’re trying to do things concurrently or from the Shadow Workspace if you’re trying to do things concurrently. That’s such an exciting future by the way. It’s a bit of a tangent, but to allow a model to change files, it’s scary for people but it’s really cool, to be able to just let the agent do a set of tasks and you come back the next day and kind of observe like it’s a colleague or something like that. e files, it’s scary for people but it’s really cool, to be able to just let the agent do a set of tasks and you come back the next day and kind of observe like it’s a colleague or something like that. And I think there may be different versions of runability where, for the simple things where you’re doing things in the span of a few minutes on behalf of the user as they’re programming, it makes sense to make something work locally in their machine. I think for the more aggressive things where you’re making larger changes that take longer periods of time, you’ll probably want to do this in some sandbox remote environment and that’s another incredibly tricky problem of how do you exactly reproduce or mostly reproduce to the point of it being effectively equivalent for running code the user’s environment with this remote sandbox. I’m curious what kind of agents you want for coding? Do you want them to find bugs? Do you want them to implement new features? What agents do you want? So by the way, when I think about agents, I don’t think just about coding. I think so for this particular podcast, there’s video editing and a lot of… If you look in Adobe, a lot… There’s code behind. It’s very poorly documented code, but you can interact with Premiere, for example, using code, and basically all the uploading, everything I do on YouTube, everything as you could probably imagine, I do all of that through code and including translation and overdubbing, all of this. So I envision all of those kinds of tasks. So automating many of the tasks that don’t have to do directly with the editing, so that. Okay, that’s what I was thinking about. But in terms of coding, I would be fundamentally thinking about bug finding, many levels of kind of bug finding and also bug finding like logical bugs, not logical like spiritual bugs or something. Ones like big directions of implementation, that kind of stuff. Magical [inaudible 01:11:39] and bug finding. Yeah. I mean, it’s really interesting that these models are so bad at bug finding when just naively prompted to find a bug. They’re incredibly poorly calibrated. Even the smartest models. Exactly, even o1. How do you explain that? Is there a good intuition? I think these models are really strong reflection of the pre-training distribution, and I do think they generalize as the loss gets lower and lower, but I don’t think the loss and the scale is quite… The loss is low enough such that they’re really fully generalizing on code. The things that we use these things for, the frontier models that they’re quite good at, are really code generation and question answering. And these things exist in massive quantities in pre-training with all of the code in GitHub on the scale of many, many trillions of tokens and questions and answers on things like stack overflow and maybe GitHub issues. ist in massive quantities in pre-training with all of the code in GitHub on the scale of many, many trillions of tokens and questions and answers on things like stack overflow and maybe GitHub issues. And so when you try to push one of these things that really don’t exist very much online, like for example, the Cursor Tab objective of predicting the next edit given the edits done so far, the brittleness kind of shows. And then bug detection is another great example, where there aren’t really that many examples of actually detecting real bugs and then proposing fixes and the models just kind of really struggle at it. But I think it’s a question of transferring the model in the same way that you get this fantastic transfer from pre-trained models just on code in general to the Cursor Tab objective. You’ll see a very, very similar thing with generalized models that are really good at code to bug detection. It just takes a little bit of kind nudging in that direction. Look to be clear, I think they sort of understand code really well. While they’re being pre-trained, the representation that’s being built up almost certainly like somewhere in the stream, the model knows that maybe there’s something sketchy going on. It sort of has some sketchiness but actually eliciting the sketchiness to actually… Part of it is that humans are really calibrated on which bugs are really important. It’s not just actually saying there’s something sketchy. It’s like it’s this sketchy trivial, it’s this sketchy like you’re going to take the server down. Part of it is maybe the cultural knowledge of why is a staff engineer is good because they know that three years ago someone wrote a really sketchy piece of code that took the server down and as opposed to maybe you just… This thing is an experiment. So a few bugs are fine, you’re just trying to experiment and get the feel of the thing. And so if the model gets really annoying when you’re writing an experiment, that’s really bad, but if you’re writing something for super production, you’re writing a database. You’re writing code in Postgres or Linux or whatever. You’re Linus Torvalds. It’s sort of unacceptable to have even an edge case and just having the calibration of how paranoid is the user and like- But even then if you’re putting in a maximum paranoia, it still just doesn’t quite get it. Yeah, yeah. Yeah. I mean, but this is hard for humans too to understand which line of code is important, which is not. I think one of your principles on a website says if a code can do a lot of damage, one should add a comment that say, “This line of code is dangerous.” And all caps, repeated 10 times. No, you say for every single line of code inside the function you have to… And that’s quite profound, that says something about human beings because the engineers move on, even the same person might just forget how it can sink the Titanic a single function. You might not intuit that quite clearly by looking at the single piece of code.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.