High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Speaker unverified: belief

6 Oct 2024 Lex Fridman Podcast #447 – Cursor Team: Future of Programming with AI

“I don’t think we think that that’s the case because a lot of programming, a lot of the value is in iterating, or you don’t actually want to specify something upfront because you don’t really know what you want until you have seen an initial version and then you want to iterate on that and then you provide more information.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
belief
Recorded
6 Oct 2024
Publisher
Lex Fridman Podcast

Transcript context

…that you should just do whatever is the most natural thing for you, and then our job is to figure out how do we actually retrieve the relative event things so that your thinking actually makes sense? Well, this is the discussion I had with Aravind of Perplexity is his whole idea is you should let the person be as lazy as he wants. That’s a beautiful thing, but I feel like you’re allowed to ask more of programmers, right? Yes. So if you say, “Just do what you want,” humans are lazy. There’s a tension between just being lazy versus provide more as be prompted… Almost like the system pressuring you or inspiring you to be articulate. Not in terms of the grammar of the sentences, but in terms of the depth of thoughts that you convey inside the prompts. I think even as a system gets closer to some level of perfection, often when you ask the model for something, not enough intent is conveyed to know what to do. And there are a few ways to resolve that intent. One is the simple thing of having the model just ask you, “I’m not sure how to do these parts based on your query. Could you clarify that?” I think the other could be maybe if there are five or six possible generations, “Given the uncertainty present in your query so far, why don’t we just actually show you all of those and let you pick them?” How hard is it for the model to choose to talk back versus generally… It’s hard, how deal with the uncertainty. Do I choose to ask for more information to reduce the ambiguity? So one of the things we do, it’s like a recent addition, is try to suggest files that you can add. And while you’re typing, one can guess what the uncertainty is and maybe suggest that maybe you’re writing your API and we can guess using the commits that you’ve made previously in the same file that the client and the server is super useful and there’s a hard technical problem of how do you resolve it across all commits? Which files are the most important given your current prompt? And we’re still initial version is ruled out and I’m sure we can make it much more accurate. It’s very experimental, but then the idea is we show you, do you just want to add this file, this file, this file also to tell the model to edit those files for you? Because if maybe you’re making the API, you should also edit the client and the server that is using the API and the other one resolving the API. So that would be cool as both there’s the phase where you’re writing a prompt and there’s… Before you even click, “Enter,” maybe we can help resolve some of the uncertainty. To what degree do you use agentic approaches? How useful are agents? We think agents are really, really cool. Okay. … Before you even click, “Enter,” maybe we can help resolve some of the uncertainty. To what degree do you use agentic approaches? How useful are agents? We think agents are really, really cool. Okay. I think agents, it’s like resembles like a human… You can feel that you’re getting closer to AGI because you see a demo where it acts as a human would and it’s really, really cool. I think agents are not yet super useful for many things. I think we’re getting close to where they will actually be useful. And so I think there are certain types of tasks where having an agent would be really nice. I would love to have an agent. For example, if we have a bug where you sometimes can’t Command+C and Command+V inside our chat input box, and that’s a task that’s super well specified. I just want to say in two sentences, “This does not work, please fix it.” And then I would love to have an agent that just goes off, does it, and then a day later, I come back and I review the thing. You mean it goes, finds the right file? Yeah, it finds the right files, it tries to reproduce the bug, it fixes the bug and then it verifies that it’s correct. And this could be a process that takes a long time. And so I think I would love to have that. And then I think a lot of programming, there is often this belief that agents will take over all of programming. I don’t think we think that that’s the case because a lot of programming, a lot of the value is in iterating, or you don’t actually want to specify something upfront because you don’t really know what you want until you have seen an initial version and then you want to iterate on that and then you provide more information. And so for a lot of programming, I think you actually want a system that’s instant, that gives you an initial version instantly back and then you can iterate super, super quickly. What about something like that recently came out, replica agent, that does also setting up the development environment and solving software packages, configuring everything, configuring the databases and actually deploying the app. Is that also in the set of things you dream about? I think so. I think that would be really cool. For certain types of programming, it would be really cool. Is that within scope of Cursor? Yeah, we aren’t actively working on it right now, but it’s definitely… We want to make the programmer’s life easier and more fun and some things are just really tedious and you need to go through a bunch of steps and you want to delegate that to an agent. And then some things you can actually have an agent in the background while you’re working. Let’s say you have a PR that’s both backend and frontend, and you’re working in the frontend and then you can have a background agent that doesn’t work and figure out what you’re doing. And then when you get to the backend part of your PR, then you have some initial piece of code that you can iterate on. And so that would also be really cool. ’t work and figure out what you’re doing. And then when you get to the backend part of your PR, then you have some initial piece of code that you can iterate on. And so that would also be really cool. One of the things we already talked about is speed, but I wonder if we can just linger on that some more in the various places that the technical details involved in making this thing really fast. So every single aspect of Cursor, most aspects of Cursor feel really fast. Like I mentioned, the Apply is probably the slowest thing. And for me from… I’m sorry, the pain on Arvid’s face as I say that. I know. It’s a pain. It’s a pain that we’re feeling and we’re working on fixing it. Yeah, it says something that feels… I don’t know what it is, like one second or two seconds, that feels slow. That means that actually shows that everything else is just really, really fast. So is there some technical details about how to make some of these models, how to make the chat fast, how to make the diffs fast? Is there something that just jumps to mind? Yeah. So we can go over a lot of the strategies that we use. One interesting thing is cache warming. And so what you can do is if as the user’s typing, you can have… You’re probably going to use some piece of context and you can know that before the user’s done typing. So as we discussed before, reusing the KV cache results in lower latency, lower costs, cross requests. So as the user starts typing, you can immediately warm the cache with let’s say the current file contents, and then when they press enter, there’s very few tokens it actually has to pre-fill and compute before starting the generation. This will significantly lower TTFT. Can you explain how KV cache works? Yeah, so the way transformers work. I like it. One of the mechanisms that allow transformers to not just independently… The mechanism that allows transformers to not just independently look at each token, but see previous tokens are the keys and values to attention. And generally, the way attention works is you have at your current token some query, and then you’ve all the keys and values of all your previous tokens, which are some kind of representation that the model stores internally of all the previous tokens in the prompt. And by default, when you’re doing a chat, the model has to, for every single token, do this forward pass through the entire model. That’s a lot of matrix multiplies that happen, and that is really, really slow. Instead, if you have already done that and you stored the keys and values and you keep that in the GPU, then when I… Let’s say I have to sort it for the last N tokens. If I now want to compute the output token for the N+1nth token, I don’t need to pass those first N tokens through the entire model because I already have all those keys and values. And so you just need to do the forward pass through that last token. And then when you’re doing attention, you’re reusing those keys and values that have been computed, which is the only kind of sequential part or sequentially dependent part of the transformer. Is there higher level caching of caching of the prompts or that kind of stuff that could help?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence