Evidence receipt / belief
Published · transcript-backedSpeaker unverified: belief
6 Oct 2024 Lex Fridman Podcast #447 – Cursor Team: Future of Programming with AI
“You thought through everything, which you didn’t actually think through everything. But I think for that particular system, we’ve… So for concrete details, the thing we do is obviously we upload when… We chunk up all of your code, and then we send up the code for embedding and we embed the code.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- belief
- Recorded
- 6 Oct 2024
- Publisher
- Lex Fridman Podcast
Transcript context
…art yet, so we suspect something like this could make a lot of sense. There’s a question of whether it happens in the foreground too or if it happens in the background like what we’ve been discussing. Sure. The background’s pretty cool. I could be running the code in different ways. Plus there’s a database side to this, which how do you protect it from not modifying the database, but okay. I mean, there’s certainly cool solutions there. There’s this new API that is being developed for… It’s not in AWS, but it certainly… I think it’s in PlanetScale. I don’t know if PlanetScale was the first one to you add it. It’s this ability sort of add branches to a database, which is like if you’re working on a feature and you want to test against the broad database, but you don’t actually want to test against the broad database, you could sort of add a branch to the database. And the way they do that is they add a branch to the write-ahead log. And there’s obviously a lot of technical complexity in doing it correctly. I guess database companies need new things to do. They have good databases now. And I think turbopuffer, which is one of the databases we use, is going to add maybe branching to the write-ahead log. So maybe the AI agents will use branching, they’ll test against some branch, and it’s sort of going to be a requirement for the database to support branching or something. It would be really interesting if you could branch a file system, right? Yeah. I feel like everything needs branching. It’s like- Yeah. Yeah. The problem with the multiverse, right? If you branch on everything that’s like a lot. There’s obviously these super clever algorithms to make sure that you don’t actually use a lot of space or CPU or whatever. Okay. This is a good place to ask about infrastructure. So you guys mostly use AWS, what are some interesting details? What are some interesting challenges? Why’d you choose AWS? Why is AWS still winning? Hashtag. AWS is just really, really good. It is really good. Whenever you use an AWS product, you just know that it’s going to work. It might be absolute hell to go through the steps to set it up. Why is the interface so horrible? Because it’s- It’s just so good. It doesn’t need to- It’s the nature of winning. I think it’s exactly. It’s just nature they’re winning. Yeah, yeah. But AWS we can always trust, it will always work. And if there is a problem, it’s probably your problem. Yeah. Okay. Is there some interesting challenges to… You guys are pretty new startup to scaling, to so many people and- Yeah, I think that it has been an interesting journey adding each extra zero to the request per second. You run into all of these with the general components you’re using for caching and databases, run into issues as you make things bigger and bigger, and now we’re at the scale where we get into overflows on our tables and things like that. And then also there have been some custom systems that we’ve built. For instance, our retrieval system for computing, a semantic index of your code base and answering questions about a code base that have, continually, I feel like been one of the trickier things to scale. instance, our retrieval system for computing, a semantic index of your code base and answering questions about a code base that have, continually, I feel like been one of the trickier things to scale. … that have continually, I feel like, been one of the trickier things to scale. I have a few friends who are super senior engineers and one of their lines is, it’s very hard to predict where systems will break when you scale them. You can try to predict in advance, but there’s always something weird that’s going to happen when you add these extras here. You thought through everything, which you didn’t actually think through everything. But I think for that particular system, we’ve… So for concrete details, the thing we do is obviously we upload when… We chunk up all of your code, and then we send up the code for embedding and we embed the code. And then we store the embeddings in a database, but we don’t actually store any of the code. And then there’s reasons around making sure that we don’t introduce client bugs because we’re very, very paranoid about client bugs. We store much of the details on the server. Everything is encrypted. So one of the technical challenges is always making sure that the local index, the local code base state is the same as the state that is on the server. The way, technically, we ended up doing that is, for every single file you can keep this hash, and then for every folder you can keep a hash, which is the hash of all of its children. You can recursively do that until the top. Why do something complicated? One thing you could do is you could keep a hash for every file and every minute, you could try to download the hashes that are on the server, figure out what are the files that don’t exist on the server. Maybe you just created a new file, maybe you just deleted a file, maybe you checked out a new branch, and try to reconcile the state between the client and the server. But that introduces absolutely ginormous network overhead both on the client side. Nobody really wants us to hammer their WiFi all the time if you’re using Cursor. But also, it would introduce ginormous overhead on the database. It would be reading these tens of terabytes database, approaching 20 terabytes or something data base every second. That’s just crazy. You definitely don’t want to do that. So what you do, you just try to reconcile the single hash, which is at the root of the project. And then if something mismatches, then you go, you find where all the things disagree. Maybe you look at the children and see if the hashes match. If the hashes don’t match, go look at their children and so on. But you only do that in the scenario where things don’t match. For most people, most of the time, the hashes match. So it’s like a hierarchical reconciliation- Yeah. … of hashes- Something like that. Yeah, it’s called a Merkle tree. Yeah, Merkle. Yeah. Yeah, this is cool to see that you have to think through all these problems. a hierarchical reconciliation- Yeah. … of hashes- Something like that. Yeah, it’s called a Merkle tree. Yeah, Merkle. Yeah. Yeah, this is cool to see that you have to think through all these problems. The reason it’s gotten hard is just because the number of people using it and some of your customers have really, really large code bases to the point where… We originally reordered dark code base, which is big, but it’s just not the size of some company that’s been there for 20 years and has a ginormous number of files and you want to scale that across programmers. There’s all these details where building the simple thing is easy, but scaling it to a lot of people, a lot of companies is obviously a difficult problem, which is independent of, actually… so that there’s part of this scaling. Our current solution is also coming up with new ideas that, obviously, we’re working on, but then scaling all of that in the last few weeks, months. Yeah. There are a lot of clever things, additional things that go into this indexing system. For example, the bottleneck in terms of costs is not soaring things in the vector database or the database. It’s actually embedding the code. You don’t want to re-embed the code base for every single person in a company that is using the same exact code except for maybe they’re a different branch with a few different files or they’ve made a few local changes. Because again, embeddings are the bottleneck, you can do this one clever trick and not have to worry about the complexity of dealing with branches and the other databases where you just have some cash on the actual vectors computed from the hash of a given chunk. So this means that when the nth person at a company goes and embed their code base, it’s really, really fast. You do all this without actually storing any code on our servers at all. No code data is stored. We just store the vectors in the vector database and the vector cache. What’s the biggest gains at this time you get from indexing the code base? Just out of curiosity, what benefit do users have? It seems like longer term, there’ll be more and more benefit, but in the short term, just asking questions of the code base, what’s the usefulness of that? I think the most obvious one is just, you want to find out where something is happening in your large code base, and you have a fuzzy memory of, “Okay, I want to find the place where we do X,” but you don’t exactly know what to search for in a normal text search. So you ask a chat, you hit command enter to ask with the code base chat. And then very often, it finds the right place that you were thinking of. Like you mentioned, in the future, I think there’s only going to get more and more powerful, where we’re working a lot on improving the quality of our retrieval. I think the ceiling for that is really, really much higher than people give the credit for.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.