Evidence receipt / belief
Published · transcript-backedSpeaker unverified: belief
6 Oct 2024 Lex Fridman Podcast #447 – Cursor Team: Future of Programming with AI
“Why do you think it’s different than cloud providers? Because I think a lot of this data would never have gone to the cloud providers in the first place where this is often… You want to give more data to the AI models, you want to give personal data that you would never have put online in the first place to these companies or to these models.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- belief
- Recorded
- 6 Oct 2024
- Publisher
- Lex Fridman Podcast
Transcript context
…ng to be really, really hard to run for almost all people, locally. Don’t you want the most capable model? You want [inaudible 01:38:55] too? And also o1- I like how you’re pitching me. o1 is another- Would you be satisfied with an inferior model? Listen, yes, I’m one of those, but there’s some people that like to do stuff locally, especially like… Really, there’s a whole obviously open source movement that resists. It’s good that they exist actually because you want to resist the power centers that are growing our- There’s actually an alternative to local models that I am particularly fond of. I think it’s still very much in the research stage, but you could imagine to do homomorphic encryption for language model inference. So you encrypt your input on your local machine, then you send that up, and then the server can use loss of computation. They can run models that you cannot run locally on this encrypted data, but they cannot see what the data is, and then they send back the answer and you decrypt the answer and only you can see the answer. So I think that’s still very much research and all of it is about trying to make the overhead lower because right now, the overhead is really big, but if you can make that happen, I think that would be really, really cool, and I think it would be really, really impactful because I think one thing that’s actually worrisome is that, as these models get better and better, they’re going to become more and more economically useful. So more and more of the world’s information and data will flow through one or two centralized actors. And then there are worries about, there can be traditional hacker attempts, but it also creates this scary part where if all of the world’s information is flowing through one node in plaintext, you can have surveillance in very bad ways. Sometimes that will happen for… Initially, will be good reasons. People will want to try to protect against bad actors using AI models in bad ways, and then you will add in some surveillance code. And then someone else will come in and you’re on a slippery slope, and then you start doing bad things with a lot of the world’s data. So I am very hopeful that we can solve homomorphic encryption for- Yeah, and- … language model inference. … doing privacy, preserving machine learning. But I would say, that’s the challenge we have with all software these days. It’s like there’s so many features that can be provided from the cloud and all us increasingly rely on it and make our life awesome. But there’s downsides, and that’s why you rely on really good security to protect from basic attacks. But there’s also only a small set of companies that are controlling that data, and they obviously have leverage and they could be infiltrated in all kinds of ways. That’s the world we live in. So it’s- ut there’s also only a small set of companies that are controlling that data, and they obviously have leverage and they could be infiltrated in all kinds of ways. That’s the world we live in. So it’s- Yeah, the thing I’m just actually quite worried about is the world where… Anthropic has this responsible scaling policy where we’re the low ASLs, which is the Anthropic security level or whatever of the models. But as we get to ASL-3, ASL-4, whatever models which are very powerful… But for mostly reasonable security reasons, you would want to monitor all the prompts. But I think that’s reasonable and understandable where everyone is coming from. But man, it’d be really horrible if all the world’s information is monitored that heavily, it’s way too centralized. It’s like this really fine line you’re walking where on the one side, you don’t want the models to go rogue. On the other side, humans like… I don’t know if I trust all the world’s information to pass through three model providers. Why do you think it’s different than cloud providers? Because I think a lot of this data would never have gone to the cloud providers in the first place where this is often… You want to give more data to the AI models, you want to give personal data that you would never have put online in the first place to these companies or to these models. It also centralizes control where right now, for cloud, you can often use your own encryption keys, and AWS can’t really do much. But here, it’s just centralized actors that see the exact plain text of everything. Yeah. On the topic of a context, that’s actually been a friction for me. When I’m writing code in Python, there’s a bunch of stuff imported. You could probably intuit the kind of stuff I would like to include in the context. How hard is it to auto figure out the context? It’s tricky. I think we can do a lot better at computing the context automatically in the future. One thing that’s important to note is, there are trade-offs with including automatic context. So the more context you include for these models, first of all, the slower they are and the more expensive those requests are, which means you can then do less model calls and do less fancy stuff in the background. Also, for a lot of these models, they get confused if you have a lot of information in the prompt. So the bar for accuracy and for relevance of the context you include should be quite high. Already, we do some automatic context in some places within the product. It’s definitely something we want to get a lot better at. I think that there are a lot of cool ideas to try there, both on the learning better retrieval systems, like better embedding models, better rerankers. nitely something we want to get a lot better at. I think that there are a lot of cool ideas to try there, both on the learning better retrieval systems, like better embedding models, better rerankers. I think that there are also cool academic ideas, stuff we’ve tried out internally, but also the field is grappling with writ large about, can you get language models to a place where you can actually just have the model itself understand a new corpus of information? The most popular talked about version of this is can you make the context windows infinite? Then if you make the context windows infinite, can you make the model actually pay attention to the infinite context? And then after you can make it pay attention to the infinite context to make it somewhat feasible to actually do it, can you then do caching for that infinite context? You don’t have to recompute that all the time. But there are other cool ideas that are being tried, that are a little bit more analogous to fine-tuning of actually learning this information in the weights of the model. It might be that you actually get a qualitative lead different type of understanding if you do it more at the weight level than if you do it at the in-context learning level. I think the jury’s still a little bit out on how this is all going to work in the end? But in the interim, us as a company, we are really excited about better retrieval systems and picking the parts of the code base that are most relevant to what you’re doing, and we could do that a lot better. One interesting proof of concept for the learning this knowledge directly in the weights is with VS Code. So we’re in a VS Code fork and VS Code. The code is all public. So these models in pre-training have seen all the code. They’ve probably also seen questions and answers about it. And then they’ve been fine-tuned and RLHFed to be able to answer questions about code in general. So when you ask it a question about VS Code, sometimes it’ll hallucinate, but sometimes it actually does a pretty good job at answering the question. I think this is just by… It happens to be okay, but what if you could actually specifically train or post-train a model such that it really was built to understand this code base? It’s an open research question, one that we’re quite interested in. And then there’s also uncertainty of, do you want the model to be the thing that end-to-end is doing everything, i.e. it’s doing the retrieval in its internals and then answering a question, creating the code, or do you want to separate the retrieval from the frontier model, where maybe you’ll get some really capable models that are much better than the best open source ones in a handful of months? And then you’ll want to separately train a really good open source model to be the retriever, to be the thing that feeds in the context to these larger models. Can you speak a little more to post-training a model to understand the code base? What do you mean by that? Is this a synthetic data direction? Is this-…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.