Evidence receipt / prediction
Published · transcript-backedSpeaker unverified: prediction
6 Oct 2024 Lex Fridman Podcast #447 – Cursor Team: Future of Programming with AI
“Meaning, it’s going to be hard to continue scaling up this regime. So scaling up test time compute is an interesting way, if now increasing the number of inference time flops that we use but still getting… Yeah, as you increase the number of flops you use inference time getting corresponding improvements in the performance of these models.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- prediction
- Recorded
- 6 Oct 2024
- Publisher
- Lex Fridman Podcast
Transcript context
…nitely something we want to get a lot better at. I think that there are a lot of cool ideas to try there, both on the learning better retrieval systems, like better embedding models, better rerankers. I think that there are also cool academic ideas, stuff we’ve tried out internally, but also the field is grappling with writ large about, can you get language models to a place where you can actually just have the model itself understand a new corpus of information? The most popular talked about version of this is can you make the context windows infinite? Then if you make the context windows infinite, can you make the model actually pay attention to the infinite context? And then after you can make it pay attention to the infinite context to make it somewhat feasible to actually do it, can you then do caching for that infinite context? You don’t have to recompute that all the time. But there are other cool ideas that are being tried, that are a little bit more analogous to fine-tuning of actually learning this information in the weights of the model. It might be that you actually get a qualitative lead different type of understanding if you do it more at the weight level than if you do it at the in-context learning level. I think the jury’s still a little bit out on how this is all going to work in the end? But in the interim, us as a company, we are really excited about better retrieval systems and picking the parts of the code base that are most relevant to what you’re doing, and we could do that a lot better. One interesting proof of concept for the learning this knowledge directly in the weights is with VS Code. So we’re in a VS Code fork and VS Code. The code is all public. So these models in pre-training have seen all the code. They’ve probably also seen questions and answers about it. And then they’ve been fine-tuned and RLHFed to be able to answer questions about code in general. So when you ask it a question about VS Code, sometimes it’ll hallucinate, but sometimes it actually does a pretty good job at answering the question. I think this is just by… It happens to be okay, but what if you could actually specifically train or post-train a model such that it really was built to understand this code base? It’s an open research question, one that we’re quite interested in. And then there’s also uncertainty of, do you want the model to be the thing that end-to-end is doing everything, i.e. it’s doing the retrieval in its internals and then answering a question, creating the code, or do you want to separate the retrieval from the frontier model, where maybe you’ll get some really capable models that are much better than the best open source ones in a handful of months? And then you’ll want to separately train a really good open source model to be the retriever, to be the thing that feeds in the context to these larger models. Can you speak a little more to post-training a model to understand the code base? What do you mean by that? Is this a synthetic data direction? Is this- at feeds in the context to these larger models. Can you speak a little more to post-training a model to understand the code base? What do you mean by that? Is this a synthetic data direction? Is this- Yeah, there are many possible ways you could try doing it. There’s certainly no shortage of ideas. It’s just a question of going in and trying all of them and being empirical about which one works best. One very naive thing is to try to replicate what’s done with VS Code and these frontier models. So let’s continue pre-training. Some kind of continued pre-training that includes general code data but also throws in of the data of some particular repository that you care about. And then in post-training, meaning in… Let’s just start with instruction fine-tuning. You have a normal instruction fine-tuning data set about code. Then you throw in a lot of questions about code in that repository. So you could either get ground truth ones, which might be difficult or you could do what you hinted at or suggested using synthetic data, i.e. having the model ask questions about various recent pieces of the code. So you take the pieces of the code, then prompt the model or have a model propose a question for that piece of code, and then add those as instruction fine-tuning data points. And then in theory, this might unlock the model’s ability to answer questions about that code base. Let me ask you about OpenAI o1. What do you think is the role of that kind of test time compute system in programming? I think test time compute is really, really interesting. So there’s been the pre-training regime which will, as you scale up the amount of data and the size of your model, get you better and better performance both on loss and then on downstream benchmarks and just general performance. So we use it for coding or other tasks. We’re starting to hit a bit of a data wall. Meaning, it’s going to be hard to continue scaling up this regime. So scaling up test time compute is an interesting way, if now increasing the number of inference time flops that we use but still getting… Yeah, as you increase the number of flops you use inference time getting corresponding improvements in the performance of these models. Traditionally, we just had to literally train a bigger model that always used that many more flops, but now, we could perhaps use the same size model and run it for longer to be able to get an answer at the quality of a much larger model. So the really interesting thing I like about this is there are some problems that perhaps require 100 trillion parameter model intelligence trained on 100 trillion tokens. But that’s maybe 1%, maybe 0.1% of all queries. So are you going to spend all of this effort, all of this compute training a model that costs that much and then run it so infrequently? It feels completely wasteful when instead you get the model that can… You train the model that is capable of doing the 99.9% of queries, then you have a way of inference time running it longer for those few people that really, really want max intelligence. del that can… You train the model that is capable of doing the 99.9% of queries, then you have a way of inference time running it longer for those few people that really, really want max intelligence. How do you figure out which problem requires what level of intelligence? Is that possible to dynamically figure out when to use GPT-4, when to use a small model and when you need the o1? Yeah, that’s an open research problem, certainly. I don’t think anyone’s actually cracked this model routing problem quite well. We have initial implementations of this for something like Cursor Tab, but at the level of going between 4o sonnet to o1, it’s a bit trickier. There’s also a question like, what level of intelligence do you need to determine if the thing is too hard for the four level model? Maybe you need the o1 level model. It’s really unclear. But you mentioned this. So there’s a pre-training process then there’s post-training, and then there’s test time compute. Is that fair to separate? Where’s the biggest gains? Well, it’s weird because test time compute, there’s a whole training strategy needed to get test time compute to work. The other really weird thing about this is outside of the big labs and maybe even just OpenAI, no one really knows how it works. There’ve been some really interesting papers that show hints of what they might be doing. So perhaps they’re doing something with tree search using process reward models. But yeah, I think the issue is we don’t quite know exactly what it looks like, so it would be hard to comment on where it fits in. I would put it in post-training, but maybe the compute spent for this kind of… forgetting test time compute to work for a model is going to dwarf pre-training eventually. So we don’t even know if o1 is using just chain of thought or we don’t know how they’re using any of these? We don’t know anything? It’s fun to speculate. If you were to build a competing model, what would you do? Yeah. So one thing to do would be, I think you probably need to train a process reward model, which is… So maybe we can get into reward models and outcome reward models versus process reward models. Outcome reward models are the traditional reward models that people are trained for language modeling, and it’s just looking at the final thing. So if you’re doing some math problem, let’s look at that final thing. You’ve done everything and let’s assign a grade to it, how likely we think… What’s the reward for this outcome? Process reward models instead try to grade the chain of thought. So OpenAI had preliminary paper on this, I think, last summer where they use human labelers to get this pretty large several hundred thousand data set of creating chains of thought. Ultimately, it feels like I haven’t seen anything interesting in the ways that people use process reward models outside of just using it as a means of affecting how we choose between a bunch of samples.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.