Evidence receipt / belief
Published · transcript-backedSpeaker unverified: belief
6 Oct 2024 Lex Fridman Podcast #447 – Cursor Team: Future of Programming with AI
“The interesting work that I think has been done is figuring out how to properly train the process… Or the interesting work that has been open sourced and people I think talk about is how to train the process reward models, maybe in a more automated way.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- belief
- Recorded
- 6 Oct 2024
- Publisher
- Lex Fridman Podcast
Transcript context
…del that can… You train the model that is capable of doing the 99.9% of queries, then you have a way of inference time running it longer for those few people that really, really want max intelligence. How do you figure out which problem requires what level of intelligence? Is that possible to dynamically figure out when to use GPT-4, when to use a small model and when you need the o1? Yeah, that’s an open research problem, certainly. I don’t think anyone’s actually cracked this model routing problem quite well. We have initial implementations of this for something like Cursor Tab, but at the level of going between 4o sonnet to o1, it’s a bit trickier. There’s also a question like, what level of intelligence do you need to determine if the thing is too hard for the four level model? Maybe you need the o1 level model. It’s really unclear. But you mentioned this. So there’s a pre-training process then there’s post-training, and then there’s test time compute. Is that fair to separate? Where’s the biggest gains? Well, it’s weird because test time compute, there’s a whole training strategy needed to get test time compute to work. The other really weird thing about this is outside of the big labs and maybe even just OpenAI, no one really knows how it works. There’ve been some really interesting papers that show hints of what they might be doing. So perhaps they’re doing something with tree search using process reward models. But yeah, I think the issue is we don’t quite know exactly what it looks like, so it would be hard to comment on where it fits in. I would put it in post-training, but maybe the compute spent for this kind of… forgetting test time compute to work for a model is going to dwarf pre-training eventually. So we don’t even know if o1 is using just chain of thought or we don’t know how they’re using any of these? We don’t know anything? It’s fun to speculate. If you were to build a competing model, what would you do? Yeah. So one thing to do would be, I think you probably need to train a process reward model, which is… So maybe we can get into reward models and outcome reward models versus process reward models. Outcome reward models are the traditional reward models that people are trained for language modeling, and it’s just looking at the final thing. So if you’re doing some math problem, let’s look at that final thing. You’ve done everything and let’s assign a grade to it, how likely we think… What’s the reward for this outcome? Process reward models instead try to grade the chain of thought. So OpenAI had preliminary paper on this, I think, last summer where they use human labelers to get this pretty large several hundred thousand data set of creating chains of thought. Ultimately, it feels like I haven’t seen anything interesting in the ways that people use process reward models outside of just using it as a means of affecting how we choose between a bunch of samples. timately, it feels like I haven’t seen anything interesting in the ways that people use process reward models outside of just using it as a means of affecting how we choose between a bunch of samples. So what people do in all these papers is they sample a bunch of outputs from the language model, and then use the process reward models to grade all those generations alongside maybe some other heuristics and then use that to choose the best answer. The really interesting thing that people think might work and people want to work is tree search with these process reward models. Because if you really can grade every single step of the chain of thought, then you can branch out and explore multiple paths of this chain of thought and then use these process reward models to evaluate how good is this branch that you’re taking. Yeah. When the quality of the branch is somehow strongly correlated with the quality of the outcome at the very end, so you have a good model of knowing which branch to take. So not just in the short term, in the long term? Yeah. The interesting work that I think has been done is figuring out how to properly train the process… Or the interesting work that has been open sourced and people I think talk about is how to train the process reward models, maybe in a more automated way. I could be wrong here, could not be mentioning some papers. I haven’t seen anything super that seems to work really well for using the process reward models creatively to do tree search and code. This is an AI safety, maybe a bit of a philosophy question. So OpenAI says that they’re hiding the chain of thought from the user, and they’ve said that that was a difficult decision to make. Instead of showing the chain of thought, they’re asking the model to summarize the chain of thought. They’re also in the background saying they’re going to monitor the chain of thought to make sure the model is not trying to manipulate the user, which is a fascinating possibility. But anyway, what do you think about hiding the chain of thought? One consideration for OpenAI, and this is completely speculative, could be that they want to make it hard for people to distill these capabilities out of their model. It might actually be easier if you had access to that hidden chain of thought to replicate the technology, because pretty important data, like seeing the steps that the model took to get to the final results. So you could probably train on that also? hat hidden chain of thought to replicate the technology, because pretty important data, like seeing the steps that the model took to get to the final results. So you could probably train on that also? And there was a mirror situation with this, with some of the large language model providers, and also this is speculation, but some of these APIs used to offer easy access to log probabilities for all the tokens that they’re generating and also log probabilities over the prompt tokens. And then some of these APIs took those away. Again, complete speculation, but one of the thoughts is that the reason those were taken away is if you have access to log probabilities similar to this hidden chain of thought, that can give you even more information to try and distill these capabilities out of the APIs, out of these biggest models and to models you control. As an asterisk on also the previous discussion about us integrating o1, I think that we’re still learning how to use this model. So we made o1 available in Cursor because when we got the model, we were really interested in trying it out. I think a lot of programmers are going to be interested in trying it out. o1 is not part of the default Cursor experience in any way up, and we still haven’t found a way to yet integrate it into the editor in a way that we reach for every hour, maybe even every day. So I think that the jury’s still out on how to use the model, and we haven’t seen examples yet of people releasing things where it seems really clear like, oh, that’s now the use case. The obvious one to turn to is maybe this can make it easier for you to have these background things running, to have these models and loops, to have these models be agentic. But we’re still discovering, To be clear, we have ideas. We just need to try and get something incredibly useful before we put it out there. But it has these significant limitations. Even barring capabilities, it does not stream. That means it’s really, really painful to use for things where you want to supervise the output. Instead, you’re just waiting for the wall text to show up. Also, it does feel like the early innings of test time, compute and search where it’s just a very, very much a v0, and there’s so many things that don’t feel quite right. I suspect in parallel to people increasing the amount of pre-training data and the size of the models and pre-training and finding tricks there, you’ll now have this other thread of getting search to work better and better. So let me ask you about strawberry tomorrow eyes. So it looks like GitHub Copilot might be integrating o1 in some kind of way, and I think some of the comments are saying, does this mean Cursor is done? I think I saw one comment saying that. It’s a time to shut down Cursor. Yeah. Time to shut down Cursor. [inaudible 01:58:38]. Thank you. So is it time to shut down Cursor?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.