Evidence receipt / preference
Published · transcript-backedKevin Weil: preference
10 Apr 2025 Lenny's Podcast OpenAI’s CPO on how AI changes must-have skills, moats, coding, startup playbooks, more | Kevin Weil (CPO at OpenAI, ex-Instagram, Twitter)
“I certainly have better ideas when I get in a room and brainstorm with other people because they think differently than me. So anyways, there's just all these situations where you can actually reason about it like a group of humans or an individual human and it works, which I don't know, maybe I shouldn't have been surprised but I was.”
Source trail
Everything needed to verify it.
- Speaker
- Kevin Weil
- Attribution
- Verified speaker
- Claim type
- preference
- Recorded
- 10 Apr 2025
- Publisher
- Lenny's Podcast
Transcript context
…I don't know, maybe I should have expected this, but one of the things that's been funny for me is the extent to which you're trying to figure out how some product should work with AI, or even why some AI thing happens to be true, you can often reason about it the way you would reason about another human and it works. So maybe a couple examples. When we were first launching our reasoning model, we were the first to build a model that could reason, that could, instead of giving you just a quick system one answer right away to every question you asked, it was the third Emperor of the Holy Roman Empire, here's an answer. You could ask it hard questions and it would reason. The same way that if I asked you to do a crossword puzzle, you couldn't just snap fill in everything. You would be, "Well, okay. On this one across, I think it could be one of these two, but that means there's an A here. So that one has to be this, away, back track, step-by-step build up from where you are." Same way you answer any difficult logistical problem, any scientific problem. So this reasoning breakthrough was big, but it was also the first time that a model needed to sit and think. And that's a weird paradigm for a consumer product. You don't normally have something where you might need to hang out for 25 seconds after you ask a question. So we were trying to figure out what's the UI for this? With deep research where the model's going to go and think for 25 minutes sometimes, it's actually not that hard because you're not going to sit and watch it for 25 minutes. You're going to go do something else. You're going to go to another tab or go get lunch or whatever, and then you'll come back and it's done when it's like 20, 25 seconds or 10 seconds, it's a long time to wait, but it's not long enough to go to do something else. So you can think, if you asked me something that I needed to think for 20 seconds to answer, what would I do? I wouldn't just go mute and not say anything and shut down for 20 seconds and then come back. So we shouldn't do that. We shouldn't just have a slider sitting there. That's annoying. But I also wouldn't just start babbling every single thought that I had. So we probably shouldn't just expose the whole chain of thought as the model's thinking, but I might go like, "That's a good question. All right." I might approach it like that and then think. You're maybe giving little updates and that's actually what we ended up shipping. You have similar things where you can find situations where you get better thinking sometimes out of a group of models that all try and attack the same problem, and then you have a model that's looking at all their outputs and integrating it and then giving you a single answer at the end. I mean, sounds a little bit like brainstorming. I certainly have better ideas when I get in a room and brainstorm with other people because they think differently than me. a single answer at the end. I mean, sounds a little bit like brainstorming. I certainly have better ideas when I get in a room and brainstorm with other people because they think differently than me. So anyways, there's just all these situations where you can actually reason about it like a group of humans or an individual human and it works, which I don't know, maybe I shouldn't have been surprised but I was. That is so interesting because when I see these models operate, I never even thought about you guys designing that experience. To me, it just feels like this is what the LLM does. It just sits there and tells me what it's thinking. And I love this point you're making of let's make it feel like a human operating and well, how does a human operate? Well, they just talk aloud. They think, here's the thing I should explore. And I love that deep sequence to the extreme of that where they're just like, "Here's everything I'm doing and thinking." And people actually like that too, I guess. Was that surprising to you, "Maybe that could work too. People seem to like everything?"…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.