Evidence receipt / evaluation
Published · transcript-backedChip Huyen: evaluation
23 Oct 2025 Lenny's Podcast Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix)
“We go from a text chatbot to voice chatbot. It's like the consoles are completely different because now with voice chatbot, we need to think about latency because I think multiple steps, first have voice to text, text to text, text question into text answer and then text to voice answer.”
Source trail
Everything needed to verify it.
- Speaker
- Chip Huyen
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 23 Oct 2025
- Publisher
- Lenny's Podcast
Transcript context
…I think in a lot of organizations they don't move that fast, but at the same time, they move faster than I expected because again, I think it's like bias and don't work with dinosaur companies who don't care. I think a lot of executives who come to me are very forward-looking. So maybe for me, I'm very biased towards organizations is move fast. So yeah, I think one big change I see just in organizational structure. I think this a lot of value plays in... So before we have a lot of disjointed teams. We have very clear engineering team, product team, but then there's a question of who should write eval? Who should own the metrics? And it turns out, eval, it's not a separate problem. It's a system problem because you need to look into different components, how they interact with each other. You need user behaviors because you need to know what users care about so that you can write eval reflect what users care about. So all of that you can sort it from you look into different component architectures, place guardrails and stuff. So it's just engineering, but understanding users is what product. So because of a lot of things and eval is extremely important. So the kind of bring product team and engineering team, even marketing team like user acquisition, very close to each other. So yes, since in a ways if people are structuring, so that's more communications between previously very distinct functions. Another thing is I also see as teams, of course, I think about what can be automated in the next few years and what work cannot be automated. And I seen that people already shedding, actually it's a little bit scary to think about it, but I also think it's the teams, they would've told me, it's just like okay, this is good and you and me, but we have got rid of these functions for a lot of things like previously outsourced, for example. Traditionally, it's a business outsourcing that's not core to them and can be in a more systematized. So with that, you can actually use AI to automate a lot of that. And so as a separation people thinking more of what is the value of junior engineers or senior engineers, how should we restructure engineering org for that? Yeah, so I do definitely think that is one thing to successful organization. People are just moving pieces around and thinking about use cases, whether you need to spin out new use cases and who would lead a new effort. That is one big change. Another thing in terms of AI, I think there's, I'm not sure how true this is. I guess, I'm also on the camp of thinking that it has merit, is a camp of okay, base models we have probably not quite maxed out, but we're unlikely to see really, really strong, crazily strong model. So you remember when we have GPT, right? And then GPT2, which is a big step up, an [inaudible 01:00:49] better than GPT and then GPT3, which much, much bigger than GPT4, much, much bigger. rong model. So you remember when we have GPT, right? And then GPT2, which is a big step up, an [inaudible 01:00:49] better than GPT and then GPT3, which much, much bigger than GPT4, much, much bigger. And then of course, GPT5, but it's GPT5, that scale of much bigger step jump compared to the previous, I think it's debatable. So I think that we had disappointment, the base model performance improvement is not going to be mind-blowing. It was in the last three years. So I think there's a lot of improvements when I see in the post-training phase, in the application building phase. And yes, also I think that's where I feel I would see a lot of improvement there. I also very interest in multimodality. So we've seen a lot of text base, but I think there's a lot of audio, videos use cases that is very, very exciting. And I think audios is not quite as solved. Well, I think because I do work with a couple of voice startups and when it comes to, think about voice, it's an entirely different beast. So let's say have chatbot. We go from a text chatbot to voice chatbot. It's like the consoles are completely different because now with voice chatbot, we need to think about latency because I think multiple steps, first have voice to text, text to text, text question into text answer and then text to voice answer. So you have multiple hops and latency become very important. And there's a question, what does it make you sound natural? So for example, people think of in AI and humans, when humans talk to each other, if I say, you try to interrupt me and say, Chip [inaudible 01:02:36]. I would pause and I try to hear you out. But sometime even if I just like say some word, like acknowledge when I, mm-hmm, mm-hmm, that I shouldn't stop. It's just continue. So the question of forced interruption and whether it's, should I stop or not, it's a big in what perceived as natural conversations. And that's also regulations because a lot of time, people want to build AI chatbot, voice chatbots that sound like humans, try to trick users into thinking that they're talking to humans, but also maybe potential regulation saying okay, you have to disclose to users when you talk, if the bot is talking to is human or AI. So I think this a whole space, I think it's not quite as solved as you think. But it's not quite like an AI foundation model problem because a human interruption detection, it's actually a classical machining problem. It's a different framing, but you can give classifier for that. Or the question of latency, actually a massive engineering challenge, not an AI challenge. Of course, it can be an AI challenge because people are trying to build voice-to-voice model. So instead of having to firstly transcribe the voice from me into text and then get a model [inaudible 01:03:54] text answer and get another model should turn from text to speech, you can just do voice-to-voice directly. So that is something we're working on, but it's very hard. Yeah. So yeah, so even audio, I think of it's the easier than video because video have both image and voice. It's already pretty hard. So I think there's a lot of challenges in that space. That was an awesome list of things. Let me mirror them back real quick. So what you're predicting in the next few years, things that will change in the way we work, and these actually resonate with so many conversations I've had on this podcast. So says, just kind of doubling down on where things are heading. One is the blurring of lines between different functions instead of just design engineering. Everyone's going to be doing a lot of different things now. Two is, just more of work being automated with agents and all these AI tools and just in theory, productivity going up. Third is, a shifting from pre-training models to post-training, fine-tuning and things like that because to your point, models maybe are slowing down in how smart they're getting. Although, I'll point folks to the, I had a chat with the co-founder of Anthropic. He made a really good point here. He's like, we're really bad at understanding what exponentials feel like when we're in the middle of that. And also, models are being released more often. So the difference between them we may not notice because they're just happening more often versus GPT3 came out a year before after GPT2. Maybe true, maybe not. And then the fourth point you made is this idea of multimodal, investing in multimodal experiences. I cannot wait for ChatGPT voice mode to get better at interruption, exactly what you're saying. I'm just talking to it and then someone makes a little sound and it's like [inaudible 01:05:33]. Okay. And then you have to, and then it's like, and then it stops talking. It's so annoying.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.