High Signal Podcasts Evidence ledger
Method
Browse
← All source episodes

Lenny's Podcast / episode intelligence

OpenAI researcher on why soft skills are the future of work | Karina Nguyen (Research at OpenAI, ex-Anthropic)

9 Feb 2025 19 published claims 2 attributable people

Speakers in the public record

Claim mix

belief 6evaluation 5preference 3uncertainty 2commitment 1observation 1recommendation 1

Evidence policy

Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.

Claim ledger

The useful parts, with receipts.

19 published records

01 / evaluation

Actually, I think I learned this so much from Anthropic, is people spend so much time prompting models and where quality's a really bad batch all the time, and you actually get a lot of new ideas of how do you make the model better?

“Actually, I think I learned this so much from Anthropic, is people spend so much time prompting models and where quality's a really bad batch all the time, and you actually get a lot of new ideas of how do you make the model better?”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

02 / belief

I think there's still a lot of multimodal research that's going on, but I think language scaled so much easier compared to multimodal because of that.

“I think there's still a lot of multimodal research that's going on, but I think language scaled so much easier compared to multimodal because of that.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

03 / belief

I think one of my earliest inspiration is one of the last episodes of Westworld where, I don't want to spoil, but where Dolores comes to her work at that time and she comes to this new workspace and she starts writing a story.

“I think one of my earliest inspiration is one of the last episodes of Westworld where, I don't want to spoil, but where Dolores comes to her work at that time and she comes to this new workspace and she starts writing a story.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

05 / belief

I think Canvas team has still have really cool front engineers that are really people who really care about interaction, design, interacting experience.

“I think Canvas team has still have really cool front engineers that are really people who really care about interaction, design, interacting experience.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

10 / preference

And then you need to have a trade-off between what methods you want to use. And the reason why I really love relying purely on synthetic data instead of collecting data from humans is because it's much more scalable, it's cheap, less than half.

“And then you need to have a trade-off between what methods you want to use. And the reason why I really love relying purely on synthetic data instead of collecting data from humans is because it's much more scalable, it's cheap, less than half.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

11 / evaluation

And we are actually getting saturated in all benchmarks. So I think the bottleneck is actually in evaluations that we don't have all the frontier, like evals like, I don't know, GPQA, which is a Google-proof question answering, PhD level intelligence.

“And we are actually getting saturated in all benchmarks. So I think the bottleneck is actually in evaluations that we don't have all the frontier, like evals like, I don't know, GPQA, which is a Google-proof question answering, PhD level intelligence.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

12 / preference

And the most robust evals is the one where prompted baselines get the lowest score or something. And then because then you know if you're trained a good model, then it should just hill climb on that eval all the time, while not also regressing on other intelligence evals.

“And the most robust evals is the one where prompted baselines get the lowest score or something. And then because then you know if you're trained a good model, then it should just hill climb on that eval all the time, while not also regressing on other intelligence evals.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

13 / commitment

I would say the first project at OpenAI, where researchers and applying engineers started working together from the very beginning of the product development cycle. And I think there's a lot of things that we have learned on the way, but I definitely came with the mindset of, "We need to do a really rapid model situation such that it would be much easier for engineers to work with the latest model possible, but also learn from user feedback or early internal dog food.

“I would say the first project at OpenAI, where researchers and applying engineers started working together from the very beginning of the product development cycle. And I think there's a lot of things that we have learned on the way, but I definitely came with the mindset of, "We need to do a really rapid model situation such that it would be much easier for engineers to work with the latest model possible, but also learn from user feedback or early internal dog food.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

14 / evaluation

Because the models are so general giving something familiar to people that notifications is very familiar, having reminders is very familiar.

“Because the models are so general giving something familiar to people that notifications is very familiar, having reminders is very familiar.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

15 / preference

When I first came to Anthropic and I was like, "Oh my God, I really love front-end engineering." And then the reason why I switched to research is because I realized, "Oh my God, Claude is getting better at front-end.

“When I first came to Anthropic and I was like, "Oh my God, I really love front-end engineering." And then the reason why I switched to research is because I realized, "Oh my God, Claude is getting better at front-end.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

18 / evaluation

When we were launching tasks, for example, how do you make correct schedules is actually really hard for the model. But we built out some of the deterministic evaluations that is like, "Okay, if the user says 7:00 PM, the model should say 7:00 PM.

“When we were launching tasks, for example, how do you make correct schedules is actually really hard for the model. But we built out some of the deterministic evaluations that is like, "Okay, if the user says 7:00 PM, the model should say 7:00 PM.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

19 / recommendation

It's a rapid model iteration for similar product outcomes. And we can dive more into it, but the way we made Canvas and tasks and new product features for ChatGPT was mostly done by synthetic training.

“It's a rapid model iteration for similar product outcomes. And we can dive more into it, but the way we made Canvas and tasks and new product features for ChatGPT was mostly done by synthetic training.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast
Search evidence