High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Karina Nguyen

Published podcast speaker

Claims
36
Episodes
2
Shows
2
Named items
1

Books, apps, and tools

The evidenced stack.

Browse the grouped index →

tool / uses

Clip

“But like the earlier prototypes was mostly like I used like Clip.”

Latent Space · 1 Feb 2025

Evidence receipt · Source ↗

Claim ledger

What Karina said.

36 transcript-backed records

01 / evaluation

Actually, I think I learned this so much from Anthropic, is people spend so much time prompting models and where quality's a really bad batch all the time, and you actually get a lot of new ideas of how do you make the model better?

“Actually, I think I learned this so much from Anthropic, is people spend so much time prompting models and where quality's a really bad batch all the time, and you actually get a lot of new ideas of how do you make the model better?”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

02 / belief

I think there's still a lot of multimodal research that's going on, but I think language scaled so much easier compared to multimodal because of that.

“I think there's still a lot of multimodal research that's going on, but I think language scaled so much easier compared to multimodal because of that.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

03 / belief

I think one of my earliest inspiration is one of the last episodes of Westworld where, I don't want to spoil, but where Dolores comes to her work at that time and she comes to this new workspace and she starts writing a story.

“I think one of my earliest inspiration is one of the last episodes of Westworld where, I don't want to spoil, but where Dolores comes to her work at that time and she comes to this new workspace and she starts writing a story.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

05 / belief

I think Canvas team has still have really cool front engineers that are really people who really care about interaction, design, interacting experience.

“I think Canvas team has still have really cool front engineers that are really people who really care about interaction, design, interacting experience.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

09 / preference

And then you need to have a trade-off between what methods you want to use. And the reason why I really love relying purely on synthetic data instead of collecting data from humans is because it's much more scalable, it's cheap, less than half.

“And then you need to have a trade-off between what methods you want to use. And the reason why I really love relying purely on synthetic data instead of collecting data from humans is because it's much more scalable, it's cheap, less than half.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

10 / evaluation

And we are actually getting saturated in all benchmarks. So I think the bottleneck is actually in evaluations that we don't have all the frontier, like evals like, I don't know, GPQA, which is a Google-proof question answering, PhD level intelligence.

“And we are actually getting saturated in all benchmarks. So I think the bottleneck is actually in evaluations that we don't have all the frontier, like evals like, I don't know, GPQA, which is a Google-proof question answering, PhD level intelligence.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

11 / preference

And the most robust evals is the one where prompted baselines get the lowest score or something. And then because then you know if you're trained a good model, then it should just hill climb on that eval all the time, while not also regressing on other intelligence evals.

“And the most robust evals is the one where prompted baselines get the lowest score or something. And then because then you know if you're trained a good model, then it should just hill climb on that eval all the time, while not also regressing on other intelligence evals.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

12 / commitment

I would say the first project at OpenAI, where researchers and applying engineers started working together from the very beginning of the product development cycle. And I think there's a lot of things that we have learned on the way, but I definitely came with the mindset of, "We need to do a really rapid model situation such that it would be much easier for engineers to work with the latest model possible, but also learn from user feedback or early internal dog food.

“I would say the first project at OpenAI, where researchers and applying engineers started working together from the very beginning of the product development cycle. And I think there's a lot of things that we have learned on the way, but I definitely came with the mindset of, "We need to do a really rapid model situation such that it would be much easier for engineers to work with the latest model possible, but also learn from user feedback or early internal dog food.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

13 / evaluation

Because the models are so general giving something familiar to people that notifications is very familiar, having reminders is very familiar.

“Because the models are so general giving something familiar to people that notifications is very familiar, having reminders is very familiar.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

14 / preference

When I first came to Anthropic and I was like, "Oh my God, I really love front-end engineering." And then the reason why I switched to research is because I realized, "Oh my God, Claude is getting better at front-end.

“When I first came to Anthropic and I was like, "Oh my God, I really love front-end engineering." And then the reason why I switched to research is because I realized, "Oh my God, Claude is getting better at front-end.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

16 / evaluation

When we were launching tasks, for example, how do you make correct schedules is actually really hard for the model. But we built out some of the deterministic evaluations that is like, "Okay, if the user says 7:00 PM, the model should say 7:00 PM.

“When we were launching tasks, for example, how do you make correct schedules is actually really hard for the model. But we built out some of the deterministic evaluations that is like, "Okay, if the user says 7:00 PM, the model should say 7:00 PM.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

17 / recommendation

It's a rapid model iteration for similar product outcomes. And we can dive more into it, but the way we made Canvas and tasks and new product features for ChatGPT was mostly done by synthetic training.

“It's a rapid model iteration for similar product outcomes. And we can dive more into it, but the way we made Canvas and tasks and new product features for ChatGPT was mostly done by synthetic training.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

18 / evaluation

3 had a lot of hallucinations actually. So I think there was like, one of the concerns is like, I don't think like the leadership was convinced, had the conviction that this is the model that you need to like, you want to like deploy or something.

“3 had a lot of hallucinations actually. So I think there was like, one of the concerns is like, I don't think like the leadership was convinced, had the conviction that this is the model that you need to like, you want to like deploy or something.”
Speaker
Karina Nguyen
Publisher
Latent Space

20 / belief

I think on OpenAI there was a team called Model Design. It's John, the PM. She's leading that team and I work very closely with those teams that we were working on, like actually writing improvements that we did with ChaiGPT last year.

“I think on OpenAI there was a team called Model Design. It's John, the PM. She's leading that team and I work very closely with those teams that we were working on, like actually writing improvements that we did with ChaiGPT last year.”
Speaker
Karina Nguyen
Publisher
Latent Space

22 / belief

I feel like we are bottlenecked by like human creativity on like completely changing the way we think about the internet or like some of the, the way you think about software, like AI right now is pushes us to like rethink everything that we've done before in my view.

“I feel like we are bottlenecked by like human creativity on like completely changing the way we think about the internet or like some of the, the way you think about software, like AI right now is pushes us to like rethink everything that we've done before in my view.”
Speaker
Karina Nguyen
Publisher
Latent Space

24 / belief

Like writing use case in terms of like, how do we make it easier for people to format or like do some of the editing? I think there is still a lot of room for improvement, to be honest.

“Like writing use case in terms of like, how do we make it easier for people to format or like do some of the editing? I think there is still a lot of room for improvement, to be honest.”
Speaker
Karina Nguyen
Publisher
Latent Space

25 / belief

I think in terms of like other operating system, I think a lot about it these days, it's just like, if you're entering this like task oriented, like operating system or something, where also a generative OS, like in my opinion, like people in like few years will click on like websites way less.

“I think in terms of like other operating system, I think a lot about it these days, it's just like, if you're entering this like task oriented, like operating system or something, where also a generative OS, like in my opinion, like people in like few years will click on like websites way less.”
Speaker
Karina Nguyen
Publisher
Latent Space

26 / belief

I think that ideally we learn from like the user behavior and ideally the model will just be more proactive in suggesting of like, oh, I can either do this for you every day because I've observed that you do that every day or something.

“I think that ideally we learn from like the user behavior and ideally the model will just be more proactive in suggesting of like, oh, I can either do this for you every day because I've observed that you do that every day or something.”
Speaker
Karina Nguyen
Publisher
Latent Space

28 / belief

What you're looking for, basically the model will be, will have a much easier time to like kind of like select the candidates and match like the candidate that is most like fulfilled the criteria that you gave. And I think there's a class of problems like this that O1 excels at.

“What you're looking for, basically the model will be, will have a much easier time to like kind of like select the candidates and match like the candidate that is most like fulfilled the criteria that you gave. And I think there's a class of problems like this that O1 excels at.”
Speaker
Karina Nguyen
Publisher
Latent Space

29 / uncertainty

I think like the way you train the model is like very different, obviously. But I think what I've learned is that like you will end up with like, I don't know, like 70 models and every model will have its own like brain damage.

“I think like the way you train the model is like very different, obviously. But I think what I've learned is that like you will end up with like, I don't know, like 70 models and every model will have its own like brain damage.”
Speaker
Karina Nguyen
Publisher
Latent Space

31 / belief

I think like right now like with Canvas we are thinking about like this paradigm of like real time collaboration to like asynchronous collaboration.

“I think like right now like with Canvas we are thinking about like this paradigm of like real time collaboration to like asynchronous collaboration.”
Speaker
Karina Nguyen
Publisher
Latent Space

32 / recommendation

Will knows this. So you actually need to like go back to like, I don't know, like GPT-4 model card and like read the appendix just to like make sure that like the settings are the same as you're running the settings too.

“Will knows this. So you actually need to like go back to like, I don't know, like GPT-4 model card and like read the appendix just to like make sure that like the settings are the same as you're running the settings too.”
Speaker
Karina Nguyen
Publisher
Latent Space

33 / evaluation

Like, I think like the first like 50,000 code of lines without any reviews at that time, because there's no one, um, yeah, it was like very small team.

“Like, I think like the first like 50,000 code of lines without any reviews at that time, because there's no one, um, yeah, it was like very small team.”
Speaker
Karina Nguyen
Publisher
Latent Space

34 / evaluation

I think like at that time I was like in product engineering team and then I switched to like research team and the product engineering team grew so much.

“I think like at that time I was like in product engineering team and then I switched to like research team and the product engineering team grew so much.”
Speaker
Karina Nguyen
Publisher
Latent Space

35 / preference

Like we don't know how each other's like working style, like what I prefer, what do they prefer, how do they prefer to communicate, et cetera, et cetera.

“Like we don't know how each other's like working style, like what I prefer, what do they prefer, how do they prefer to communicate, et cetera, et cetera.”
Speaker
Karina Nguyen
Publisher
Latent Space
Search evidence