High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Karina Nguyen

Published podcast speaker

Claims
36
Episodes
2
Shows
2
Named items
1

Books, apps, and tools

The evidenced stack.

Browse the grouped index →

tool / uses

Clip

“But like the earlier prototypes was mostly like I used like Clip.”

Latent Space · 1 Feb 2025

Evidence receipt · Source ↗

Claim ledger

What Karina said.

8 transcript-backed records

01 / evaluation

Actually, I think I learned this so much from Anthropic, is people spend so much time prompting models and where quality's a really bad batch all the time, and you actually get a lot of new ideas of how do you make the model better?

“Actually, I think I learned this so much from Anthropic, is people spend so much time prompting models and where quality's a really bad batch all the time, and you actually get a lot of new ideas of how do you make the model better?”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

02 / evaluation

And we are actually getting saturated in all benchmarks. So I think the bottleneck is actually in evaluations that we don't have all the frontier, like evals like, I don't know, GPQA, which is a Google-proof question answering, PhD level intelligence.

“And we are actually getting saturated in all benchmarks. So I think the bottleneck is actually in evaluations that we don't have all the frontier, like evals like, I don't know, GPQA, which is a Google-proof question answering, PhD level intelligence.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

03 / evaluation

Because the models are so general giving something familiar to people that notifications is very familiar, having reminders is very familiar.

“Because the models are so general giving something familiar to people that notifications is very familiar, having reminders is very familiar.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

05 / evaluation

When we were launching tasks, for example, how do you make correct schedules is actually really hard for the model. But we built out some of the deterministic evaluations that is like, "Okay, if the user says 7:00 PM, the model should say 7:00 PM.

“When we were launching tasks, for example, how do you make correct schedules is actually really hard for the model. But we built out some of the deterministic evaluations that is like, "Okay, if the user says 7:00 PM, the model should say 7:00 PM.”
Speaker
Karina Nguyen
Publisher
Lenny's Podcast

06 / evaluation

3 had a lot of hallucinations actually. So I think there was like, one of the concerns is like, I don't think like the leadership was convinced, had the conviction that this is the model that you need to like, you want to like deploy or something.

“3 had a lot of hallucinations actually. So I think there was like, one of the concerns is like, I don't think like the leadership was convinced, had the conviction that this is the model that you need to like, you want to like deploy or something.”
Speaker
Karina Nguyen
Publisher
Latent Space

07 / evaluation

Like, I think like the first like 50,000 code of lines without any reviews at that time, because there's no one, um, yeah, it was like very small team.

“Like, I think like the first like 50,000 code of lines without any reviews at that time, because there's no one, um, yeah, it was like very small team.”
Speaker
Karina Nguyen
Publisher
Latent Space

08 / evaluation

I think like at that time I was like in product engineering team and then I switched to like research team and the product engineering team grew so much.

“I think like at that time I was like in product engineering team and then I switched to like research team and the product engineering team grew so much.”
Speaker
Karina Nguyen
Publisher
Latent Space
Search evidence