Speakers in the public record
Claim mix
belief 6evaluation 5preference 3uncertainty 2commitment 1observation 1recommendation 1
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
19 published records
“Actually, I think I learned this so much from Anthropic, is people spend so much time prompting models and where quality's a really bad batch all the time, and you actually get a lot of new ideas of how do you make the model better?”
- Publisher
- Lenny's Podcast
“I think there's still a lot of multimodal research that's going on, but I think language scaled so much easier compared to multimodal because of that.”
- Publisher
- Lenny's Podcast
“I think one of my earliest inspiration is one of the last episodes of Westworld where, I don't want to spoil, but where Dolores comes to her work at that time and she comes to this new workspace and she starts writing a story.”
- Publisher
- Lenny's Podcast
“Education I think I would have dreamt if I had the tool like ChatGPT when I was young and would learn so much.”
- Publisher
- Lenny's Podcast
“I think Canvas team has still have really cool front engineers that are really people who really care about interaction, design, interacting experience.”
- Publisher
- Lenny's Podcast
“Kevin Weill/Kevin Weill, I don't know exactly how to pronounce his last name, the CPO of OpenAI.”
- Publisher
- Lenny's Podcast
“I think there was something in there that was kind like, "Oh my god, this is a really cool new capability.”
- Publisher
- Lenny's Podcast
“I know that ChatGPT had Slackbar tools. I don't know, maybe it will come back sometime.”
- Publisher
- Lenny's Podcast
“I think the form factor of Claude and Slack was kind of constrained a little bit when you want to talk about new features.”
- Publisher
- Lenny's Podcast
“And then you need to have a trade-off between what methods you want to use. And the reason why I really love relying purely on synthetic data instead of collecting data from humans is because it's much more scalable, it's cheap, less than half.”
- Publisher
- Lenny's Podcast
“And we are actually getting saturated in all benchmarks. So I think the bottleneck is actually in evaluations that we don't have all the frontier, like evals like, I don't know, GPQA, which is a Google-proof question answering, PhD level intelligence.”
- Publisher
- Lenny's Podcast
“And the most robust evals is the one where prompted baselines get the lowest score or something. And then because then you know if you're trained a good model, then it should just hill climb on that eval all the time, while not also regressing on other intelligence evals.”
- Publisher
- Lenny's Podcast
“I would say the first project at OpenAI, where researchers and applying engineers started working together from the very beginning of the product development cycle. And I think there's a lot of things that we have learned on the way, but I definitely came with the mindset of, "We need to do a really rapid model situation such that it would be much easier for engineers to work with the latest model possible, but also learn from user feedback or early internal dog food.”
- Publisher
- Lenny's Podcast
“Because the models are so general giving something familiar to people that notifications is very familiar, having reminders is very familiar.”
- Publisher
- Lenny's Podcast
“When I first came to Anthropic and I was like, "Oh my God, I really love front-end engineering." And then the reason why I switched to research is because I realized, "Oh my God, Claude is getting better at front-end.”
- Publisher
- Lenny's Podcast
“I would say what I've learned from Anthropic is this real care and craft towards model behavior, model craft, model training.”
- Publisher
- Lenny's Podcast
“It's interesting, there's these milestones that kind of open up our view of what is happening and where things are going.”
- Publisher
- Lenny's Podcast
“When we were launching tasks, for example, how do you make correct schedules is actually really hard for the model. But we built out some of the deterministic evaluations that is like, "Okay, if the user says 7:00 PM, the model should say 7:00 PM.”
- Publisher
- Lenny's Podcast
“It's a rapid model iteration for similar product outcomes. And we can dive more into it, but the way we made Canvas and tasks and new product features for ChatGPT was mostly done by synthetic training.”
- Publisher
- Lenny's Podcast