Evidence receipt / evaluation
Published · transcript-backedDiane Penn: evaluation
26 Jul 2026 Lenny's Podcast Anthropic’s first technical PM on token maxing, the jagged edge, and living in the future | Dianne Penn
“I think there's some really interesting graphs in the original scaling law papers. And I think folks are very familiar with the scaling loss in the lens of as you add in more compute and data, what's called loss, AKA the loss from next token prediction goes down.”
Source trail
Everything needed to verify it.
- Speaker
- Diane Penn
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 26 Jul 2026
- Publisher
- Lenny's Podcast
Transcript context
…So what I'm hearing here is you almost don't know what will be possible with every model release. And so the important things to focus on is being adaptable as things emerge. To your point, the product itself has to catch up to what is possible. To your point again, just like it can do so much, but people may not understand how to do it and may not be able to do it. So the product making it easy and even just telling you, here's something you could do feels like an important part. Is that roughly what you're describing? I think so. I think there's some really interesting graphs in the original scaling law papers. And I think folks are very familiar with the scaling loss in the lens of as you add in more compute and data, what's called loss, AKA the loss from next token prediction goes down. And so it's a very smooth linear curve of the models get more intelligent as you scale them up. What's actually also interesting in that paper is there are these very different emerging capability graphs. And so for example, as you add in more data and you train the models with more compute, you essentially see these actually discontinuous emerging capabilities jump. So the models go from one plus one being a thing that it can't calculate to a thing that it could reliably calculate. And so these emerging capabilities, this some nature of predictability is not necessarily everyone knows the exact moment. You need the evals to be able to assess that has actually always been a part of how this technology works. And also what makes things like safety harder. Because unless you have the evals, unless you have the systems to test, these jumps might actually happen and you don't know. That's so interesting that you may have developed this AI brain that can do something you're not even aware of. And so part of the job is just uncovering, "Wow, we just got really good at this thing. What can we do with that?"…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.