Evidence receipt / preference
Published · transcript-backedBeth Barnes: preference
4 May 2026 Machine Learning Street Talk The AI Models Smart Enough to Know They're Cheating — Beth Barnes & David Rein [METR]
“There there has been a bit of an obsession, I think, with, headline accuracy when we do evaluations so that I'm a huge fan of Melanie Mitchell, for example, when she speaks about construct validity.”
Source trail
Everything needed to verify it.
- Speaker
- Beth Barnes
- Attribution
- Verified speaker
- Claim type
- preference
- Recorded
- 4 May 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…not good enough? For me, it was mostly thinking about this problem of scalable oversight. As models get more capable, it just gets harder to evaluate their capabilities. If we imagine that models are able to complete tasks that take people a long time to complete or require expertise that you don't necessarily have. You need a method for still being confident in their outputs and trusting their outputs. And so thinking about that problem was a lot of motivation actually for for GPQA and, was what kind of started, got me started thinking about, evaluations. To me, I'd say there's some big picture thing of thinking that AI seems important and sort of navigating it well seems important. And, you know, clearly, we we don't have a great understanding of what is going on with that. And, you know, people generally disagreeing very, strongly about what what to expect. And maybe if there's a particular moment informing time horizon, maybe just the the sense that people really couldn't agree on what the capabilities of current models are, let alone extrapolating to the future and trying to sort of think about how could you characterize, the ways in which models are and aren't highly capable and when it's sort of like, okay. And, you know, in some sense, they're they're expert level, at some kinds of things, like like question answering, and in some sense, they're below average human at at some other actually being useful somehow. You know? But the the I guess, a 0.1 few years ago where it sort of in theory, the benchmarks say that they're PhD level, but when you try to do anything, it's like, this isn't helpful. There there has been a bit of an obsession, I think, with, headline accuracy when we do evaluations so that I'm a huge fan of Melanie Mitchell, for example, when she speaks about construct validity. And she had a really good blog post out recently. And she said that there are 4 big problems, right? So there's like data contamination where the benchmark appears in the training data. Approximate retrieval where the LLMs interpolate from similar training examples without possessing the actual capability to come up with it themselves. Shortcuts, so doing the right things for the wrong reasons. And just more more broadly, not really testing for things like consistency and and robustness and generalization or or the mechanism. It's so much focused just on on the accuracy itself. I mean, how how do you folks think about those kind of problems with benchmark?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.
Named in this claim