High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Beth Barnes: belief

4 May 2026 Machine Learning Street Talk The AI Models Smart Enough to Know They're Cheating — Beth Barnes & David Rein [METR]

“I think, you know, there's a good chance that this, you know, makes our lives a lot better or a lot worse, and people disagree, you know, about even what what current models can do, let alone where we're heading.”

— Beth Barnes

Source trail

Everything needed to verify it.

Speaker
Beth Barnes
Attribution
Verified speaker
Claim type
belief
Recorded
4 May 2026
Publisher
Machine Learning Street Talk

Transcript context

…Let's get few quality examples in. Let's get the right humans in to get the right quality of human feedback in. So so we're we're trying to make human data or human feedback. We treat it as an infrastructure We try to make it accessible. We make it cheaper. We effectively democratize access to this data. Yeah. So, yeah, I'm I'm super excited to talk, to you, Tim, about, the time horizon graph and and and and meter. Yeah. I think the world does not have a good understanding of what is happening with AI, and I think it should have a better understanding. I think, you know, there's a good chance that this, you know, makes our lives a lot better or a lot worse, and people disagree, you know, about even what what current models can do, let alone where we're heading. So, at MEDA, we're, you know, trying to give the world a kind of better understanding of what is up with AI capabilities and and risks and and forecasts. Have a bunch of different research angles on this, both on the pessimistic and optimistic or positive and negative estimations of capabilities, and excited to to talk about that. I'm so excited about having you both on. So, you both have incredibly impressive backgrounds. So, Beth, you were, an ex OpenAI alignment researcher, and, you started ArchieVowse in 2022 with Paul Cristiano, and you spun out, you know, that that out as METER in in, December 2023. You've been featured on the Times Top 100 AI Profiles. And David, you're the creator of the GPQA, so the graduate level Google proof QA benchmark, which is used by every single major AI lab as a capability benchmark. And you're the co author on HCAST, which we'll talk about today, and the Time Horizons paper and the Developer Productivity RCT. Incredible to have you both in here. But maybe we should just start as a bit of a question to both of you. So Beth, you left OpenAI to build META. What was the moment that each of you realized that existing evaluation approaches were fundamentally…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence