High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Beth Barnes

Published podcast speaker

Claims
13
Episodes
1
Shows
1
Named items
1

Books, apps, and tools

The evidenced stack.

Browse the grouped index →

person / likes

Melanie Mitchell

“There there has been a bit of an obsession, I think, with, headline accuracy when we do evaluations so that I'm a huge fan of Melanie Mitchell, for example, when she speaks about construct validity.”

Machine Learning Street Talk · 4 May 2026

Evidence receipt · Source ↗

Claim ledger

What Beth said.

4 transcript-backed records

01 / prediction

Like, that that is the the thing that we're trying to predict. And then the question is, like, how do we, you know, how how can we predict that given the observations we do have of, like, well, we've never put it in that situation, and, you know, we just have this, behavior, which is maybe indistinguishable between, oh, it was a totally nice model doing, you know, what we wanted, and it's just gonna continue to do what we want in a kind of predictable way versus, ah, yes.

“Like, that that is the the thing that we're trying to predict. And then the question is, like, how do we, you know, how how can we predict that given the observations we do have of, like, well, we've never put it in that situation, and, you know, we just have this, behavior, which is maybe indistinguishable between, oh, it was a totally nice model doing, you know, what we wanted, and it's just gonna continue to do what we want in a kind of predictable way versus, ah, yes.”
Speaker
Beth Barnes
Publisher
Machine Learning Street Talk

02 / prediction

You know, the the main reason you expect bad code to be bad is, like, you can't actually build something that sophisticated because it you know, you get bugs, you it's all too complicated, and you can't figure out how to fix it. So in some sense, if we see models building things that do actually work that are very complicated, we you know, it's like, well, it's less interesting exactly how they're doing that, but it it it's maybe bad for for human observability, and it also maybe gets into this thing of you know, we expect models to be able to do much better at well specified tasks, we sort of, you know, to the extent that we have things that we can measure, the those things will go up.

“You know, the the main reason you expect bad code to be bad is, like, you can't actually build something that sophisticated because it you know, you get bugs, you it's all too complicated, and you can't figure out how to fix it. So in some sense, if we see models building things that do actually work that are very complicated, we you know, it's like, well, it's less interesting exactly how they're doing that, but it it it's maybe bad for for human observability, and it also maybe gets into this thing of you know, we expect models to be able to do much better at well specified tasks, we sort of, you know, to the extent that we have things that we can measure, the those things will go up.”
Speaker
Beth Barnes
Publisher
Machine Learning Street Talk

03 / prediction

I give a slightly different number, but, yeah, it it I'm like, this seems very unlikely to happen this year, but it's not, you know, not un unlikely enough to rule out. And I think that basically looks like maybe we we would see accelerating trend in time horizon on it like like, easily held climbable tasks, and it turns out that was actually a, you know, a much more general capability, and that was just, you know, a bit of something you needed to do to sort of, like, elicit, it on on these less less health and climateable tasks, but sort of, you know, fundamentally, they they are using the same capabilities in a model that was just sort of, you know, what what you trained on that was affecting the difference we're seeing.

“I give a slightly different number, but, yeah, it it I'm like, this seems very unlikely to happen this year, but it's not, you know, not un unlikely enough to rule out. And I think that basically looks like maybe we we would see accelerating trend in time horizon on it like like, easily held climbable tasks, and it turns out that was actually a, you know, a much more general capability, and that was just, you know, a bit of something you needed to do to sort of, like, elicit, it on on these less less health and climateable tasks, but sort of, you know, fundamentally, they they are using the same capabilities in a model that was just sort of, you know, what what you trained on that was affecting the difference we're seeing.”
Speaker
Beth Barnes
Publisher
Machine Learning Street Talk

04 / prediction

Like, I think 1 thing we've done less of is is sort of being like, oh, I think, you know, the real bottleneck is this, like, I don't know, some some, like, you know, reasoning about novel some, you know, some specific skill, and you're like, oh, we're gonna build a benchmark to capture that and, like, target that because that's the, like, real thing that humans can do that models can't.

“Like, I think 1 thing we've done less of is is sort of being like, oh, I think, you know, the real bottleneck is this, like, I don't know, some some, like, you know, reasoning about novel some, you know, some specific skill, and you're like, oh, we're gonna build a benchmark to capture that and, like, target that because that's the, like, real thing that humans can do that models can't.”
Speaker
Beth Barnes
Publisher
Machine Learning Street Talk
Search evidence