person / likes
Melanie Mitchell
“There there has been a bit of an obsession, I think, with, headline accuracy when we do evaluations so that I'm a huge fan of Melanie Mitchell, for example, when she speaks about construct validity.”
Public evidence record
Published podcast speaker
Books, apps, and tools
person / likes
“There there has been a bit of an obsession, I think, with, headline accuracy when we do evaluations so that I'm a huge fan of Melanie Mitchell, for example, when she speaks about construct validity.”
Claim ledger
13 transcript-backed records
01 / observation
“We actually ended up with a negative correlation between years of experience and, and and performance because, like, the the sort of people who are in network, like, our friends who are kind of doing really well, and the people who are more qualified were, like, actually not doing that great.”
02 / observation
“there is a sort of adversarial selection going on where people are trying to make some benchmark, like, cheaply subject to the constraint that current models do badly on it, which means you, you know, you can't use a lot of expensive human labor, it has to be something that's either automatically checkable or that you can, like, create with kind of cheap human labor.”
03 / belief
“Think he's he's probably, like, you know, more more confident on some things where I'm I'm more uncertain, and I think some of the, like, AI futures project models are, like, more sensitive to the meteor time horizon metrics than they should be.”
04 / belief
“I think may maybe there's some difference in, like, how much you've you think that this has improved between, like, g b d 2 and where we are now, or I would say sort of in, you know, the amount of adapting to new things that are happening that models could do now does seem like it's much higher and, you know, they're much better at, like, editing their own scaffolding or or, you know, sort of reasoning about their, like, you know, their their sort of embodiment, like, this thing of, like, knowing not to kill your own process or no like, you know, stuff stuff like that where it's like there is like, yes, they are limited, but there's also, some trend of improvement.”
05 / belief
“Like, these 2 things can coexist, and I think people often sort of you know, positions are surprisingly correlated on some axis of, like, how, you know, how soon you think AIs or how good you think AIs or something.”
06 / belief
“I think people u usually use scheming to refer to the model is doing what it's currently doing in in service of some long term goal and is deliberately doing things like appearing aligned or or getting a high score, like, in service of, you know, eventually accomplishing that goal versus you can be reward hacking both you know, you could be reward hacking in some extremely dumb way, like the boat example where it's just like, this is what RL kind of found or, like, this is what, you know, like, a star search found.”
07 / belief
“I think, you know, there's a good chance that this, you know, makes our lives a lot better or a lot worse, and people disagree, you know, about even what what current models can do, let alone where we're heading.”
08 / prediction
“Like, that that is the the thing that we're trying to predict. And then the question is, like, how do we, you know, how how can we predict that given the observations we do have of, like, well, we've never put it in that situation, and, you know, we just have this, behavior, which is maybe indistinguishable between, oh, it was a totally nice model doing, you know, what we wanted, and it's just gonna continue to do what we want in a kind of predictable way versus, ah, yes.”
09 / prediction
“You know, the the main reason you expect bad code to be bad is, like, you can't actually build something that sophisticated because it you know, you get bugs, you it's all too complicated, and you can't figure out how to fix it. So in some sense, if we see models building things that do actually work that are very complicated, we you know, it's like, well, it's less interesting exactly how they're doing that, but it it it's maybe bad for for human observability, and it also maybe gets into this thing of you know, we expect models to be able to do much better at well specified tasks, we sort of, you know, to the extent that we have things that we can measure, the those things will go up.”
10 / evaluation
“You know, operationalizing, what do we care about in the definition of intelligence is that it, like, you know, allows us to predict, like, how will the moles affect the world and what you know, you know, predict what will happen and know how to how to handle them well and things.”
11 / preference
“There there has been a bit of an obsession, I think, with, headline accuracy when we do evaluations so that I'm a huge fan of Melanie Mitchell, for example, when she speaks about construct validity.”
12 / prediction
“I give a slightly different number, but, yeah, it it I'm like, this seems very unlikely to happen this year, but it's not, you know, not un unlikely enough to rule out. And I think that basically looks like maybe we we would see accelerating trend in time horizon on it like like, easily held climbable tasks, and it turns out that was actually a, you know, a much more general capability, and that was just, you know, a bit of something you needed to do to sort of, like, elicit, it on on these less less health and climateable tasks, but sort of, you know, fundamentally, they they are using the same capabilities in a model that was just sort of, you know, what what you trained on that was affecting the difference we're seeing.”
13 / prediction
“Like, I think 1 thing we've done less of is is sort of being like, oh, I think, you know, the real bottleneck is this, like, I don't know, some some, like, you know, reasoning about novel some, you know, some specific skill, and you're like, oh, we're gonna build a benchmark to capture that and, like, target that because that's the, like, real thing that humans can do that models can't.”