Evidence receipt / observation
Published · transcript-backedTim Scarfe: observation
18 Oct 2025 Machine Learning Street Talk The Secret Engine of AI - Prolific [Sponsored] (Sara Saab, Enzo Blindow)
“I am amenable by the way to this idea of loss of control, you know, which is that we we we start to build systems on top of systems on top of systems, and it's a little bit like the power station.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- observation
- Recorded
- 18 Oct 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…An interesting thing that I pulled out of that Value Compass paper by Shen et al. Is the misalignment between what AIs think they are and what we think AIs are as people. And AIs think, or aspire to be, I'm going to use provocative language, autonomous thinkers. And research found that humans don't want that, right? So I think the reason I bring that up is the way you evaluate a helpful system is kind of, as you're saying, Tim, that sort of impossible problem of covering every test case in an infinite algorithm, which we will never do. But the way you evaluate a person or thinker, an autonomous being, we have loads of examples for in the world, right? Jury trials, right? Nobody expects that the moral behavior of a human is all predetermined when it's born, and we know exactly what right or wrong looks like in every case. We have loads of social structure for evaluating the behavior and agency of a person. And I do think that we kind of have to stop sort of equivocating between or maybe nobody's equivocating but me. I think I think we should assume we're building towards AGI or superintelligence or thinking creatures and work backwards as opposed to trying to sort of box in systems. Yes. Because when I was reading that anthropic paper about the agentic misalignment, 1 of my thoughts was that these are individual agents. And as you were just saying, Sara, in the real world, we know that we are fallible, which is why we build error correction systems, know. You know, 1 person can't press the red button to to drop a nuke. And we have juries and and we have, you know, various forms of of collectives to to overcome individual errors. And I'm guessing we could do the same thing with AIs. I'm not sure what that anthropic experiment would look like if you had a supervisor. You know, the KGB, they had a an expression trust but verify. You know, so you can almost have a a supervisor agent and you can have a committee of agents. But then we're almost getting into even murkier territory, right, because we're building these inscrutable things. I am amenable by the way to this idea of loss of control, you know, which is that we we we start to build systems on top of systems on top of systems, and it's a little bit like the power station. You can't just turn off a power station when when we start to increasingly, like, rely on all of this stuff. But but in principle, though, do do you think that building some kind of agentic network could overcome some of these alignment problems? I think certainly there will be networks and conditional layers when it comes to evaluation, oversight, and monitoring. We're already seeing it. I mean, this is very, very much sort of mainstream already, right? So your evals of your model will be done by sort of automated benchmarks or LLM as a judge or some kind of oracle or reward function or something, and then there will be, whether it's considered human evaluation or not, some human will verify something along the chain. And there's this kind of orchestration that's emerging between machines and people in the space of evals and oversight. So I think it feels pretty uncontroversial to say that there will be, you know, layered and orchestrated approaches like that. But what that looks like when you kind of push the dial to 12, I'm not sure.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.