High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Dan Hendrycks: evaluation

14 Aug 2025 Machine Learning Street Talk Superintelligence Strategy (Dan Hendrycks)

“Because fortunately, they're not agents yet. So basically, all this research doesn't particularly matter, with the exception of, like, dual use expert level advice, because the agents aren't capable.”

— Dan Hendrycks

Source trail

Everything needed to verify it.

Speaker
Dan Hendrycks
Attribution
Verified speaker
Claim type
evaluation
Recorded
14 Aug 2025
Publisher
Machine Learning Street Talk

Transcript context

…So just quickly touching on your utility engineering paper. So so this was you you used a type of theory from econometrics utility theory to detect coherent preferences, in in LLM. So you you found that, preference coherence correlates positively with model scale, that models exhibit measurable self preservation instincts, and that political and demographic biases emerge as coherent utility functions, which is fascinating. Well, I I don't know. These are just sort of troubling signs, and maybe we'll be able to come up with methods that can really counteract these, issues. Maybe we can design models to reliably not have self preservation instincts or pressures in that direction, even though those sort of come out of from scaling somewhat. So I think it's just sort of 1 of the other very concerning hazards that we need to research and deal with and get ahead of. Because fortunately, they're not agents yet. So basically, all this research doesn't particularly matter, with the exception of, like, dual use expert level advice, because the agents aren't capable. They can't exfiltrate themselves reliably, they, or really at all. They they can't self sustain. They can't hack by themselves or more autonomously. So this is trying to identify some of these things that could be more of a problem down the line as models become more capable and try and do research to get ahead of that. But, yeah, if we if we leave that unaddressed or if we don't fix it, yeah, that's like that's potentially sufficient for a global catastrophe. Or I think, like, that would be, like, pretty sufficient. We have some self preserving AI that's really biased toward itself over people. That would be a and if it's very capable, that would I think that'd be a problem. That'd be kind of a disaster in the making. So we have various disasters in the making, though, but, hopefully hopefully, we'll get ahead of that either technically or politically. But but just digging into that a tiny bit. Mean, first of all, it was really interesting that political and demographic biases would emerge as coherent utility functions. And I I do take umbrage with this word emergence because I think in the emergence literature, there is a little bit more nuance to how machine learning people use the words. They use it to say, oh, there's just some observer relative, macroscopically surprising change in something. In in…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence