Evidence receipt / evaluation
Published · transcript-backedDan Hendrycks: evaluation
14 Aug 2025 Machine Learning Street Talk Superintelligence Strategy (Dan Hendrycks)
“I think on the I think, generally, I think the political problems, the incentives, the giving people things to do that are incentive payable is is where more of the value is at compared to on the technical side.”
Source trail
Everything needed to verify it.
- Speaker
- Dan Hendrycks
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 14 Aug 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…AI alignment is famously difficult. It's 1 of the most intractable challenges perhaps of of a generation. And, some some things that people think of as as alignment like RLHF, for example, I'm I'm sure you would agree with the statement that it it's something that makes models behave as if they are aligned, but perhaps it's it's not really aligning them in in the way that we would want to. And, you know, just emphatically, in the next year, I mean, if you could solve a single problem in alignment, what would it be, and what impact would it have? I think on the I think, generally, I think the political problems, the incentives, the giving people things to do that are incentive payable is is where more of the value is at compared to on the technical side. I would guess if there's a way to reliably get them to tell the truth, for instance, or make them reliably honest, that would be, I think that would be very valuable. And it would be solved such that it wouldn't have a severe trade off. Like, wouldn't be much more expensive to run, or it wouldn't tank its performance in other axes. Like, it wouldn't trade off on its crystallized knowledge, for instance. But I think that would be very valuable, having it be, not overtly lie. Because then you could build standards around that as well. I don't think anybody would say like, if if you could make them very reliably not lie, then I think people would, it'd be reasonable for people to make demands that AI's not lying to them. On on on that, because I I've I've read in in your papers about this concept of deception and and lying and and whatnot. And in a sense, I think you might be projecting mentalistic properties onto AI models. So, you know, the the that they have beliefs and that that they have thinking and and so on. And, I mean, just sort of thinking critically, like, what what makes you think that we can think of them as having beliefs and telling lies?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.