Evidence receipt / belief
Published · transcript-backedDan Hendrycks: belief
14 Aug 2025 Machine Learning Street Talk Superintelligence Strategy (Dan Hendrycks)
“I think you could certainly give them an under specified goal But I just don't think they could pursue that terribly coherently or, like, learn from experiments because they just have a big short term they they just have a short term memory of, like, maybe a million tokens.”
Source trail
Everything needed to verify it.
- Speaker
- Dan Hendrycks
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 14 Aug 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Yeah. I mean, I I guess another another source of my skepticism, I'm hugely inspired by Kenneth Stanley, and he's a big open ended researcher. And he had this paper out talking all about what he called a fractured and tangled representation. So when you dig into the representations of neural network models, they don't they don't really factorize the world in a parsimonious way the way we do. And and maybe I'm being anthropocentric here. Maybe I'm kind of hanging too much weight on the way we think about things. Our brain is made out of spaghetti, basically. So, you know, maybe it's a bit of an illusion that that we have these kind of factored representations. But but certainly, think from an agency point of view, this this is important. Right? Because right now, we talk about agents as being, LLMs wired in an autonomous loop that can use tools. And to me, agency is more than autonomy going in a predefined direction. It's the ability to set your own direction. And what happens now when we build these, you know, quote unquote, agentic AIs is that they don't do any you know, certainly when they set their own direction, they don't do anything particularly valuable. They require supervision constantly, certainly in terms of setting a new direction. And that makes me think of AI as as a kind of cultural technology, a bit like Photoshop or something like that. So, you know, a very creative graphic designer could use Photoshop and make beautiful images, whereas like a complete noob using Photoshop, it it just wouldn't, know, that they would just reuse the same effects, and they wouldn't create very beautiful images. And and and in a sense, like, AI sans humans is kind of like that, I think, because it doesn't have these very deep factored, you you know, representations of of the world. Would you kind of agree with that? I think for the current technology, yes. They have all those sorts of limitations, but, for it being agential, like, I think it will need to get better at planning and maintaining state across long periods of time. And then it can pursue some of these sub goals in this sort of in a vague, open ended, underspecified goal, and have that add up to something. But I think retrieving these sorts of memories of what worked, what didn't, and storing those is I think a substantial chunk if not most of what's missing on that agent picture. I think you could certainly give them an under specified goal But I just don't think they could pursue that terribly coherently or, like, learn from experiments because they just have a big short term they they just have a short term memory of, like, maybe a million tokens. And then they'll just, like, keep they'll kind of summarize this stuff, but they'll they'll start tripping over themselves in their context window because they can't maintain all that in its in its short term memory. Yes. I mean, it's another 1 of those things where philosophically, I I agree with you. If such if such a recursively improving intelligence existed, I mean, God knows just to control it, we would lose control because we would have to use another recursively improving superintelligence to control the other 1, and then we would basically just be minnows in in the grand scheme of things.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.