High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Dan Hendrycks: belief

24 Jun 2025 Machine Learning Street Talk Three Red Lines We're About to Cross Toward AGI (Daniel Kokotajlo, Gary Marcus, Dan Hendrycks)

“The reason this is and for fluid intelligence, such as what we see in the Arc AGI thing and Raven's regressive Matrices test, that still seems very deficient as well. So I think when thinking about these it's important to split up one's notion of cognition because if there is a severe limitation on any of these dimensions, then the models will be fairly defective economically, or at least for many tasks.”

— Dan Hendrycks

Source trail

Everything needed to verify it.

Speaker
Dan Hendrycks
Attribution
Verified speaker
Claim type
belief
Recorded
24 Jun 2025
Publisher
Machine Learning Street Talk

Transcript context

…1 thing is on benchmark based forecasting. And I think that has limitations, in particular street light effect, where well when it gets 100% that will suggest that it has the task or that suggests it solved the task. The issue is that they often have some structural defects that are only obvious later in the or when you go fairly far out into the curve. So for instance, in video understanding, you think, wow, if you looked at the benchmarks from a few years ago, they're totally at the top of them. But then you can come up a year later with several other sorts of benchmarks that challenge them. So I think that there tends to be a gravitation toward what's very tractable for AI systems and where some interesting action is happening. That's a selection pressure on the sort of benchmarks. If one's looking at cognitive tasks, such as those that you would give kids if you're testing their intelligence, for instance, if you randomly sample many of those, the the models don't do that well. Maybe on you could it'd be a double digit percentage, maybe be almost on half of them they don't do that well. For instance, count the number of faces in this photograph. O 3 can't do that very well. Or connect the dots or fill in fill in the colors in this in this picture. Just an example for for visual ability. So there's I I don't think when Gary's pointing these out, it's it's just that he's just running the cherry picking program, and what he's going do is he's just going to keep he's just going keep cherry picking and doing God of the gaps thing for AI until it basically is is AGI, I I I do think that there is a non adversarial distribution, a difference from the cognitive science angle, you would see many of these these, sorts of issues. And I I think that there's potentially some speaking past each other in part because he's not viewing intelligence as some unidimensional thing entirely. There's a lot might be correlated together, but there are various other mental faculties that are important really online that much. We saw with the GPT series that by pre training on a lot of text we got some subcomponents of intelligence. We got reading writing ability and we got a lot of crystallized intelligence or acquired knowledge from that. And that process took several years, so it's not the case, there's often people will sort of point out once it gets traction on it then it will solve it immediately or solve it very shortly thereafter. For a lot of these core cognitive abilities, they took multiple years. Mathematical reasoning would be a recent example with Minerva. Google's Minerva system got 50% on the math benchmark, some benchmark I made some while ago, in 2022. And I think we've only recently crushed it in 2025, so it took a good 3 years. And I think the reading writing ability took got to an interesting state and now relatively complete 4 years later, I think crystallized intelligence as well is now relatively complete. good 3 years. And I think the reading writing ability took got to an interesting state and now relatively complete 4 years later, I think crystallized intelligence as well is now relatively complete. However, there's still a variety of others. I wouldn't expect visual processing ability for video to be taken care of in a year or 2. Maybe audio ability would be taken care of in a year or 2, though. We don't see almost any progress on long term memory, basically. I think that they can't really that meaningfully maintain a state across long periods of times over complex interactions. And so once we start to get traction on that, then we can start to forecast that out. The reason this is and for fluid intelligence, such as what we see in the Arc AGI thing and Raven's regressive Matrices test, that still seems very deficient as well. So I think when thinking about these it's important to split up one's notion of cognition because if there is a severe limitation on any of these dimensions, then the models will be fairly defective economically, or at least for many tasks. For instance, if a person doesn't have good long term memory, it will be very difficult to enculturate them and teach them how to be productive in a working environment and load all that context. If they have very slow reaction time, that's also a problem. Or if they have low fluid intelligence, that will limit their ability to generalize substantially. So I think the number the numbers are going up substantially or continually on many of these axes for these benchmarks. So but at the same time when those benchmarks are at their conclusion there might be another peak for those, and there's many of these. And for it to be an AGI you're going to need all of them as well. And for some of them they're fairly early on in their capabilities such as with long term memory. So think if one's doing forecasting, I think understanding intelligence somewhat more and having a more sophisticated account such as 1 would find in cognitive science can help foresee bottlenecks that will actually be fairly action relevant. I mean, completely agree with you, and I'll just mention physical intelligence and visual intelligence. You mentioned visual, we left out. I skipped physical. And spatial intelligence are really very serious limits right now. So if you ask O3 to label a diagram, for example, it would be quite poor at it. If you ask it to reason about an environment, it's going to be poor.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence