Evidence receipt / belief
Published · transcript-backedGary Marcus: belief
24 Jun 2025 Machine Learning Street Talk Three Red Lines We're About to Cross Toward AGI (Daniel Kokotajlo, Gary Marcus, Dan Hendrycks)
“Was going say the last thing, which is I first noted the problem with chess I think 2 years ago, the illegal move problem.”
Source trail
Everything needed to verify it.
- Speaker
- Gary Marcus
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 24 Jun 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…So I totally agree with those limitations about current systems. But the thing that I would say is that it seems like the trends are going up. So I would predict that if you measured this sort of thing for the last 5 years, even though the current systems might still have weaknesses, they would be less weak than the systems of 2 years ago, which are less Well, chess is an example, then I'm going go to you. Was going say the last thing, which is I first noted the problem with chess I think 2 years ago, the illegal move problem. And really I think Matthew Acre was first to Archer was first to point it out. And I spread it on Twitter and said, look, this is a serious problem. And it persists. There are a lot of problems that I feel like have persisted. But your turn. You've been fine. 1 thing is on benchmark based forecasting. And I think that has limitations, in particular street light effect, where well when it gets 100% that will suggest that it has the task or that suggests it solved the task. The issue is that they often have some structural defects that are only obvious later in the or when you go fairly far out into the curve. So for instance, in video understanding, you think, wow, if you looked at the benchmarks from a few years ago, they're totally at the top of them. But then you can come up a year later with several other sorts of benchmarks that challenge them. So I think that there tends to be a gravitation toward what's very tractable for AI systems and where some interesting action is happening. That's a selection pressure on the sort of benchmarks. If one's looking at cognitive tasks, such as those that you would give kids if you're testing their intelligence, for instance, if you randomly sample many of those, the the models don't do that well. Maybe on you could it'd be a double digit percentage, maybe be almost on half of them they don't do that well. For instance, count the number of faces in this photograph. O 3 can't do that very well. Or connect the dots or fill in fill in the colors in this in this picture. Just an example for for visual ability. So there's I I don't think when Gary's pointing these out, it's it's just that he's just running the cherry picking program, and what he's going do is he's just going to keep he's just going keep cherry picking and doing God of the gaps thing for AI until it basically is is AGI, I I I do think that there is a non adversarial distribution, a difference from the cognitive science angle, you would see many of these these, sorts of issues. And I I think that there's potentially some speaking past each other in part because he's not viewing intelligence as some unidimensional thing entirely. There's a lot might be correlated together, but there are various other mental faculties that are important really online that much. We saw with the GPT series that by pre training on a lot of text we got some subcomponents of intelligence. We got reading writing ability and we got a lot of crystallized intelligence or acquired knowledge from that. And that process took several years, so it's not the case, there's often people will sort of point out once it gets traction on it then it will solve it immediately or solve it very shortly thereafter. For a lot of these core cognitive abilities, they took multiple years. Mathematical reasoning would be a recent example with Minerva. Google's Minerva system got 50% on the math benchmark, some benchmark I made some while ago, in 2022. And I think we've only recently crushed it in 2025, so it took a good 3 years. And I think the reading writing ability took got to an interesting state and now relatively complete 4 years later, I think crystallized intelligence as well is now relatively complete.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.