Evidence receipt / uncertainty
Published · transcript-backedGary Marcus: uncertainty
24 Jun 2025 Machine Learning Street Talk Three Red Lines We're About to Cross Toward AGI (Daniel Kokotajlo, Gary Marcus, Dan Hendrycks)
“The coding tasks, I'm very concerned from a scientific perspective that we don't know how much data contamination there is.”
Source trail
Everything needed to verify it.
- Speaker
- Gary Marcus
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 24 Jun 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…So the way I would interpret this graph and I'm sure many of our audience has already seen this. But they have this suite of agentic coding tasks that are organized from shortest to longest in terms of, like, how long it takes as human beings to complete the tasks. And then they note that, you know, the AIs of 2023 could, generally speaking, do the tasks from here to here but not do the tasks above this length. But then each year, the crossover point has been lengthening. And a very natural interpretation of what's going on here is that the AIs are getting more reliable. If they have a chance of getting into some sort of catastrophic error at any given point, then if it's like a 1% chance per second, then after 50 seconds, they're going to get into an error. But then if that goes down by an order of magnitude, then they can go for more seconds and so forth. And so so the thought here, would say this is evidence that they are just getting, generally speaking, more reliable. Better at not only not making mistakes, but recovering from the mistakes they make. Not infinitely better. They're still less reliable than humans. But there's substantial progress being made year over year. I see that your argument you're making with the particular graph, I think there's a lot of problems. And some are actually relevant to this argument. 1 is it was all relative to coding tasks. It wasn't tasks in general. The coding tasks, I'm very concerned from a scientific perspective that we don't know how much data contamination there is. We don't know how much augmentation is done relative to those benchmarks and so forth. Which leads me to my next point about it, which is I can't I'm blanking on the guy's name from 80000 hours posted something on Twitter yesterday which was a parallel to that graph. I think it was also by Meter, looking at agents. And so the y axis for the coding thing was sort of minutes to hours to days or whatever. And for agents, was like seconds to tens of seconds or something like that. So it was a totally different y axis. What types of agents?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.