Evidence receipt / uncertainty
Published · transcript-backedTim Scarfe: uncertainty
14 Aug 2025 Machine Learning Street Talk Superintelligence Strategy (Dan Hendrycks)
“I mean, 1 thing I was thinking about is so certainly in in enigma about them, we're looking at multistep creative reasoning. And I don't know whether, again, there was some kind of human based methodology for filtering and coming up with ideas, or maybe your frame was, I have some technical principled intuition about what the limitations of AI models are, so I'm gonna lean in in that direction.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- uncertainty
- Recorded
- 14 Aug 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Yes. So Enigma Evaluation then is a collection of puzzles that, humanities so humanities last exam is getting individual questions that are tough that, an individual with a lot of expertise could solve. Enigma evaluation, you can think of it as like MIT mystery hunts. So MIT mystery hunts are things that happen like over a weekend and a group of MIT students try to solve this puzzle. There are many steps to it. So in terms of human compute, so to speak, it takes a lot of human compute to solve it. And it takes groups to solve it as well. And there's not a very high solve rate. So this is very multi step and requires sort of group level intelligence to be able to have a shot at solving it. So we just collected some of those and I think this approximates longer horizon types of intellectual tasks. And, yeah, I don't see that. I don't think that will be solved this year at all. I I would be I would be very surprised if it if it would be. So I think we have some evaluations that can keep us aware and able to differentiate between models for a while. There are other ones. For instance, we'll soon have out a automation related benchmarks. That way, we're just directly measuring what the automation rate of things are. But I won't go into, you know, too much detail about that until it's released. But there's so I I think that there's many axes for which the the the models are actually not doing that well, even though people claim that, you know, that all the benchmarks are saturated or they get solved, you know, in a in a few months, which I I I think you can create ones that that take take on the order of a year or 2 to to solve. Yeah. I mean, 1 thing I was thinking about is so certainly in in enigma about them, we're looking at multistep creative reasoning. And I don't know whether, again, there was some kind of human based methodology for filtering and coming up with ideas, or maybe your frame was, I have some technical principled intuition about what the limitations of AI models are, so I'm gonna lean in in that direction. But more broadly, I'm I'm interested in in intelligence and what it is. Right? So for me, it's about doing more with less. Right? You know, intelligence is about taking hard problems and making them simple. And and stupidity, ironically, is the other way around. It's about taking simple problems and and making them hard. And, LLMs, I'm kind of quoting, David Krakauer here, who's the director of the Santa Fe Institute. You know, he said that LLMs are doing more with more. So, they already know everything. Right? And and they and they can take these shortcuts, and and then that's why he thinks that they're not really Mhmm. Intelligent. And, you know, so we're we're left with this quandary, really, because arguably, these entities can take shortcuts, and they're not really doing things the way that that we are. And when we make more complex benchmarks, do you think that that increases the fog of war of how we evaluate these things? I think that they can definitely, prioritize some axes that aren't the key bottleneck capabilities. So for instance, Ximeni's last exam, gets at mathematical ability, But that's quite separate from various other abilities that it has. I mean, it, of course, is a combination of many different skills. But I think it very much gets at quantitative and mathematical ability. That is not necessarily a bottleneck for agency at all. So in thinking about intelligence, I tend to think about it on like 10 or so dimensions instead of monolithic definition or 1 key metric. And I think some of these benchmarks just get at different parts of that. So those dimensions would be things like fluid intelligence, what the the, ARC, stuff does and what Rayman's progressive matrices does. There's crystallized intelligence or acquired knowledge, which is what MMLU largely gets at. Does it know a lot about different things? Image classification, is it able to name lots of different species and objects, is also a a facet of crystallized intelligence. There's reading writing ability is its own, and the scaling substantially helped with that. There's its visual processing ability. So how well can it count things and objects Can it or in images? Can it discern the the latent pattern in an image, for instance? Is it able to generate images with sort of precise specifications? Like, can it, like, cross out the middle of some different segments? Can it determine the angle or line up the determine the angle of different things in an image? Like, is this a obtuse angle or an acute angle? There's audio audio processing ability. There's short term memory. There's long term memory. There's input processing speed. There's output processing speed. All these different things. If if you lack any of these, if you can't read and write for instance, it'd be severely limited. If you don't have a long term memory, you'll be severely limited. You'll be very difficult to employ. So I think there are several bottlenecks that, these benchmarks don't particularly get at. And so when it gets to 100%, people are, oh, why well, we still don't have something extremely economically valuable. And that that's a consequence of it just measuring a different facet. So yeah. But hopefully, that adds some type of clarity to it. And I think you just have to get all those axes to get something that is human level or at the level of a typical human on on cognitive tasks, that might be it could be thought AGI.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.