Evidence receipt / evaluation
Published · transcript-backedTim Scarfe: evaluation
25 Jan 2026 Machine Learning Street Talk VAEs Are Energy-Based Models? [Dr. Jeff Beck]
“The ARC was actually really amazing because it's the only intelligence benchmark that has survived for 5 years before being defeated.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 25 Jan 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…Yes. And that is a laudable goal, right? And I certainly share it, right? The last thing you wanna do is I mean, you know, fortunately, like, networks are so big, we don't really run the risk of of, like, overfitting so as much as we used to. But the last thing you wanna do is throw is is train your network to toss information that you might need down the road. That said, the vast majority of what, you know, the brain does just like these neural networks is decide what information is currently task irrelevant. But that's all the more reason to do things in a self supervised or unsupervised way. Right? Because you're basically not telling it this is the important. You know, you're not telling it like what's all task relevant and task irrelevant. So I interviewed about the version 2 of the ARC challenge. And 1 thing that struck me is I think of intelligence as being multidimensional. So version 1 got saturated. The ARC was actually really amazing because it's the only intelligence benchmark that has survived for 5 years before being defeated. Since the advent of these thinking models, it has been defeated very quickly. But they're working on version 3, and there'll be version 4, there'll be version 5. Will there always just be something left over? That sounds like another philosophical. So yes is my answer. There will always be there will always be something left over. In the sense that like, you know, you know, we we we have this this has been the trajectory things have been going for a really long time. Right? It's sort of like, we get algorithms that do amazing new cool things, and then someone comes along and says, yeah, but it can't build me. It it can't pull a rabbit out of a hat. Right? And then and then of course, what does someone do? They oh, they they figure out the new training protocol, a slightly different architecture, or they just train it to pull rabbits out of hats, and then suddenly it can. And then someone proposes a new challenge, and a new challenge, and a new challenge. And it's always this game of like 1 upmanship. So the question becomes, well, what's the point at which there are no more new challenges? I'm not entirely certain we're ever gonna get there. Right? It may very well be the case that we get, you know, these sort of algorithms that are capable of replicating the complete suite of human behaviors, and then someone will come up with some criticism like, yeah, but it's not really doing x, it's just faking it. Right? This is just the direction things go because people really do think they're important.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.