Evidence receipt / belief
Published · transcript-backedMichal Tesnar: belief
1 Jul 2026 Machine Learning Street Talk The Benchmark With No Instructions — ARC-AGI-3 (winning team!)
“I think ARC is an example for the case that you can do this, at least to some level, because we see the Frontier LLM scoring quite well on this end of day, wouldn't be able to if these core priors would be breaking it.”
Source trail
Everything needed to verify it.
- Speaker
- Michal Tesnar
- Attribution
- Verified speaker
- Claim type
- belief
- Recorded
- 1 Jul 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…So like with the glider for example, we almost don't need to know how it came about because the new description encompasses, it's a complete description of the new phenomena. It's almost like it's the start of a new piece of knowledge. And a lot of our knowledge is like that. It's such a good compression of what went underneath that we can start from it. Like for example like So calculus for example. So calculus is within the closure of calculus or probability theory. We can do a lot of things. We don't need to know how it came about. So so but but then is is that a new layer of knowledge and we don't care how it came about? It's just kind of like saying do we need to start from the bottom? Or in many cases, you know, this this has come about and we can just use it you know, higher up the abstraction mountain. Yeah. I think ARC is an example for the case that you can do this, at least to some level, because we see the Frontier LLM scoring quite well on this end of day, wouldn't be able to if these core priors would be breaking it. But I think what is really nice about Arc is that there are some games that are so different and are still easy for humans, but they somehow like break this concept. Right? You talk about this closure of like, and I gave the example of Maze, But like you can maybe change the core prior concepts, and they would still make sense for a human, but they would somehow break adversarially this this like, let's say, calculus like, maze like representation that helps solve the game. You can like move them around a little bit, then you suddenly break it, and you still have a valid game for a human, but not for the abstracted LLM intelligence. And that's what we see also with the with some of the difficult games. But I think it's really difficult to not move the priors too much so that you can still make it human solvable. And that sometimes ARK unfortunately fails at this. I want to talk a little bit about RKGI 3. So the 1st 2 versions of the ARK challenge, they were quite abstract because Charle was trying to idealize intelligence on its own in the most abstract kind of legible way. RKGI 3, it introduces in my opinion, and we can talk about this, the concept of agency. So agency in my in my definition is the ability of an agent to have goals, to plan and realize those goals. And the more ambitious the goals are and the more you realize the goals, the more agency you have. So there's a kind of low level no nonsense definition of agency which is just a thing that can sense and act. So that's like if I'm doing computer programming that's what an agent is but the more kind of cognitive science definition of an agent is 1 that sort of has future point in control. I have these big goals in the future and I can realize them and that makes me an agent. So I you know I think that RKGI 3 introduces agency not just in terms of realizing goals but acquiring goals.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.