← All source episodes Machine Learning Street Talk / episode intelligence
The Benchmark With No Instructions — ARC-AGI-3 (winning team!)
1 Jul 2026 27 published claims 6 attributable people
Speakers in the public record
Claim mix
belief 10evaluation 7prediction 3uncertainty 3recommendation 2preference 1commitment 1
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
27 published records
“Then if you train it using reasoning, you can actually do way more because now it has more time to actually reason and figure out what you're actually asking and form new extractions for solving a specific problem case as opposed to just regurgitating what it's seen on the internet?”
- Publisher
- Machine Learning Street Talk
“So the main constraint was we only had 3 games, 3 public games and 3 private games to evaluate on. So we had to make sure or that to make sure that whatever the design is is is you don't assume too much about the the environments and the games because it on the private leaderboard or like the public leaderboard beforehand, you could actually see that I wasn't even close to the top because it was super easy to overfit.”
- Publisher
- Machine Learning Street Talk
“We tried other methods such as directly predicting like more of a transductive method where you just have the input frames as context into a long sequence model and you just predict the actions. But that doesn't also it doesn't generalize well and it doesn't really make intuitive sense a lot because if you play the game, for example, the 1st game here, the maze level, you would intuitively play 1 or 2 actions and you know think about your path and you'll go to the end position.”
- Publisher
- Machine Learning Street Talk
“they have the intention, they have something that the coding agents don't have because you know, the bull case, wouldn't it be amazing if you could stick the requirements in and the agents themselves would kind of understand, oh I see what they meant and they could evolve the requirements.”
- Publisher
- Machine Learning Street Talk
“You mentioned transductive as well which is quite interesting because roughly speaking I think of transduction as you're making a prediction about the specific test instance.”
- Publisher
- Machine Learning Street Talk
“Does that really count as the old models can play games now if it takes them so much effort? I would say that is a real gap.”
- Publisher
- Machine Learning Street Talk
“I just only I think I can base my answer on the on these examples and I completely agree with you that these that there are some abstractions like that.”
- Publisher
- Machine Learning Street Talk
“I don't know if it's intentional or just the speed at which they developed them, but it does seem pretty clear.”
- Publisher
- Machine Learning Street Talk
“For example, you have this reasoning tokens which has like abstract representation of objects, which we then manifest or write down in like an English language textual tokens. So I would say that's more abstract, you identify objects, you try and find out what the mechanics is, dynamics of the game, what the goal is.”
- Publisher
- Machine Learning Street Talk
“I don't know. I guess the main idea is a bunch of bright people in the room and do good research together.”
- Publisher
- Machine Learning Street Talk
“I think the 1st 4 places was basically brute force algorithms that just search over a large space of actions but do some basic form of filtering and you can get a very good score.”
- Publisher
- Machine Learning Street Talk
“I think ARC is an example for the case that you can do this, at least to some level, because we see the Frontier LLM scoring quite well on this end of day, wouldn't be able to if these core priors would be breaking it.”
- Publisher
- Machine Learning Street Talk
“I think I think that's the interesting part about ARC that it kind of tries to remove as much as possible the prior that you get from language or from human knowledge and kind of strip them to the minimum and really only test for intelligence.”
- Publisher
- Machine Learning Street Talk
“Like, for instance, if you give them an auto research task, they might overfocus on details and try kind of not see the big picture and not kind of zoom out and see things from Fireball. And they might, I don't know, start optimizing some hyperparameters for, at 0.”
- Publisher
- Machine Learning Street Talk
“I think we have a clear benchmark, which we know humans, which is general in some sub domain, which we can say whatever that domain is can score that score and that is the medium score.”
- Publisher
- Machine Learning Street Talk
“Yeah, I think the reason we're putting it back is because we're specifically focusing on language models which has been extensively trained on language.”
- Publisher
- Machine Learning Street Talk
“I don't really know what he thinks, but you know, I've I've got a pretty good simulation of Charle in my mind.”
- Publisher
- Machine Learning Street Talk
“because there's so little time to learn and you really, if you wanna get the 100% which is the which is the grand prize, then you really you you have to outcompete a human which really can take a lot of time just to think.”
- Publisher
- Machine Learning Street Talk
“Exactly what I did was I just used a action model that learns, yeah, which frames are valid for a given transition and then it was more about the engineering of being able to learn within 100,000 actions because that was kind of the max action limit we could get out in the time limit we had.”
- Publisher
- Machine Learning Street Talk
“Yeah. And not yet I will mention, even with our simple With GWEN, we can already get sometimes 100% on certain games.”
- Publisher
- Machine Learning Street Talk
“Like we have to make sure it's working, but it is a case that we're gradually we are understanding less and less of our own code base. And we're struggling even with reviewing some of the changes is you might use a codex to help review some of it because it's such a broad change or we need to split it up.”
- Publisher
- Machine Learning Street Talk
“The the best recipe we have today to build the intelligent systems is scaling up these large language models. I think Sam said you should definitely not be trying to to train LLMs yourself.”
- Publisher
- Machine Learning Street Talk
“Right? But that doesn't work because it doesn't really it cannot really it we have tried that and that has not really led to improvements because it doesn't really understand where the agent fails.”
- Publisher
- Machine Learning Street Talk
“The question is whether this holds also for the for the withheld private test, but it seems that goal setting is not the bottlenecks, rather the action efficiency and the accumulation of the knowledge over a very long context because you need millions hundreds of thousands, if not millions of tokens to solve this.”
- Publisher
- Machine Learning Street Talk
“You play a game, you solve it as in like a large compute or action budget and then you have to play it again and you have to do like a speed run through it to improve it. So I think that is important, but I also think action efficiency is a practical step to reduce just brute force solutions.”
- Publisher
- Machine Learning Street Talk
“my career has been going on a bit longer, and let's say most of it, for most of my career, there were no coding agents, which means I also have a bit of an opportunity to still make use of patterns that have been useful in the past.”
- Publisher
- Machine Learning Street Talk
“I think in this world, we can still use intelligence and we can still like acquire abstractions and descriptions for these high level phenomena.”
- Publisher
- Machine Learning Street Talk