High Signal Podcasts Evidence ledger
Method
Browse
← All source episodes

Machine Learning Street Talk / episode intelligence

The Benchmark With No Instructions — ARC-AGI-3 (winning team!)

1 Jul 2026 27 published claims 6 attributable people

Speakers in the public record

Claim mix

belief 10evaluation 7prediction 3uncertainty 3recommendation 2preference 1commitment 1

Evidence policy

Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.

Claim ledger

The useful parts, with receipts.

27 published records

01 / evaluation

Then if you train it using reasoning, you can actually do way more because now it has more time to actually reason and figure out what you're actually asking and form new extractions for solving a specific problem case as opposed to just regurgitating what it's seen on the internet?

“Then if you train it using reasoning, you can actually do way more because now it has more time to actually reason and figure out what you're actually asking and form new extractions for solving a specific problem case as opposed to just regurgitating what it's seen on the internet?”
Speaker
Dries Smit
Publisher
Machine Learning Street Talk

02 / evaluation

So the main constraint was we only had 3 games, 3 public games and 3 private games to evaluate on. So we had to make sure or that to make sure that whatever the design is is is you don't assume too much about the the environments and the games because it on the private leaderboard or like the public leaderboard beforehand, you could actually see that I wasn't even close to the top because it was super easy to overfit.

“So the main constraint was we only had 3 games, 3 public games and 3 private games to evaluate on. So we had to make sure or that to make sure that whatever the design is is is you don't assume too much about the the environments and the games because it on the private leaderboard or like the public leaderboard beforehand, you could actually see that I wasn't even close to the top because it was super easy to overfit.”
Speaker
Dries Smit
Publisher
Machine Learning Street Talk

03 / evaluation

We tried other methods such as directly predicting like more of a transductive method where you just have the input frames as context into a long sequence model and you just predict the actions. But that doesn't also it doesn't generalize well and it doesn't really make intuitive sense a lot because if you play the game, for example, the 1st game here, the maze level, you would intuitively play 1 or 2 actions and you know think about your path and you'll go to the end position.

“We tried other methods such as directly predicting like more of a transductive method where you just have the input frames as context into a long sequence model and you just predict the actions. But that doesn't also it doesn't generalize well and it doesn't really make intuitive sense a lot because if you play the game, for example, the 1st game here, the maze level, you would intuitively play 1 or 2 actions and you know think about your path and you'll go to the end position.”
Speaker
Dries Smit
Publisher
Machine Learning Street Talk

04 / prediction

they have the intention, they have something that the coding agents don't have because you know, the bull case, wouldn't it be amazing if you could stick the requirements in and the agents themselves would kind of understand, oh I see what they meant and they could evolve the requirements.

“they have the intention, they have something that the coding agents don't have because you know, the bull case, wouldn't it be amazing if you could stick the requirements in and the agents themselves would kind of understand, oh I see what they meant and they could evolve the requirements.”
Speaker
Dries Smit
Publisher
Machine Learning Street Talk

05 / belief

You mentioned transductive as well which is quite interesting because roughly speaking I think of transduction as you're making a prediction about the specific test instance.

“You mentioned transductive as well which is quite interesting because roughly speaking I think of transduction as you're making a prediction about the specific test instance.”
Speaker
Tim Scarfe
Publisher
Machine Learning Street Talk

09 / belief

For example, you have this reasoning tokens which has like abstract representation of objects, which we then manifest or write down in like an English language textual tokens. So I would say that's more abstract, you identify objects, you try and find out what the mechanics is, dynamics of the game, what the goal is.

“For example, you have this reasoning tokens which has like abstract representation of objects, which we then manifest or write down in like an English language textual tokens. So I would say that's more abstract, you identify objects, you try and find out what the mechanics is, dynamics of the game, what the goal is.”
Speaker
Dries Smit
Publisher
Machine Learning Street Talk

11 / belief

I think the 1st 4 places was basically brute force algorithms that just search over a large space of actions but do some basic form of filtering and you can get a very good score.

“I think the 1st 4 places was basically brute force algorithms that just search over a large space of actions but do some basic form of filtering and you can get a very good score.”
Speaker
Dries Smit
Publisher
Machine Learning Street Talk

12 / belief

I think ARC is an example for the case that you can do this, at least to some level, because we see the Frontier LLM scoring quite well on this end of day, wouldn't be able to if these core priors would be breaking it.

“I think ARC is an example for the case that you can do this, at least to some level, because we see the Frontier LLM scoring quite well on this end of day, wouldn't be able to if these core priors would be breaking it.”
Speaker
Michal Tesnar
Publisher
Machine Learning Street Talk

13 / belief

I think I think that's the interesting part about ARC that it kind of tries to remove as much as possible the prior that you get from language or from human knowledge and kind of strip them to the minimum and really only test for intelligence.

“I think I think that's the interesting part about ARC that it kind of tries to remove as much as possible the prior that you get from language or from human knowledge and kind of strip them to the minimum and really only test for intelligence.”
Speaker
Stefano Viel
Publisher
Machine Learning Street Talk

14 / uncertainty

Like, for instance, if you give them an auto research task, they might overfocus on details and try kind of not see the big picture and not kind of zoom out and see things from Fireball. And they might, I don't know, start optimizing some hyperparameters for, at 0.

“Like, for instance, if you give them an auto research task, they might overfocus on details and try kind of not see the big picture and not kind of zoom out and see things from Fireball. And they might, I don't know, start optimizing some hyperparameters for, at 0.”
Speaker
Stefano Viel
Publisher
Machine Learning Street Talk

15 / belief

I think we have a clear benchmark, which we know humans, which is general in some sub domain, which we can say whatever that domain is can score that score and that is the medium score.

“I think we have a clear benchmark, which we know humans, which is general in some sub domain, which we can say whatever that domain is can score that score and that is the medium score.”
Speaker
Dries Smit
Publisher
Machine Learning Street Talk

18 / preference

because there's so little time to learn and you really, if you wanna get the 100% which is the which is the grand prize, then you really you you have to outcompete a human which really can take a lot of time just to think.

“because there's so little time to learn and you really, if you wanna get the 100% which is the which is the grand prize, then you really you you have to outcompete a human which really can take a lot of time just to think.”
Speaker
Michal Tesnar
Publisher
Machine Learning Street Talk

19 / commitment

Exactly what I did was I just used a action model that learns, yeah, which frames are valid for a given transition and then it was more about the engineering of being able to learn within 100,000 actions because that was kind of the max action limit we could get out in the time limit we had.

“Exactly what I did was I just used a action model that learns, yeah, which frames are valid for a given transition and then it was more about the engineering of being able to learn within 100,000 actions because that was kind of the max action limit we could get out in the time limit we had.”
Speaker
Dries Smit
Publisher
Machine Learning Street Talk

21 / evaluation

Like we have to make sure it's working, but it is a case that we're gradually we are understanding less and less of our own code base. And we're struggling even with reviewing some of the changes is you might use a codex to help review some of it because it's such a broad change or we need to split it up.

“Like we have to make sure it's working, but it is a case that we're gradually we are understanding less and less of our own code base. And we're struggling even with reviewing some of the changes is you might use a codex to help review some of it because it's such a broad change or we need to split it up.”
Speaker
Dries Smit
Publisher
Machine Learning Street Talk

22 / recommendation

The the best recipe we have today to build the intelligent systems is scaling up these large language models. I think Sam said you should definitely not be trying to to train LLMs yourself.

“The the best recipe we have today to build the intelligent systems is scaling up these large language models. I think Sam said you should definitely not be trying to to train LLMs yourself.”
Publisher
Machine Learning Street Talk

23 / evaluation

Right? But that doesn't work because it doesn't really it cannot really it we have tried that and that has not really led to improvements because it doesn't really understand where the agent fails.

“Right? But that doesn't work because it doesn't really it cannot really it we have tried that and that has not really led to improvements because it doesn't really understand where the agent fails.”
Speaker
Michal Tesnar
Publisher
Machine Learning Street Talk

24 / evaluation

The question is whether this holds also for the for the withheld private test, but it seems that goal setting is not the bottlenecks, rather the action efficiency and the accumulation of the knowledge over a very long context because you need millions hundreds of thousands, if not millions of tokens to solve this.

“The question is whether this holds also for the for the withheld private test, but it seems that goal setting is not the bottlenecks, rather the action efficiency and the accumulation of the knowledge over a very long context because you need millions hundreds of thousands, if not millions of tokens to solve this.”
Speaker
Michal Tesnar
Publisher
Machine Learning Street Talk

25 / recommendation

You play a game, you solve it as in like a large compute or action budget and then you have to play it again and you have to do like a speed run through it to improve it. So I think that is important, but I also think action efficiency is a practical step to reduce just brute force solutions.

“You play a game, you solve it as in like a large compute or action budget and then you have to play it again and you have to do like a speed run through it to improve it. So I think that is important, but I also think action efficiency is a practical step to reduce just brute force solutions.”
Speaker
Dries Smit
Publisher
Machine Learning Street Talk

26 / evaluation

my career has been going on a bit longer, and let's say most of it, for most of my career, there were no coding agents, which means I also have a bit of an opportunity to still make use of patterns that have been useful in the past.

“my career has been going on a bit longer, and let's say most of it, for most of my career, there were no coding agents, which means I also have a bit of an opportunity to still make use of patterns that have been useful in the past.”
Speaker
Jeroen Cottaar
Publisher
Machine Learning Street Talk
Search evidence