High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Tim Scarfe: belief

1 Jul 2026 Machine Learning Street Talk The Benchmark With No Instructions — ARC-AGI-3 (winning team!)

“Yeah, I think the reason we're putting it back is because we're specifically focusing on language models which has been extensively trained on language.”

— Tim Scarfe

Source trail

Everything needed to verify it.

Speaker
Tim Scarfe
Attribution
Verified speaker
Claim type
belief
Recorded
1 Jul 2026
Publisher
Machine Learning Street Talk

Transcript context

…it's it's still hard to make a harness that solve these games and it's still and only starting, like and training is also hard. Like, it's not as simple as just throwing computer. There are still many details that you have to get right. But for sure, as you said, I think it brought up a lot level needed to enter the competition. Like, we're seeing that many people are stuck in the kind of template solution and there are as far as we are aware, at least in the competition, not many trying kind of LM approaches because it's so computationally expensive. So it's definitely made it harder for kind of the average person to enter the competition, but I wouldn't say that it's just a matter of throwing a lot of compute at it and brute forcing it. Yeah. It it's not like most other benchmarks. Right? It it there's there's no language, there's no instructions. Does that make it quite distinct as well? I think I think that's the interesting part about ARC that it kind of tries to remove as much as possible the prior that you get from language or from human knowledge and kind of strip them to the minimum and really only test for intelligence. I think, yeah, that's a cool feature. Language is the basis of how we do a lot of thinking. And what you folks have done is you've gone to language and then gone back again. So we've got this kind of loop. So we go to language, we do some reasoning and then we might use that to do active fine tuning or like do reinforcement learning or whatever. And we've got this virtuous cycle. So he took the language away and we're kind of putting it back again. Yeah, I think the reason we're putting it back is because we're specifically focusing on language models which has been extensively trained on language. I guess if there's other approaches like neural guided search, they might not use language at all. You don't need language for this, at least a human textual language, but it's just in our case, we're heavily leveraging reasoning models. So it's just natural to bring them back into the domain they're experts in and what they're doing while we wanna actually just our harness should shape that and encourage that behavior. So, I guess that's the main reason we're moving back. I think it's also very difficult…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence