High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / preference

Published · transcript-backed

Tim Scarfe: preference

27 Sept 2025 Machine Learning Street Talk New top score on ARC-AGI-2-pub (29.4%) - Jeremy Berman

“There's this phylogeny of of knowledge, and you need to respect it as much as possible because if you don't respect it, you're not grounded anymore. So it kind of feels to me that intuitively code is great because it means that I'm actually respecting the constraints and and the semantics are correct and it's grounded in in the real world.”

— Tim Scarfe

Source trail

Everything needed to verify it.

Speaker
Tim Scarfe
Attribution
Verified speaker
Claim type
preference
Recorded
27 Sept 2025
Publisher
Machine Learning Street Talk

Transcript context

…So I actually don't think it would be so much better. For some reason, and this is what I talk about in my blog post, these language models are very spiky in certain things where they were trained heavily on. Right? So I think what happened with Grok is there was a distribution of similar, shape tasks, right, grid tasks, just reasoning in the type of gen general direction that allowed Grok to have a special capability in this area. And I actually, like, tested each model. So I tested Grok versus GPT. You know? I didn't just go by the leaderboard, and Grok definitely outperformed. The problem is, you know, for my v 1 solution, you also have to generate code, And Sonnet 3 5 is really good at thinking about code and generating code, and I I prefer Sonnet to Grok, for code generation. So my guess actually would be that if you use my v 1 solution, it's highly possible, you know, Opus 4 1 would be the best. I haven't tested that. Would be very expensive on Opus 4.1, but maybe it's worth testing. I I think that the general idea is that, these networks are are very spiky. When you get into specific domains, the net it it actually very much matters which model you use. And the ARC is a great example of this. Right? Like, leaderboard is super spiky, in ways that, other benchmarks are not. I I did an interesting interview at NeurIPS last year, with the Google guys, and and they were talking about this adaptive temperature in language models for reasoning. Because, you know, there's this constant trade off between reasoning, we wanna be quite constrained. Right? So so we actually want to kind of, like, go go a particular pathway. We want to be constrained by our knowledge. And when we're being quite creative and flexible, we want to we want to be able to go in in different places. And I'm I'm really interested in creativity, for example, and and and I think creativity is like you you you it's very similar to reasoning as Charley talks about. You know, you're composing together these constraints. There's this phylogeny of of knowledge, and you need to respect it as much as possible because if you don't respect it, you're not grounded anymore. So it kind of feels to me that intuitively code is great because it means that I'm actually respecting the constraints and and the semantics are correct and it's grounded in in the real world. Do you feel in any way that by using these natural language descriptions that you're kind of creating something which might by dint of chance or search find the right solution, but is isn't correct and verifiable? Yes. Okay. Yes. Tell tell me more. Yes. I I for sure. I think generally, when models think in natural language and they output a natural language, they, they are higher entropy. Right? Yeah. I think the when you the second you start prompting with code, they go into code mode. And this is you know, there are lot of papers that show, right, just by prompting it in a certain direction, it activates certain weights that are, you know, just naturally lower entropy. But that was part of the thing that I wanted. I actually wanted to introduce entropy because, you know, still most arc tasks for v 2, the models don't get close. Right? You know, my solution was the top, and it's at 30%. So I actually wanted to inject as much entropy as possible, which is partially why my, prompts are so broad. You know, I could definitely improve my accuracy on a few tasks by making the prompts more specific, but I wanted to just constantly berate it. More entropy. More entropy. So I actually found that to be a a positive, not a negative.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence