High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / recommendation

Published · transcript-backed

Likes Sonnet.

27 Sept 2025 Machine Learning Street Talk New top score on ARC-AGI-2-pub (29.4%) - Jeremy Berman

“The problem is, you know, for my v 1 solution, you also have to generate code, And Sonnet 3 5 is really good at thinking about code and generating code, and I I prefer Sonnet to Grok, for code generation.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
recommendation
Recorded
27 Sept 2025
Publisher
Machine Learning Street Talk

Transcript context

…The the other important thing is you are now using Grok 4, which is very, very powerful. I I assume you you chose Grok 4 because it just happened to be the the winner on on the leaderboard for for the base ARC v 2. And how much uplift is coming from that? So for example, if you used Grok 4 on your on your solution last year, how much better would it be? So I actually don't think it would be so much better. For some reason, and this is what I talk about in my blog post, these language models are very spiky in certain things where they were trained heavily on. Right? So I think what happened with Grok is there was a distribution of similar, shape tasks, right, grid tasks, just reasoning in the type of gen general direction that allowed Grok to have a special capability in this area. And I actually, like, tested each model. So I tested Grok versus GPT. You know? I didn't just go by the leaderboard, and Grok definitely outperformed. The problem is, you know, for my v 1 solution, you also have to generate code, And Sonnet 3 5 is really good at thinking about code and generating code, and I I prefer Sonnet to Grok, for code generation. So my guess actually would be that if you use my v 1 solution, it's highly possible, you know, Opus 4 1 would be the best. I haven't tested that. Would be very expensive on Opus 4.1, but maybe it's worth testing. I I think that the general idea is that, these networks are are very spiky. When you get into specific domains, the net it it actually very much matters which model you use. And the ARC is a great example of this. Right? Like, leaderboard is super spiky, in ways that, other benchmarks are not. I I did an interesting interview at NeurIPS last year, with the Google guys, and and they were talking about this adaptive temperature in language models for reasoning. Because, you know, there's this constant trade off between reasoning, we wanna be quite constrained. Right? So so we actually want to kind of, like, go go a particular pathway. We want to be constrained by our knowledge. And when we're being quite creative and flexible, we want to we want to be able to go in in different places. And I'm I'm really interested in creativity, for example, and and and I think creativity is like you you you it's very similar to reasoning as Charley talks about. You know, you're composing together these constraints. There's this phylogeny of of knowledge, and you need to respect it as much as possible because if you don't respect it, you're not grounded anymore. So it kind of feels to me that intuitively code is great because it means that I'm actually respecting the constraints and and the semantics are correct and it's grounded in in the real world. Do you feel in any way that by using these natural language descriptions that you're kind of creating something which might by dint of chance or search find the right solution, but is isn't correct and verifiable?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Named in this claim

Books, apps, tools, and people.

Search evidence