Evidence receipt / belief
Published · transcript-backedSpeaker unverified: belief
27 Sept 2025 Machine Learning Street Talk New top score on ARC-AGI-2-pub (29.4%) - Jeremy Berman
“I think generally, when models think in natural language and they output a natural language, they, they are higher entropy.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- belief
- Recorded
- 27 Sept 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…I I did an interesting interview at NeurIPS last year, with the Google guys, and and they were talking about this adaptive temperature in language models for reasoning. Because, you know, there's this constant trade off between reasoning, we wanna be quite constrained. Right? So so we actually want to kind of, like, go go a particular pathway. We want to be constrained by our knowledge. And when we're being quite creative and flexible, we want to we want to be able to go in in different places. And I'm I'm really interested in creativity, for example, and and and I think creativity is like you you you it's very similar to reasoning as Charley talks about. You know, you're composing together these constraints. There's this phylogeny of of knowledge, and you need to respect it as much as possible because if you don't respect it, you're not grounded anymore. So it kind of feels to me that intuitively code is great because it means that I'm actually respecting the constraints and and the semantics are correct and it's grounded in in the real world. Do you feel in any way that by using these natural language descriptions that you're kind of creating something which might by dint of chance or search find the right solution, but is isn't correct and verifiable? Yes. Okay. Yes. Tell tell me more. Yes. I I for sure. I think generally, when models think in natural language and they output a natural language, they, they are higher entropy. Right? Yeah. I think the when you the second you start prompting with code, they go into code mode. And this is you know, there are lot of papers that show, right, just by prompting it in a certain direction, it activates certain weights that are, you know, just naturally lower entropy. But that was part of the thing that I wanted. I actually wanted to introduce entropy because, you know, still most arc tasks for v 2, the models don't get close. Right? You know, my solution was the top, and it's at 30%. So I actually wanted to inject as much entropy as possible, which is partially why my, prompts are so broad. You know, I could definitely improve my accuracy on a few tasks by making the prompts more specific, but I wanted to just constantly berate it. More entropy. More entropy. So I actually found that to be a a positive, not a negative. Interesting. Interesting. On the efficiency of the solution, so the o 3 model from OpenAI, that was about $200 per task. And that was I think it was did we ever find out? I think it was sampling. Right? So they just sampled it a bunch of times. They had a basic verifier. Is is that correct?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.