Evidence receipt / evaluation
Published · transcript-backedTim Scarfe: evaluation
27 Sept 2025 Machine Learning Street Talk New top score on ARC-AGI-2-pub (29.4%) - Jeremy Berman
“The reason for that, as we discuss in today's show, is that current AI does not understand the world in a grounded way. It doesn't have a deep abstract understanding of the world, which is why the only way that we can make AI work effectively is by grounding the generation and supervising the training of AI models with human data.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 27 Sept 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…I get actually even more fundamentally like the ideal system would be we have a set of data. Our language model is bad at a certain thing. We can just give it this data and then all of a sudden it keeps all of its knowledge and then also gets really good at this new thing. We we are not there yet and that to me is like a fundamental, missing part. Really what you want is more expressive program. And so that's why I switched from Python to English, which is a much more expressive program. You can language you can always teach a language model skill, Right? But it's the meta skill. It's the skill to create the skills that is AGI. And to me, that's reasoning. Like, reasoning is that meta skill. And so, to put it another way, I think if you fundamentally learn the skill of reasoning, you should be able to then, apply that skill to learn all the other skills. That is the meta skill. You know, kick whatever weights out you need to, align the model to reason, and then from there, you have a foundation from which you can actually build general intelligence. Okay, folks. Hot off the press. Many of you would have seen last week that Jeremy Berman, who is a research scientist at Reflection AI, is now the winner of the ArcGI v 2 leaderboard, the public version of the leaderboard. He's using an evolutionary approach. Now, remember last year in December, he published a similar approach. Generating Python functions and then refining those functions in a kind of iterative loop. His new architecture is generating descriptions of algorithms rather than code. And iteratively, in an evolutionary sense, refining those ones and discarding the ones that don't work. He's now at the top of the leaderboard. It's a really really cool and elegant algorithm. And by the way, he works for reflection AI. So he's doing reinforcement learning with verifiable feedback. And he's trying to address the biggest gap in AI at the moment, which is that we want systems that can synthesize new knowledge and new understanding. Current systems just get trained with a whole bunch of data and they only know what they've been trained on. They can't kind of think outside the box by creatively synthesizing new knowledge. Prolific are really focused on the contributions of human data in AI. And the reason this is important is actually the dirty secret of Silicon Valley. The extent to which human data is used to evaluate and fine tune AI models. The reason for that, as we discuss in today's show, is that current AI does not understand the world in a grounded way. It doesn't have a deep abstract understanding of the world, which is why the only way that we can make AI work effectively is by grounding the generation and supervising the training of AI models with human data. Prolific are putting together a report on how human data is being used in AI systems and they need volunteers. You can just go and fill out this form to help them produce this report and you will get privileged access to see the report before anyone else. The link is in the description. Oh, there was an amazing part in, I think it was in your first paper where you said, a parrot that lives in a courthouse will regurgitate more correct statements than a parrot that lives in a madhouse. Thank you. Thank you. My sister, who doesn't know anything about language models or AI, she pointed that out and said that was a great line. So at least I have that.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.