High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Speaker unverified: belief

27 Sept 2025 Machine Learning Street Talk New top score on ARC-AGI-2-pub (29.4%) - Jeremy Berman

“I think OpenAI, generally, they wanna do the right thing, and they want their solutions to be very general and broad.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
belief
Recorded
27 Sept 2025
Publisher
Machine Learning Street Talk

Transcript context

…Oh, very interesting. Yeah. I remember there was that big, hoo at the time that they you know, it was scandalous that they were training on the training set. But, anyway, that's 1 side. There's also, like, the thought occurs that if they did did if they did do something like what you are doing, so, you know, like this approach of iterative refinement with verification at every single step, would they have done even better? For sure. Yeah. For sure. I think OpenAI, generally, they wanna do the right thing, and they want their solutions to be very general and broad. This And is the sense I get. I think it's to their culture. I spoke to the OpenAI guys. They did include the training data in that o 3 model, but I think that's fair game. Right? They also so I don't think they fine tune on it. Right? So it's just part of the corpus that went into pretraining, which to me is fair game. Right? This is this is totally fine. Okay. Very cool. So so on ARC v 1, their their efficiency was $200 per task. What was your efficiency?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence