Evidence receipt / prediction
Published · transcript-backedMicah-Hill Smith: prediction
8 Jan 2026 Latent Space Artificial Analysis: Independent LLM Evals as a Service — with George Cameron and Micah-Hill Smith
“Almost every model coming out because doing things like writing the Python functions and human evil is now pretty trivial.”
Source trail
Everything needed to verify it.
- Speaker
- Micah-Hill Smith
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 8 Jan 2026
- Publisher
- Latent Space
Transcript context
…And I think that reflects on the change in the industry. Right. So that’s a nice way to tell that story. Well, V1 would be completely saturated right now. Almost every model coming out because doing things like writing the Python functions and human evil is now pretty trivial. It’s easy to forget, actually, I think how much progress has been made in the last two years. Like we obviously play the game constantly of like the today’s version versus last week’s version and the week before and all of the small changes in the horse race between the current frontier and who has the best like smaller than 10B model like right now this week. Right. And that’s very important to a lot of developers and people and especially in this particular city of San Francisco. But when you zoom out a couple of years ago, literally most of what we were doing to evaluate the models then would all be 100% solved by even pretty small models today. And that’s been one of the key things, by the way, that’s driven down the cost of intelligence at every tier of intelligence. We can talk about more in a bit. So V1, V2, V3, we made things harder. We covered a wider range of use cases. And we tried to get closer to things developers care about as opposed to like just the Q&A type stuff that MMLU and GPQA represented. Yeah. I don’t know if you have anything to add there. Or we could just go right into showing people the benchmark and like looking around and asking questions about it. Yeah.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.