High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / commitment

Published · transcript-backed

Micah-Hill Smith: commitment

8 Jan 2026 Latent Space Artificial Analysis: Independent LLM Evals as a Service — with George Cameron and Micah-Hill Smith

“We will make sure that within each version number that there’s no parliamentary issue.”

— Micah-Hill Smith

Source trail

Everything needed to verify it.

Speaker
Micah-Hill Smith
Attribution
Verified speaker
Claim type
commitment
Recorded
8 Jan 2026
Publisher
Latent Space

Transcript context

…Maybe that is a direction that Chadji opening eyes leaning into a lot. So if you manage to solve that, you should definitely talk to Fiji and Roon. Oh, okay. Yeah. So what is going to be included in, let’s say like a V3 of the intelligence index, because obviously you’re going to saturate in March. Why don’t we break it now? How soon is the podcast going to come out? Whenever you want. Okay. So we’re at V3 right now. So the, so the, the, the version that we, that’s going inside is, is, is, is V3 V4 is what we’re going to call the next, you know, major of it. Surprise, surprise. We’re going to be adding several of the things that we’ve actually talked about today that we’ve launched over the last few weeks. So it’s not, that’s not going to be wildly shocking, but some of the things that are most exciting is that adding GDP value is going to give us this general agentic performance in a really strong way in intelligence index and in critical point, the, um, physics, EBL, George was talking about similar to frontier math. Yeah. That’s very interesting. That gives us completely new view with a brand new data set of very, very hard research problems. We are going to be using Omniscience and we are going to be using hallucination rate. The exact way is that all of those are going to come together. Um, The waitings is going to be hard because the numbers are different. Yeah. We’re going to make sure that we don’t do anything to cause odd distortions and stuff that could be misleading. But every time you version it, you have a one-time reset of the Exactly. Yeah. That’s exactly how we think about it. We will make sure that within each version number that there’s no parliamentary issue. No drift in any of the scores so that people can rely on them and reference them. You just have to watch out for that version number. Once it’s v4.1, those numbers won’t be compatible with v4. Of course. There’s a little bit of debate over the accuracy of TileBench. I don’t know if you’re clued in to what’s going on. Apparently, a very high number of TileBench tests are impossible.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence