High Signal Podcasts Evidence ledger
Method
Browse
← All source episodes

Latent Space / episode intelligence

Artificial Analysis: Independent LLM Evals as a Service — with George Cameron and Micah-Hill Smith

8 Jan 2026 41 published claims 3 attributable people

Speakers in the public record

Claim mix

belief 18uncertainty 9commitment 3evaluation 3observation 3preference 3recommendation 1prediction 1

Evidence policy

Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.

Claim ledger

The useful parts, with receipts.

41 published records

01 / recommendation

It depends what you’re looking at, right? Because you can, if you’re trying to see whether or not it can solve a particular type of reasoning problem, and you don’t want to test it on its ability to do answer formatting at the same time, then you might want to use an LLM as answer extractor approach to make sure that you get the answer out no matter how unanswered.

“It depends what you’re looking at, right? Because you can, if you’re trying to see whether or not it can solve a particular type of reasoning problem, and you don’t want to test it on its ability to do answer formatting at the same time, then you might want to use an LLM as answer extractor approach to make sure that you get the answer out no matter how unanswered.”
Publisher
Latent Space

02 / uncertainty

I don’t know. I’m not used to it. Once upon a time, we did call it Quality Index, and we would talk about quality, performance, and price, but we changed it to intelligence.

“I don’t know. I’m not used to it. Once upon a time, we did call it Quality Index, and we would talk about quality, performance, and price, but we changed it to intelligence.”
Publisher
Latent Space

06 / belief

Over the last couple of years, the best way to think about that is that the cost for each terror of intelligence has been dropping the, like one fact on that is that you can get intelligence at the level of GPT-4 for over a hundred times cheaper than GPT-4 was at launch right now. I think my number is a thousand actually.

“Over the last couple of years, the best way to think about that is that the cost for each terror of intelligence has been dropping the, like one fact on that is that you can get intelligence at the level of GPT-4 for over a hundred times cheaper than GPT-4 was at launch right now. I think my number is a thousand actually.”
Publisher
Latent Space

08 / belief

The tools that are available have actually diverged in my opinion, a fair bit across the major chatbot apps and the amount of data sources that you can connect them to have gone up a lot, meaning that your experience and the way you’re using the model is more different than ever.

“The tools that are available have actually diverged in my opinion, a fair bit across the major chatbot apps and the amount of data sources that you can connect them to have gone up a lot, meaning that your experience and the way you’re using the model is more different than ever.”
Publisher
Latent Space

12 / belief

You can, uh, you can partially blame us and how we define intelligence having until now not defined hallucination as a negative in the way that we think about intelligence.

“You can, uh, you can partially blame us and how we define intelligence having until now not defined hallucination as a negative in the way that we think about intelligence.”
Publisher
Latent Space

13 / commitment

One of the reasons this is cool, right, is that if you’re trying to understand the holistic picture of the models and what you can do with all the stuff the company’s contributing, this gives you that picture. And so we are going to keep it up to date alongside all the models that we do intelligence index on, on the site.

“One of the reasons this is cool, right, is that if you’re trying to understand the holistic picture of the models and what you can do with all the stuff the company’s contributing, this gives you that picture. And so we are going to keep it up to date alongside all the models that we do intelligence index on, on the site.”
Publisher
Latent Space

17 / belief

I think our view is that hallucination rate makes sense in this context where it’s around knowledge, but in many cases, people want the models to hallucinate, to have a go.

“I think our view is that hallucination rate makes sense in this context where it’s around knowledge, but in many cases, people want the models to hallucinate, to have a go.”
Speaker
George Cameron
Publisher
Latent Space

22 / belief

I did give you shit for missing fireworks, and how do you have a model benchmarking thing without fireworks? But you had together, you had perplexity, and I think we just started chatting there.

“I did give you shit for missing fireworks, and how do you have a model benchmarking thing without fireworks? But you had together, you had perplexity, and I think we just started chatting there.”
Speaker
Shawn Wang
Publisher
Latent Space

23 / belief

Our accuracy, benchmark as part of a omniscience, it’s very correlated with total. It’s not correlated with, with active, uh, parameters, which I think is very at all, which is very, very interesting.

“Our accuracy, benchmark as part of a omniscience, it’s very correlated with total. It’s not correlated with, with active, uh, parameters, which I think is very at all, which is very, very interesting.”
Speaker
George Cameron
Publisher
Latent Space

24 / uncertainty

I don’t know if you have anything to add there. Or we could just go right into showing people the benchmark and like looking around and asking questions about it.

“I don’t know if you have anything to add there. Or we could just go right into showing people the benchmark and like looking around and asking questions about it.”
Speaker
Shawn Wang
Publisher
Latent Space

27 / belief

Do this. Because, yeah, the simplest thing that, like, our opinion is, is that there is a lot of advantage to having, like, an official OSI license like MIT or Apache 2, because then the box is just checked.

“Do this. Because, yeah, the simplest thing that, like, our opinion is, is that there is a lot of advantage to having, like, an official OSI license like MIT or Apache 2, because then the box is just checked.”
Publisher
Latent Space

28 / belief

I’m pretty confident that Blackwell has delivered pretty enormous gains and that the next couple of years of NVIDIA’s roadmap are going to continue to deliver quite enormous gains and that those will actually come through as lower total cost per token to the companies that are running models on them and will allow bigger models will allow way more tokens to be made for lower cost and that that’s gonna continue these things also stack on all of the software and model improvements. So basically like my prediction across like both sides of that, like smile chart, uh, that we’re gonna see the left-hand side continue to be true and probably like for another order of magnitude and the right-hand side continue to be true for another order of magnitude, and that’s gonna enable a whole lot of things.

“I’m pretty confident that Blackwell has delivered pretty enormous gains and that the next couple of years of NVIDIA’s roadmap are going to continue to deliver quite enormous gains and that those will actually come through as lower total cost per token to the companies that are running models on them and will allow bigger models will allow way more tokens to be made for lower cost and that that’s gonna continue these things also stack on all of the software and model improvements. So basically like my prediction across like both sides of that, like smile chart, uh, that we’re gonna see the left-hand side continue to be true and probably like for another order of magnitude and the right-hand side continue to be true for another order of magnitude, and that’s gonna enable a whole lot of things.”
Publisher
Latent Space

29 / belief

Let’s pick on hardware efficiency since you also have, you also track hardware stuff. And I think the general assertion or the message is that the efficiency from next gen Nvidia chips is actually not 4X.

“Let’s pick on hardware efficiency since you also have, you also track hardware stuff. And I think the general assertion or the message is that the efficiency from next gen Nvidia chips is actually not 4X.”
Speaker
Shawn Wang
Publisher
Latent Space

31 / belief

I think to some extent, I’m mixed opinion on that one because to some extent, your target audience is not people in AI Grants who are obviously at the frontier.

“I think to some extent, I’m mixed opinion on that one because to some extent, your target audience is not people in AI Grants who are obviously at the frontier.”
Speaker
Shawn Wang
Publisher
Latent Space

32 / preference

I mean, I think, I think for hallucinations specifically, there are a bunch of different things that you might care about reasonably, and that you’d measure quite differently, like we’ve called this a amnesty and solutionation rate, not trying to declare the, like, it’s humanity’s last hallucination.

“I mean, I think, I think for hallucinations specifically, there are a bunch of different things that you might care about reasonably, and that you’d measure quite differently, like we’ve called this a amnesty and solutionation rate, not trying to declare the, like, it’s humanity’s last hallucination.”
Publisher
Latent Space

33 / evaluation

The reason is that for this one here specifically, it would be very, very easy to like have data contamination because it is just factual knowledge questions.

“The reason is that for this one here specifically, it would be very, very easy to like have data contamination because it is just factual knowledge questions.”
Publisher
Latent Space

35 / evaluation

One extra, um, point regarding, um, GDP Val AA is that on the basis of the overperformance of the models compared to the chatbots turns out, we realized that, oh, like our reference harness that we built actually white works quite well on like gen generalist agentic tasks.

“One extra, um, point regarding, um, GDP Val AA is that on the basis of the overperformance of the models compared to the chatbots turns out, we realized that, oh, like our reference harness that we built actually white works quite well on like gen generalist agentic tasks.”
Speaker
George Cameron
Publisher
Latent Space

36 / observation

We don’t actually publish anything on this right now, but have tracked it a bunch internally in our internal analytics on evals across all the models that we run, where we look at the difficulty to questions and the correlation between token usage and difficulty and net net, surprise, surprise, like models have got.

“We don’t actually publish anything on this right now, but have tracked it a bunch internally in our internal analytics on evals across all the models that we run, where we look at the difficulty to questions and the correlation between token usage and difficulty and net net, surprise, surprise, like models have got.”
Publisher
Latent Space

38 / observation

Interestingly in Tal, um, Tal2Bench Telecom, it’s cheaper to run, you know, on a per token basis, more expensive models like a GBD5 compared to some smaller open source models, because the, um, some of the GBD5, for instance, uh, got to the answer faster.

“Interestingly in Tal, um, Tal2Bench Telecom, it’s cheaper to run, you know, on a per token basis, more expensive models like a GBD5 compared to some smaller open source models, because the, um, some of the GBD5, for instance, uh, got to the answer faster.”
Speaker
George Cameron
Publisher
Latent Space

41 / preference

What I love talking to people like you who sit across the ecosystem is, well, I have theories about what people want, but you have data and that’s obviously more relevant.

“What I love talking to people like you who sit across the ecosystem is, well, I have theories about what people want, but you have data and that’s obviously more relevant.”
Speaker
Shawn Wang
Publisher
Latent Space
Search evidence