Evidence receipt / observation
Published · transcript-backedGeorge Cameron: observation
8 Jan 2026 Latent Space Artificial Analysis: Independent LLM Evals as a Service — with George Cameron and Micah-Hill Smith
“Interestingly in Tal, um, Tal2Bench Telecom, it’s cheaper to run, you know, on a per token basis, more expensive models like a GBD5 compared to some smaller open source models, because the, um, some of the GBD5, for instance, uh, got to the answer faster.”
Source trail
Everything needed to verify it.
- Speaker
- George Cameron
- Attribution
- Verified speaker
- Claim type
- observation
- Recorded
- 8 Jan 2026
- Publisher
- Latent Space
Transcript context
…TalBench Retail, TalBench Airline. Yeah. Interestingly in Tal, um, Tal2Bench Telecom, it’s cheaper to run, you know, on a per token basis, more expensive models like a GBD5 compared to some smaller open source models, because the, um, some of the GBD5, for instance, uh, got to the answer faster. And so it was able to resolve the customer’s query faster and fewer turns. And maybe it used more tokens per turn, but it certainly- It’s not going to cost more per token. So you would always rather use GBD5 in, in, in, in that scenario. And so I think that’s what, that’s where we’re getting to. I think number of turns is, it’s going to be a metric that we’re going to be talking about a lot more. And, uh, I think it’ll be something that people want to really start to think about, uh, a lot more. There’s a trade-off in benchmarking here where most benchmarks needs to be one turn to be autonomous, to be parallelized and all that. But most, a lot of real life use cases need to be multi-turn and especially like quick multi-turns. So you can align. Yeah.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.