Evidence receipt / preference
Published · transcript-backedShawn Wang: preference
8 Jan 2026 Latent Space Artificial Analysis: Independent LLM Evals as a Service — with George Cameron and Micah-Hill Smith
“What I love talking to people like you who sit across the ecosystem is, well, I have theories about what people want, but you have data and that’s obviously more relevant.”
Source trail
Everything needed to verify it.
- Speaker
- Shawn Wang
- Attribution
- Verified speaker
- Claim type
- preference
- Recorded
- 8 Jan 2026
- Publisher
- Latent Space
Transcript context
…We built it because we needed it as people building in the space and thought, Oh, other people might find it useful too. So we’ll buy domain and link it to the Vercel deployment that we had and tweet about it. And, but very quickly it started getting attention. Thank you, Swyx for, I think doing an initial retweet and spotlighting it there. This project that we released. And then very quickly though, it was useful to others, but very quickly it became more useful as the number of models released accelerated. We had Mixtrel 8x7B and it was a key. That’s a fun one. Yeah. Like a open source model that really changed the landscape and opened up people’s eyes to other serverless inference providers and thinking about speed, thinking about cost. And so that was a key. And so it became more useful quite quickly. Yeah. What I love talking to people like you who sit across the ecosystem is, well, I have theories about what people want, but you have data and that’s obviously more relevant. But I want to stay on the origin story a little bit more. When you started out, I would say, I think the status quo at the time was every paper would come out and they would report their numbers versus competitor numbers. And that’s basically it. And I remember I did the legwork. I think everyone has some knowledge. I think there’s some version of Excel sheet or a Google sheet where you just like copy and paste the numbers from every paper and just post it up there. And then sometimes they don’t line up because they’re independently run. And so your numbers are going to look better than... Your reproductions of other people’s numbers are going to look worse because you don’t hold their models correctly or whatever the excuse is. I think then Stanford Helm, Percy Liang’s project would also have some of these numbers. And I don’t know if there’s any other source that you can cite. The way that if I were to start artificial analysis at the same time you guys started, I would have used the Luther AI’s eval framework harness. Yup. Yup. That was some cool stuff. At the end of the day, running these evals, it’s like if it’s a simple Q&A eval, all you’re doing is asking a list of questions and checking if the answers are right, which shouldn’t be that crazy. But it turns out there are an enormous number of things that you’ve got control for. And I mean, back when we started the website. Yeah. Yeah. Like one of the reasons why we realized that we had to run the evals ourselves and couldn’t just take rules from the labs was just that they would all prompt the models differently. And when you’re competing over a few points, then you can pretty easily get- You can put the answer into the model. Yeah. That in the extreme. And like you get crazy cases like back when I’m Googled a Gemini 1.0 Ultra and needed a number that would say it was better than GPT-4 and like constructed, I think never published like chain of thought examples. 32 of them in every topic in MLU to run it, to get the score, like there are so many things that you- They never shipped Ultra, right? That’s the one that never made it up. Not widely. Yeah. Yeah. Yeah. I mean, I’m sure it existed, but yeah. So we were pretty sure that we needed to run them ourselves and just run them in the same way across all the models. Yeah. And we were, we also did certain from the start that you couldn’t look at those in isolation. You needed to look at them alongside the cost and performance stuff. Yeah.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.