Evidence receipt / commitment
Published · transcript-backedGeorge Cameron: commitment
8 Jan 2026 Latent Space Artificial Analysis: Independent LLM Evals as a Service — with George Cameron and Micah-Hill Smith
“We built it because we needed it as people building in the space and thought, Oh, other people might find it useful too.”
Source trail
Everything needed to verify it.
- Speaker
- George Cameron
- Attribution
- Verified speaker
- Claim type
- commitment
- Recorded
- 8 Jan 2026
- Publisher
- Latent Space
Transcript context
…That’s actually true. I don’t even think we’d pause like, like George had an acquittance job. I didn’t quit working on my legal AI thing. Like it was genuinely a side project. We built it because we needed it as people building in the space and thought, Oh, other people might find it useful too. So we’ll buy domain and link it to the Vercel deployment that we had and tweet about it. And, but very quickly it started getting attention. Thank you, Swyx for, I think doing an initial retweet and spotlighting it there. This project that we released. And then very quickly though, it was useful to others, but very quickly it became more useful as the number of models released accelerated. We had Mixtrel 8x7B and it was a key. That’s a fun one. Yeah. Like a open source model that really changed the landscape and opened up people’s eyes to other serverless inference providers and thinking about speed, thinking about cost. And so that was a key. And so it became more useful quite quickly. Yeah. What I love talking to people like you who sit across the ecosystem is, well, I have theories about what people want, but you have data and that’s obviously more relevant. But I want to stay on the origin story a little bit more. When you started out, I would say, I think the status quo at the time was every paper would come out and they would report their numbers versus competitor numbers. And that’s basically it. And I remember I did the legwork. I think everyone has some knowledge. I think there’s some version of Excel sheet or a Google sheet where you just like copy and paste the numbers from every paper and just post it up there. And then sometimes they don’t line up because they’re independently run. And so your numbers are going to look better than... Your reproductions of other people’s numbers are going to look worse because you don’t hold their models correctly or whatever the excuse is. I think then Stanford Helm, Percy Liang’s project would also have some of these numbers. And I don’t know if there’s any other source that you can cite. The way that if I were to start artificial analysis at the same time you guys started, I would have used the Luther AI’s eval framework harness. Yup.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.