Evidence receipt / evaluation
Published · transcript-backedMicah-Hill Smith: evaluation
8 Jan 2026 Latent Space Artificial Analysis: Independent LLM Evals as a Service — with George Cameron and Micah-Hill Smith
“I mean, the computer stuff doesn’t work quite well enough and so on and so on, but yeah.”
Source trail
Everything needed to verify it.
- Speaker
- Micah-Hill Smith
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 8 Jan 2026
- Publisher
- Latent Space
Transcript context
…I mean, it’s in the data set. Like be a YouTuber? It’s a marketing video. Oh, wow. What? Like model has to go find clips on the internet and try to put it together. The models are not that good at doing that one, for now, to be clear. It’s pretty hard to do that with a code editor. I mean, the computer stuff doesn’t work quite well enough and so on and so on, but yeah. And so there’s no kind of ground truth, necessarily, to compare against, to work out percentage correct. It’s hard to come up with correct or incorrect there. And so it’s on a relative basis. And so we use an ELO approach to compare outputs from each of the models between the task.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.