High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / recommendation

Published · transcript-backed

Micah-Hill Smith: recommendation

8 Jan 2026 Latent Space Artificial Analysis: Independent LLM Evals as a Service — with George Cameron and Micah-Hill Smith

“It depends what you’re looking at, right? Because you can, if you’re trying to see whether or not it can solve a particular type of reasoning problem, and you don’t want to test it on its ability to do answer formatting at the same time, then you might want to use an LLM as answer extractor approach to make sure that you get the answer out no matter how unanswered.”

— Micah-Hill Smith

Source trail

Everything needed to verify it.

Speaker
Micah-Hill Smith
Attribution
Verified speaker
Claim type
recommendation
Recorded
8 Jan 2026
Publisher
Latent Space

Transcript context

…And so for people who haven’t done this work, literally parsing the responses is a whole thing, right? Like because sometimes the models, the models can answer any way they feel fit and sometimes they actually do have the right answer, but they just returned the wrong format and they will get a zero for that unless you work it into your parser. And that involves more work. And so, I mean, but there’s an open question whether you should give it points for not following your instructions on the format. It depends what you’re looking at, right? Because you can, if you’re trying to see whether or not it can solve a particular type of reasoning problem, and you don’t want to test it on its ability to do answer formatting at the same time, then you might want to use an LLM as answer extractor approach to make sure that you get the answer out no matter how unanswered. But these days, it’s mostly less of a problem. Like, if you instruct a model and give it examples of what the answers should look like, it can get the answers in your format, and then you can do, like, a simple regex. Yeah, yeah. And then there’s other questions around, I guess, sometimes if you have a multiple choice question, sometimes there’s a bias towards the first answer, so you have to randomize the responses. All these nuances, like, once you dig into benchmarks, you’re like, I don’t know how anyone believes the numbers on all these things. It’s so dark magic.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence