Evidence receipt / preference
Published · transcript-backedWilliam Beauchamp: preference
26 Jan 2025 Latent Space Outlasting Noam Shazeer, crowdsourcing Chai AI with >1.4m DAU, and becoming the "Western DeepSeek" — with William Beauchamp, Chai Research
“I think Elo is a fantastic north star and the reason for it, or like it's the main one we want to see go up because it's this human feedback.”
Source trail
Everything needed to verify it.
- Speaker
- William Beauchamp
- Attribution
- Verified speaker
- Claim type
- preference
- Recorded
- 26 Jan 2025
- Publisher
- Latent Space
Transcript context
…So yeah, like Elo cannot be the only eval. You must have internal evals. You mentioned evals. I think Elo is a fantastic north star and the reason for it, or like it's the main one we want to see go up because it's this human feedback. The humans know what they want. It's beautiful because when you come up with an eval, you're further removing yourself away from the true problem. Right? So whatever it is you're trying to optimize or figure out, you kind of have to, have to slice it. And then you've got this, it's like a snapshot. Like as soon as you saturate one eval, you need to figure out a new eval. But with, by saying to humans, just which is better, A or B, it's super robust. It's super generalizable. It just keeps, keeps scaling. So we've in the past used evals to get through a, to get through a blocker. I mean, a great example is, you know, is like having like a safety filter or something. Yeah. Where you want to make sure your models, because listen, users find, you'll be shocked the correlation between not family friendly content, whether that's just like swearing, like people find it funny when the AI swears. So if you have two completions, A or B, like if you give me any LLM, I can make it 20% funnier just by training it to throw in swear words. So the issue with that is it's like, how are we measuring like quality improvements? Are we measuring superficial improvements? Right. And this actually links back to the LLM sys. They did a style control. We actually had them on the podcast.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.