Evidence receipt / commitment
Published · transcript-backedWei Lin: commitment
1 Nov 2024 Latent Space In the Arena: How LMSys changed LLM Benchmarking Forever
“On a high level, I think our goal here is to build a fast eval for everyone, and including everyone in the community can see the data board and understand, compare the models.”
— Wei Lin
Source trail
Everything needed to verify it.
- Speaker
- Wei Lin
- Attribution
- Verified speaker
- Claim type
- commitment
- Recorded
- 1 Nov 2024
- Publisher
- Latent Space
Transcript context
…O1 is a developing story. We still haven't seen the full model yet, but it's definitely a very exciting new paradigm. I think one community controversy I just wanted to give you guys space to address is the collaboration between you and the large model labs. People have been suspicious, let's just say, about how they choose to A-B test on you. I'll state the argument and let you respond, which is basically they run like five anonymous models and basically argmax their Elo on LMSYS or chatbot arena, and they release the best one. Right? What has been your end of the controversy? How have you decided to clarify your policy going forward? On a high level, I think our goal here is to build a fast eval for everyone, and including everyone in the community can see the data board and understand, compare the models. More importantly, I think we want to build the best eval also for model builders, like all these frontier labs building models. They're also internally facing a challenge, which is how do they eval the model? That's the reason why we want to partner with all the frontier lab people, and then to help them testing. That's one of the... We want to solve this technical challenge, which is eval. Yeah. I mean, ideally, it benefits everyone, right?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.