High Signal Podcasts Evidence ledger
Method
Browse
← All source episodes

Latent Space / episode intelligence

In the Arena: How LMSys changed LLM Benchmarking Forever

1 Nov 2024 13 published claims 2 attributable people

Speakers in the public record

Claim mix

evaluation 6belief 3preference 2commitment 1recommendation 1

Evidence policy

Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.

Claim ledger

The useful parts, with receipts.

13 published records

01 / evaluation

It's obviously not reflective of the general population, but it's reflective of some corner of the world of people that really care. And to some extent, maybe that's all right, because those are like the power users.

“It's obviously not reflective of the general population, but it's reflective of some corner of the world of people that really care. And to some extent, maybe that's all right, because those are like the power users.”
Speaker
Not verified from transcript
Publisher
Latent Space

02 / belief

I think one of the reasons we still do this static benchmark, we still wanted to explore, experiment whether we can automate this, because people, eventually, model developers need it to fast iterate their model.

“I think one of the reasons we still do this static benchmark, we still wanted to explore, experiment whether we can automate this, because people, eventually, model developers need it to fast iterate their model.”
Speaker
Wei Lin
Publisher
Latent Space

03 / belief

Anastasios, I actually saw you, I think at last year's NeurIPS. You were presenting a paper, which I don't really super understand, but it was some theory paper about how your method was very dominating over other sort of search methods.

“Anastasios, I actually saw you, I think at last year's NeurIPS. You were presenting a paper, which I don't really super understand, but it was some theory paper about how your method was very dominating over other sort of search methods.”
Speaker
Shawn Wang
Publisher
Latent Space

04 / evaluation

The reason why you and others in the community have that instinct is because when you look at something like a benchmark, like an image net, a static benchmark, what happens is that if I give you a million different models that are all slightly different, and I pick the best one, there's something called selection bias that plays in, which is that the performance of the winning model is overstated.

“The reason why you and others in the community have that instinct is because when you look at something like a benchmark, like an image net, a static benchmark, what happens is that if I give you a million different models that are all slightly different, and I pick the best one, there's something called selection bias that plays in, which is that the performance of the winning model is overstated.”
Speaker
Not verified from transcript
Publisher
Latent Space

05 / evaluation

They came to use free LLM inference, right? And also, a lot of users just come to the website to use direct chat, because you can chat with the model for free.

“They came to use free LLM inference, right? And also, a lot of users just come to the website to use direct chat, because you can chat with the model for free.”
Speaker
Not verified from transcript
Publisher
Latent Space

06 / evaluation

To be clear, none of the leaderboard currently is apples to apples because you have like Gemini Flash, you have, you know, all sorts of tiny models like Lama 8B, like 8B and 405B are not apples to apples.

“To be clear, none of the leaderboard currently is apples to apples because you have like Gemini Flash, you have, you know, all sorts of tiny models like Lama 8B, like 8B and 405B are not apples to apples.”
Speaker
Not verified from transcript
Publisher
Latent Space

08 / belief

I think the point is that, and you guys probably also talked about this at some point, but static benchmarks are intrinsically, to some extent, unable to measure generative model performance.

“I think the point is that, and you guys probably also talked about this at some point, but static benchmarks are intrinsically, to some extent, unable to measure generative model performance.”
Speaker
Not verified from transcript
Publisher
Latent Space

11 / commitment

On a high level, I think our goal here is to build a fast eval for everyone, and including everyone in the community can see the data board and understand, compare the models.

“On a high level, I think our goal here is to build a fast eval for everyone, and including everyone in the community can see the data board and understand, compare the models.”
Speaker
Wei Lin
Publisher
Latent Space

12 / evaluation

The classic one for human preference evaluation is humans demonstrably prefer longer contexts or longer outputs, which is actually something that we don't necessarily want. You guys, I think maybe two months ago put out some length control studies.

“The classic one for human preference evaluation is humans demonstrably prefer longer contexts or longer outputs, which is actually something that we don't necessarily want. You guys, I think maybe two months ago put out some length control studies.”
Speaker
Shawn Wang
Publisher
Latent Space

13 / recommendation

Blast from the past. It's always interesting how NeurIPS and all these academic conferences are sort of six months behind what people are actually doing, but conformal risk control, I would recommend people check it out.

“Blast from the past. It's always interesting how NeurIPS and all these academic conferences are sort of six months behind what people are actually doing, but conformal risk control, I would recommend people check it out.”
Speaker
Shawn Wang
Publisher
Latent Space
Search evidence