Speakers in the public record
Claim mix
evaluation 6belief 3preference 2commitment 1recommendation 1
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
13 published records
“It's obviously not reflective of the general population, but it's reflective of some corner of the world of people that really care. And to some extent, maybe that's all right, because those are like the power users.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“I think one of the reasons we still do this static benchmark, we still wanted to explore, experiment whether we can automate this, because people, eventually, model developers need it to fast iterate their model.”
- Publisher
- Latent Space
“Anastasios, I actually saw you, I think at last year's NeurIPS. You were presenting a paper, which I don't really super understand, but it was some theory paper about how your method was very dominating over other sort of search methods.”
- Publisher
- Latent Space
“The reason why you and others in the community have that instinct is because when you look at something like a benchmark, like an image net, a static benchmark, what happens is that if I give you a million different models that are all slightly different, and I pick the best one, there's something called selection bias that plays in, which is that the performance of the winning model is overstated.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“They came to use free LLM inference, right? And also, a lot of users just come to the website to use direct chat, because you can chat with the model for free.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“To be clear, none of the leaderboard currently is apples to apples because you have like Gemini Flash, you have, you know, all sorts of tiny models like Lama 8B, like 8B and 405B are not apples to apples.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“Hard in chatbot arena, but where it is easier is in red teaming, because in red teaming, there's an explicit game.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“I think the point is that, and you guys probably also talked about this at some point, but static benchmarks are intrinsically, to some extent, unable to measure generative model performance.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“I think my fundamental philosophical doubt is, does the router model have to be at least as smart as the smartest model?”
- Publisher
- Latent Space
“I love the notebooks you guys publish. Actually really good just for learning statistics.”
- Publisher
- Latent Space
“On a high level, I think our goal here is to build a fast eval for everyone, and including everyone in the community can see the data board and understand, compare the models.”
- Publisher
- Latent Space
“The classic one for human preference evaluation is humans demonstrably prefer longer contexts or longer outputs, which is actually something that we don't necessarily want. You guys, I think maybe two months ago put out some length control studies.”
- Publisher
- Latent Space
“Blast from the past. It's always interesting how NeurIPS and all these academic conferences are sort of six months behind what people are actually doing, but conformal risk control, I would recommend people check it out.”
- Publisher
- Latent Space