Speakers in the public record
Claim mix
evaluation 7belief 7recommendation 2commitment 2observation 1
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
The useful parts, with receipts.
19 published records
“There was the psychophancy paper of Anthropic, where basically they showed that humans tend to prefer models which go their way and which agree with them because we want people to like us and apparently we want models to like us too and to agree with us too.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“On anything like an Arenas, there can be also the problem of the lack of diversity of the annotators, because most of the users of the chatbot arenas, for example, tend to be, from what I gathered, men from the US.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“I think people should stop using LLMS judges, because they have a lot of subtle biases that they introduce in evaluation.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“I'll throw in one more, which is, I think the sample of the chatbot arena data is actually out there.”
- Publisher
- Latent Space
“I think the only other, before this, it was kind of like ML perf, that's the other big leaderboard that I can think about where, obviously, maybe AlexNet, specific competitions, specific benchmarks, but not something that aggregates across all the other benchmarks.”
- Publisher
- Latent Space
“At the same time, I won't complain too much about this because we also tend to not evaluate multi-turn for automatic benchmark.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“What stuck with me though is that geology is very much an experimental science. And I think that machine learning is very much an experimental science too, even though people want to claim that it's pure math.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“There was also this cool paper by Cohere and the University of Edinburgh, which was human feedback is not gold standard, I think.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“If they manage to learn the patterns and generalize on something which is logic-based, then we will have reached a step in reasoning, which will be very interesting.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“You could use any multi-choice evaluation set to test this. I think there are more interesting datasets to build to test that.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“I think that robustness to prompting, nobody does it because it's too expensive.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“The thing is, at the moment, we really wanted to use a EleutherAI harness because it's a big stable library that everybody uses, and we think that Elusive is doing a great job at evaluations in general.”
- Speaker
- Not verified from transcript
- Publisher
- Latent Space
“Are people just finding out recently about these problems because now the scores are getting so high that you're actually inspecting the benchmarks and maybe in the past you were scoring so badly that maybe you weren't as worried about the overall quality?”
- Publisher
- Latent Space
“You just gave me an idea that Goodreads should be a data set because these are all novels that are commentaries about the contents of the novel.”
- Publisher
- Latent Space
“I think the other thing to talk about here is whether or not humans are good at judging and evaluating these models.”
- Publisher
- Latent Space
“I think obviously MMLU Pro is the top one, just because that's the top number that a lot of people report.”
- Publisher
- Latent Space
“Yeah, so I really like this concept of evaluation. So actually, yeah, I think there's typically what I always say is like sort of 25 is random chance, 50 is average human, 75 is expert human, 90 is you're cheating.”
- Publisher
- Latent Space
“Obviously, I think you're like our second or third person from Hugging Face on the podcast and it's like the definitional sort of open AI company, maybe the real open AI.”
- Publisher
- Latent Space
“You know, people just trust the person saying the thing assertively, even though it's false, and then actually try and figure out what the truth is. So yeah, I think you mentioned that, you know, it's like a more social experiment.”
- Publisher
- Latent Space