Evidence receipt / commitment
Published · transcript-backedSpeaker unverified: commitment
12 Jul 2024 Latent Space Benchmarks 201: Why Leaderboards > Arenas >> LLM-as-Judge
“If they manage to learn the patterns and generalize on something which is logic-based, then we will have reached a step in reasoning, which will be very interesting.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- commitment
- Recorded
- 12 Jul 2024
- Publisher
- Latent Space
Transcript context
…Yes, the AGI challenge. Well, first of all, I'm curious if you think that actually is AGI, if you solve it, and just overall thoughts on the more challenge-driven things rather than evaluation, benchmarks driven. I don't think if you solve it, it's AGI. I think that focusing at the moment on trying to reach AGI is also a very bad objective, to be fair. But I'm very excited about this specific dataset. I'm looking forward to see what happens because I took a stab at some questions and basically they are great. They are pure logic. One of the things that we are missing at the moment in terms of LLM evaluations is complex logic, I think. Models are very bad at this. If they manage to learn the patterns and generalize on something which is logic-based, then we will have reached a step in reasoning, which will be very interesting. Just overall, more meta question. How do you figure out whether or not a benchmark is actually useful? Everybody wants to build benchmarks, kind of like test sets and things like that. Do you have any quick ways, kind of like you have a vibe check for model? Do you have a vibe check for benchmarks?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.