High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / recommendation

Published · transcript-backed

Shawn Wang: recommendation

12 Jul 2024 Latent Space Benchmarks 201: Why Leaderboards > Arenas >> LLM-as-Judge

“You just gave me an idea that Goodreads should be a data set because these are all novels that are commentaries about the contents of the novel.”

— Shawn Wang

Source trail

Everything needed to verify it.

Speaker
Shawn Wang
Attribution
Verified speaker
Claim type
recommendation
Recorded
12 Jul 2024
Publisher
Latent Space

Transcript context

…For the leaderboard specifically, that's why we added Muser, because it's long context reasoning. In terms of high quality long context reasoning benchmarks, I can think of two which I really like. One is called a benchmark for learning to translate a new language from one grammar book, and it's actually a very fun data set where they basically provide the LLM with a grammar book written by a linguist on a small language, which is super low resource called Kalamang. Since it's so low resource, you're sure that there is no data about it anywhere on the web. Then they ask questions about the grammar, what would be the correct form, etc. This is reasoning, this is language skills, this is super long context because it's a book. I think this data set is very interesting in terms of long context. There was also LNAI, which made a benchmark which they called a novel challenge for long context model, where basically they took full-on novels published last year. They asked people who had read it to do summaries and to do adversarial descriptions of events happening in the book, which require you to have understood the full book to answer. They prompted models with that, so also a very long context evaluation because you've got a full book and then you've got those questions that you need to answer correctly and also not contaminated because hopefully the books are not in training data yet. Yeah, it's new novels. Yeah, definitely. So I think those kind of data sets are more interesting. Yeah, go ahead. You just gave me an idea that Goodreads should be a data set because these are all novels that are commentaries about the contents of the novel. Definitely. There is definitely something to do about this.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence