Evidence receipt / evaluation
Published · transcript-backedJosh Albrecht: evaluation
25 Jun 2024 Latent Space State of the Art: Training >70B LLMs on 10,000 H100 clusters
“we thought about this, which is are emergent properties of language models a mirage? And I think if you look at that paper, it actually makes a relatively compelling case that in fact, you know, this emergent behavior that you're seeing is not really emergent behavior, but is really a function of the evaluation metrics that we're using.”
Source trail
Everything needed to verify it.
- Speaker
- Josh Albrecht
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 25 Jun 2024
- Publisher
- Latent Space
Transcript context
…a little bit of how we thought about this, which is are emergent properties of language models a mirage? And I think if you look at that paper, it actually makes a relatively compelling case that in fact, you know, this emergent behavior that you're seeing is not really emergent behavior, but is really a function of the evaluation metrics that we're using. So if you look at accuracy as a metric, what's happening is that accuracy is actually going up continually over training, but it's in log scale. So it starts out at 0.001%, 0.1, 0.1, 10. Only when you're going…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.