High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / evaluation

Published · transcript-backed

Josh Albrecht: evaluation

25 Jun 2024 Latent Space State of the Art: Training >70B LLMs on 10,000 H100 clusters

“we thought about this, which is are emergent properties of language models a mirage? And I think if you look at that paper, it actually makes a relatively compelling case that in fact, you know, this emergent behavior that you're seeing is not really emergent behavior, but is really a function of the evaluation metrics that we're using.”

— Josh Albrecht

Source trail

Everything needed to verify it.

Speaker
Josh Albrecht
Attribution
Verified speaker
Claim type
evaluation
Recorded
25 Jun 2024
Publisher
Latent Space

Transcript context

…a little bit of how we thought about this, which is are emergent properties of language models a mirage? And I think if you look at that paper, it actually makes a relatively compelling case that in fact, you know, this emergent behavior that you're seeing is not really emergent behavior, but is really a function of the evaluation metrics that we're using. So if you look at accuracy as a metric, what's happening is that accuracy is actually going up continually over training, but it's in log scale. So it starts out at 0.001%, 0.1, 0.1, 10. Only when you're going…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence