Evidence receipt / evaluation
Published · transcript-backedJosh Albrecht: evaluation
25 Jun 2024 Latent Space State of the Art: Training >70B LLMs on 10,000 H100 clusters
“I would much rather have a coding agent that will give me back a thing. And, you know, it's it's actually the code doesn't work like 10% less of the time than some other model, but it will tell me 100% of the time like when it's not sure.”
Source trail
Everything needed to verify it.
- Speaker
- Josh Albrecht
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 25 Jun 2024
- Publisher
- Latent Space
Transcript context
…But actually, one thing that we're very interested in is ambiguity itself. Like, can we detect whether a task from a user is ambiguous or whether you've, you know, completed a task successfully? Like these are actually hard, messy problems, but are really important from like the user experience of using these models. I would much rather have a coding agent that will give me back a thing. And, you know, it's it's actually the code doesn't work like 10% less of the time than some other model, but it will tell me 100% of the time like when it's not sure. Like that's so much more useful if it can communicate like, I'm not really sure about this or maybe there's some errors here. Then just like, here's some code. I have no idea if it works. And so these kind of like, you know, detecting ambiguity and detecting correctness or uncertainty,…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.