Evidence receipt / evaluation
Published · transcript-backedShawn Wang: evaluation
20 Sept 2024 Latent Space The Ultimate Guide to Prompting
“Endorse all that. And I think getting things into structured output and doing those scoring is a very core part of AI engineering that we don't talk about enough.”
Source trail
Everything needed to verify it.
- Speaker
- Shawn Wang
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 20 Sept 2024
- Publisher
- Latent Space
- Episode
- The Ultimate Guide to Prompting
Transcript context
…tter way of doing this is say, Oh, output on a scale of one through five, where one means completely fine, two means possible room for emergency, three means significant room for emergency, et cetera. So you really want to assign, make sure you assign meaning to the numbers. And there's other approaches like taking the probability of an output sequence and using that to actually evaluate the, I guess these are the log props, actually evaluate the probability. That has also been shown to be problematic. There's a couple of papers that directly analyze the technique and show it doesn't work in a lot of cases. So when you're doing these sort of evals, especially in sensitive domains like medical, you need to be robust in evaluation of your own evaluation system. Endorse all that. And I think getting things into structured output and doing those scoring is a very core part of AI engineering that we don't talk about enough. But so I wanted to make sure that we give you space to talk about it. We covered a lot.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.