High Signal Podcasts Evidence ledger
Method
Browse

Public evidence record

Enzo Blindow

Published podcast speaker

Claims
6
Episodes
1
Shows
1
Named items
0

Claim ledger

What Enzo said.

5 transcript-backed records

01 / evaluation

In fact, we we have a whole interface where they they can interact directly with some of the researchers or some of the the program coordinators, and they're they're very proactive in the type of things that they do, and that obviously this all aids in also highlighting certain edge cases or they're contributing and refining the process, for example, or these kind of things. This is more like an active participation and interest in the outcome and success of whatever is being collected, which is really good and it all contributes to this ultimately increasing the quality of the data that is being used because this data in very, very, very often the case is being used in either the training of or evaluation of very, very central core systems that have huge downstream effect on everybody on this earth most likely.

“In fact, we we have a whole interface where they they can interact directly with some of the researchers or some of the the program coordinators, and they're they're very proactive in the type of things that they do, and that obviously this all aids in also highlighting certain edge cases or they're contributing and refining the process, for example, or these kind of things. This is more like an active participation and interest in the outcome and success of whatever is being collected, which is really good and it all contributes to this ultimately increasing the quality of the data that is being used because this data in very, very, very often the case is being used in either the training of or evaluation of very, very central core systems that have huge downstream effect on everybody on this earth most likely.”
Speaker
Enzo Blindow
Publisher
Machine Learning Street Talk

02 / evaluation

There's even, I believe, some countries where patients themselves are not allowed to receive transcripts from the testing facility on the chance of the patient misinterpreting the results.

“There's even, I believe, some countries where patients themselves are not allowed to receive transcripts from the testing facility on the chance of the patient misinterpreting the results.”
Speaker
Enzo Blindow
Publisher
Machine Learning Street Talk

03 / evaluation

For example, reinforcement learning is a really, really good example of this, where, for example, web agents, where are now largely trained or in these environments that are often synthetically created, but there's programs that need to create these synthetic environments for RL agents to to explore in that need to be validated by humans. So we're seeing this interesting progression where humans are no longer directly involved in, like, the main system of interest, but sort of in this secondary, almost like a second order or third order abstraction moving outside, which actually is a welcome change because that means that we're focusing more on the right kind of tasks where humans are relevant, and we're focusing more on the right kind of high quality data.

“For example, reinforcement learning is a really, really good example of this, where, for example, web agents, where are now largely trained or in these environments that are often synthetically created, but there's programs that need to create these synthetic environments for RL agents to to explore in that need to be validated by humans. So we're seeing this interesting progression where humans are no longer directly involved in, like, the main system of interest, but sort of in this secondary, almost like a second order or third order abstraction moving outside, which actually is a welcome change because that means that we're focusing more on the right kind of tasks where humans are relevant, and we're focusing more on the right kind of high quality data.”
Speaker
Enzo Blindow
Publisher
Machine Learning Street Talk

04 / evaluation

Chatbot Arena is somewhere in the realm of in between because you're you're not actually ranking or you're not evaluating for 1 or the other. It is technically preference but you don't know quite whether the preference is because it said something wrong or whether the formatting was off or whether it didn't hit the cultural relativity or sensitivity or whether it wasn't adaptive enough, it doesn't tell you anything of the sorts.

“Chatbot Arena is somewhere in the realm of in between because you're you're not actually ranking or you're not evaluating for 1 or the other. It is technically preference but you don't know quite whether the preference is because it said something wrong or whether the formatting was off or whether it didn't hit the cultural relativity or sensitivity or whether it wasn't adaptive enough, it doesn't tell you anything of the sorts.”
Speaker
Enzo Blindow
Publisher
Machine Learning Street Talk

05 / evaluation

That's nice if we can align on the measurement that we consider success. And I think that's lacking to some degree because we have somehow inherently decided that chatbot arena for example is the measure of success, so people optimize for it.

“That's nice if we can align on the measurement that we consider success. And I think that's lacking to some degree because we have somehow inherently decided that chatbot arena for example is the measure of success, so people optimize for it.”
Speaker
Enzo Blindow
Publisher
Machine Learning Street Talk
Search evidence