01 / recommendation
Likes Andy.
“I'm I'm huge fan of Andy's, big hero of mine. You probably didn't think that they would be so relevant later on in your career.”
- Speaker
- Tim Scarfe
- Publisher
- Machine Learning Street Talk
Machine Learning Street Talk / episode intelligence
Speakers in the public record
Claim mix
Evidence policy
Every row below preserves an exact excerpt. Identified speakers are linked; unresolved voices are labeled and excluded from people counts.
Claim ledger
22 published records
01 / recommendation
“I'm I'm huge fan of Andy's, big hero of mine. You probably didn't think that they would be so relevant later on in your career.”
02 / belief
“I think people working at the very frontier of state of the art models, I think, do believe that, that that evals need to be rigorous and robust and humans have to be in the loop. But I think we are also at a very, very early stage of sort of break it and apologize later, where I think a big swathe of our industry doesn't yet think that that sort of human mediated evaluation is going to be important and perhaps will get in the way of innovation.”
03 / belief
“The first misconception is that the work is sort of grinding or boring, which actually in our experience when it comes to providing either training or evaluation or fine tuning data for SOTA models, it's not.”
04 / belief
“There's this image of it, I think, in the industry of being seedy. And I think that's not helped by some of the history of how, you know, data has been extracted from human beings whether or not they they know that their data is being used.”
05 / belief
“I'm I'm I'm really sort of vibing with the way you described that. And I think the very first thing we need to do is try to get perspective into the mix through representativeness.”
06 / belief
“Massive kudos to Anthropic for releasing the entire methodology quite openly, I think, and transparently.”
07 / belief
“You know, humans live in society and they tend to share their cultural beliefs with their tribes. And I think that's why being able to stratify the data you gather for evaluation from people, I think is quite powerful actually.”
08 / belief
“Although we are getting much better at the cross functional collaboration between researchers, academics, industry people, public bodies, and it's just such a very interesting point in, I think, the history of science and the future of science.”
09 / observation
“I am amenable by the way to this idea of loss of control, you know, which is that we we we start to build systems on top of systems on top of systems, and it's a little bit like the power station.”
10 / belief
“You know, those very, very enduring problems of what it means to be a person and a thinker are suddenly impossible to ignore as we're doing software releases and model deployments. I think yeah, that that's super unexpected for me.”
11 / belief
“We're all trying to deliver stuff and there's a, you know, accelerating sort of, you know, hot industry around us, and the idea of waiting on a human to tell us if a thing worked feels counterintuitive, and I think our approach to that is stick a really well treated, verified, diversely demographic human behind an API, essentially, and make sure that the structures and infrastructure are there to ensure that human can go fast, understand instructions, and give you something akin to deterministic human in the loop behaviors.”
12 / belief
“I think 25 was exactly the same and then I think the other 25% was nearly the same like 95% cosine distance on the embeddings or something like that.”
13 / evaluation
“In fact, we we have a whole interface where they they can interact directly with some of the researchers or some of the the program coordinators, and they're they're very proactive in the type of things that they do, and that obviously this all aids in also highlighting certain edge cases or they're contributing and refining the process, for example, or these kind of things. This is more like an active participation and interest in the outcome and success of whatever is being collected, which is really good and it all contributes to this ultimately increasing the quality of the data that is being used because this data in very, very, very often the case is being used in either the training of or evaluation of very, very central core systems that have huge downstream effect on everybody on this earth most likely.”
14 / evaluation
“So I was talking to Claude about the halting problem a few days ago, and I was like, well, do you mean the algorithm won't end, or do you mean someone won't unplug the computer? And Claude goes, no, I mean, the algorithm won't end, you know, if we cannot confirm or deny whether the algorithm would end, and I just I just thought to myself, that seems kind of nonsensical to me because, you know, the universe will end.”
15 / evaluation
“There's even, I believe, some countries where patients themselves are not allowed to receive transcripts from the testing facility on the chance of the patient misinterpreting the results.”
16 / evaluation
“For example, reinforcement learning is a really, really good example of this, where, for example, web agents, where are now largely trained or in these environments that are often synthetically created, but there's programs that need to create these synthetic environments for RL agents to to explore in that need to be validated by humans. So we're seeing this interesting progression where humans are no longer directly involved in, like, the main system of interest, but sort of in this secondary, almost like a second order or third order abstraction moving outside, which actually is a welcome change because that means that we're focusing more on the right kind of tasks where humans are relevant, and we're focusing more on the right kind of high quality data.”
17 / prediction
“You know, if we had perfect verifiers and and synthetic data generators, probably we wouldn't even need LLMs in the first place, right, because we've already solved all the problems.”
18 / evaluation
“You know, when when something is, you know, trivially easy to mechanize, no 1 actually thought it was intelligent. But I think the Turing test is bad because we know that in my opinion, language models aren't actually that intelligent, yet we've passed it with flying colors.”
19 / evaluation
“Chatbot Arena is somewhere in the realm of in between because you're you're not actually ranking or you're not evaluating for 1 or the other. It is technically preference but you don't know quite whether the preference is because it said something wrong or whether the formatting was off or whether it didn't hit the cultural relativity or sensitivity or whether it wasn't adaptive enough, it doesn't tell you anything of the sorts.”
20 / evaluation
“That's nice if we can align on the measurement that we consider success. And I think that's lacking to some degree because we have somehow inherently decided that chatbot arena for example is the measure of success, so people optimize for it.”
21 / prediction
“They are very, very heavy foundational models with lots and lots of capabilities, where also tons of the fine tuned models are effectively descendants of these models. So the more effort we put into the foundational model, we will ultimately benefit also all of the descendants and derivatives of these models.”
22 / preference
“I love the word orchestration because it has the root word orchestra, which is kind of this beautiful, collaborative, you know, symphonic word, and I think that to me is the, you know, correcting a 5 year old over and over again, or correcting an 18 year old or a 25 year old about life, right, over and over again.”