Evidence receipt / evaluation
Published · transcript-backedTim Scarfe: evaluation
20 Aug 2026 Machine Learning Street Talk Every Exponential Ends — Silicon Valley Forgot — Adam Becker
“Maybe we can RL train it to refactor it, and we just have to kind of and it's very dangerous keeping going because now we're messing all of our code bases up and we're we're creating all of this slop everywhere.”
Source trail
Everything needed to verify it.
- Speaker
- Tim Scarfe
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 20 Aug 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…but but, again, I mean, not to repeat myself, but I do think that the difference here is that it's language, not math. And so that makes it feel like something is is sort of thinking and conscious and talking to us. And, you know, again, that's not to diminish it, you know. I'm not diminishing the the functionality of calculators either. But it is quite deceptive, and and I think I I in terms of how to make sense of it, I think we just have to keep in mind what these systems are at the end of the day. They are for predicting the next word or whatever in the sequence, that they've been given, and and it turns out you can get pretty far doing that. You can. And I would push back a little bit on the stochastic parrot thing even though that's technically true. They they I mean, internally, they are they are acquiring during their training process some kind of structure Oh, yeah. Which allows them to, you know, extrapolate, generalize, call it what you want. So, you know, it's different to us, but there's something there. But I I think the the $2,000,000 questions are, at the moment, it needs human supervision. And these folks say, well, we we just scale is all you need. Yeah. At the moment, it needs human supervision. And at the moment, it just generates loads of spaghetti garbage code and, you know, it just creates basically, it creates more problems than than it solves. But it's deceptive because most people can't see the problems. But all we need to do is keep scaling it. You know, when when GPT 7 comes out, it will actually refactor all that code. Maybe we can RL train it to refactor it, and we just have to kind of and it's very dangerous keeping going because now we're messing all of our code bases up and we're we're creating all of this slop everywhere. Just just hold on, boys. Just wait for a couple of years, and and the the next version will kind of bring it back in check. What do you think about that? I just don't think that that's true. I don't think that there's I mean, I could be wrong, obviously, but I don't see good evidence for that. Right? I mean, I think that these systems are you know, unless there's some sort of fundamental breakthrough, right, something more than just scale, these systems are always gonna require human supervision because they are always going to end up hallucinating. You know? I I mean, that's inherent to the way that they work. I say this in the book, but, you know, I don't love the word hallucinate. I know there's been a lot of pushback because it's like, oh, hallucination implies a sort of anthropomorphization of these systems. And I don't love that either, but that's actually not my main problem. My main problem with the word hallucination is it implies that when a hallucination occurs, something different is happening than its normal functioning. And that's not the case. They really only do 1 thing. And when they're hallucinating, they're doing the same thing that they're doing when they get it right. And so I I think that unless we have some sort of major major breakthrough, and I mean, it would probably have to be a breakthrough that makes, you know, LLMs themselves look like Eliza. You know, short of that kind of really fundamental breakthrough, I don't see us getting around the need for human supervision on these things. So and and if anything, as they get better, it's going to get harder to discern when they've made these mistakes even though they're gonna keep making them and that's quite dangerous as you said. I know. And and ironically,…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.