Evidence receipt / evaluation
Published · transcript-backedMike Israetel: evaluation
24 Dec 2025 Machine Learning Street Talk "I Desperately Want To Live In The Matrix" - Dr. Mike Israetel
“It does a great job of at least giving you things to think about. But I will say, know, I think prompting guides are like being a zip disk expert in the nineties.”
Source trail
Everything needed to verify it.
- Speaker
- Mike Israetel
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 24 Dec 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…Isn't isn't there a thing though that sometimes people don't know what questions to ask? I mean, I give a couple of examples. You know, let's say I've started taking retitrutide and genuinely, there's a lot of uncertainty and chat GPT will confidently tell you a lot of things that are not true, maybe. And maybe we'll discover some new information in 6 months' time, which casts shadows on what we did know. And on this asking the right question thing, think that part of knowing something well is just understanding it with the correct perspective. Right? So a naive person will go to ChatGPT and they'll frame the question in such a way that doesn't really elicit the correct answer. And it's incredibly sycophantic because people tend to lead the witness with ChatGPT. So they say, I've got a great idea and I believe this and my friend told me this and and it will say, yeah, that's a brilliant idea. You know, keep keep doing that thing. And it just doesn't know when it doesn't know and it doesn't like really try and correct you when you get it wrong. I used to have that experience with 4 0 model a lot. Then I had the o 3 research preview and it I found it to be insanely disagreeable almost to a fault. Like it would be like that's not That assumption is stupid. I wouldn't be asking that question and I was like, motherfucker. I came here for good vibes. And I would debate it incessantly and eventually into the ground and it would like callously concede at the end. And I guess you do make a good point but I would still point out that blah blah blah and I was like fuck you. And so GPT 5 is a synthesis of models but as the open AI people will tell you it's not a true synthesis yet, they're working on it. The personalities between the just ping the LLM GPT 5 regular and the GPT 5 thinking. It's a it's absolutely o 4 under the hood or o 3 or whatever it is with a little pixie dust on it because it it acts like a different entity. Now by the time this recording is released, that could very well be a solved problem. But I have found that based on what model you use, you get different personalities out of it. And so for real serious questions, used to go to o 3 and now go to GPT 5 pro or thinking. That thing is not psychophantic. And if that's psychophancy, I need more nice people in my life. Holy shit. But the regular like 4 0 especially, oh my god. It's the best friend you'll ever have. But the psychophancy especially if you lead it on like you said gets a little out of hand. You can get it to pretend with you, which is fun but useless. I will say to Jared's point on on the marginal nature of it. It does a great job of at least giving you things to think about. But I will say, know, I think prompting guides are like being a zip disk expert in the nineties. Like it's not the thing that's gonna last long time because AI is gonna tell you how to prompt itself real soon here. But for now, what I would say is a good way to prompt AI is, hey, here's my question and here's my perspective. Can you steel man my perspective? Can you red team my perspective? And you can you give me an evidence based and logical middle ground that doesn't commit the fallacy of the middle, but actually is just your best take as a thinking machine. That's a good prompt. Otherwise, who knows what the fucking mood it's in or how you phrase the question because it goes always giving you better answers. Every month, it gets better. But, like, sometimes it really can go on wild goose chases with you. So prompting it to give you both sides and then the middle ground is really good. The thing is I have that as a selection in my personality or architecture already, because I pay for Jared and I are piece of shit we pay for the pro model. And so it automatically gives me like a ground truth perspective with ups and downs on everything. And that's awesome. The thing is a lot of people are not interested in a lot of output for their chatty PT's. They want like 5 word answer. a ground truth perspective with ups and downs on everything. And that's awesome. The thing is a lot of people are not interested in a lot of output for their chatty PT's. They want like 5 word answer. Jared and I get 1,800 page essays out of every question from that fucking thing. So we want to drink from the well of knowledge. But drinking from the well of knowledge means you gotta process a lot of fucking text. And since we pay for those goddamn tokens, stupid proscription cost a zillion dollars worth every penny by the way, then we're gonna get them. But I think what a lot of people do and this is I think interesting thing. I saw an interesting interaction between 2 AI researchers on on on Twitter and X. And they were 1 guy was like, you know, I really think the best way to get a lot out of AI in the future of AI is massive context. Like an AI that remembers every chat you've ever had, which is actually nominal to do from computer science perspective. But obviously, like, it's just like there's just not enough data centers for them to do that. They will be soon. They're they're building 10 a day. But you know, then it's really gonna know your life and the way you ask it questions is jam it full of context and you get better answers. And I was like, well, that's categorically true. And then every reply by an also very smart person that was like, bullshit. He was like, you the real real intelligence is when you need to give it almost no context and it comes up with the right answer. And thing is on vibes, his reply was really cool because it's like, man, a really brilliant person. Like if you tell Jared like, hey, I wanna prep for bodybuilding show. You've been my coach. Here's my he's just gonna cut you off and be like, I'm gonna do the intake myself. I'm gonna ask you not so many questions, believe it or not. And then when I render my expertise, you're gonna do great. That's cool. But Jared has a lot of embodied knowledge, a lot of context already. These models don't. But what I thought so so on vibes, 100%, ideally that's the system. But the way you get to that system is you shit tons of context down its throat. Eventually, it needs less. But right now, it needs a lot. So I think 1 of the ways that people get in trouble with modern AI systems is they just pretend it's an infinite knowledge machine, and they're like, hey, what's the best restaurant in London? How the fuck would you but then it's actually an untenable proposition. Right? There are 50 ways to answer that question. All of them partially wrong. And so we don't say like, oh, the best review is this. And you're like, man, I can't believe Chad GPT may be waiting in line for 2 hours for fucking oxtail and duck. Like, yeah, you fucking asshole. It just fucking found the most expensive restaurant that all you rich idiots say is the best and it actually sucks. Should've sent you to Nando's. But if you were like, hey, here's the kind of food I like. Here's the atmosphere I like. Here's my price range. Here's my area. What's a really good restaurant here and tell me the top 5? Well, that's phenomenal answer, but you're too fucking lazy to type that in.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.