Evidence receipt / belief
Published · transcript-backedSpeaker unverified: belief
2 May 2025 · 1:10 Hard Fork How Nice Is Too Nice for an AI Chatbot? | EP 134
“Like what sort of effect would I need to have on the development of AI for you to be like, "All right, well, I guess I got to do a chapter about Casey." I think I think there are a couple routes you could take.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- belief
- Recorded
- 2 May 2025 · 1:10
- Publisher
- Hard Fork
Transcript context
…as addictive as you might have found Instagram or Tik Tok, I don't think it's going to be as addictive as some sort of digital entity that is sending you text messages throughout the day, that is agreeing with everything that you say, that is much more comforting and nurturing and approving of you than anyone you know in real life. There was definitely a moment where I was sitting in the press conference hearing about like the one world money with like the decentralized one world governance scheme and I just had this sort of like moment of like the future's so weird. It's so weird. It's so weird. This is the Haley's Comet of memes and it just is about to hit us again. [Music] Well, Casey, uh, as you know, I am writing a book. Yes. And congratulations. Uh, can't wait to read it. Yeah. Uh, I can't wait to write it. Uh, so this is the book is called the AGI Chronicles. It's basically the inside story of the uh, race to creating uh, artificial general intelligence. Now, here's a question. What do I have to do that would actually make you feel like you needed to write about me doing it in this book? You know what I mean? Like what sort of effect would I need to have on the development of AI for you to be like, "All right, well, I guess I got to do a chapter about Casey." I think I think there are a couple routes you could take. One would be that you could um make some, you know, breakthrough in reinforcement learning or develop some new algorithmic optimization uh that really pushes the field forward. So, let's take that off the table. Um, the next thing you could do would be to be uh sort of a case study um in in what happens when sort of powerful AI systems are unleashed onto an unwitting populace. So, you could be sort of a hilarious case study like you could have it give you some medical advice and then follow it and end up like amputating your own leg. I don't know. Do you have any ideas? Yeah, I was going to amputate my own leg at the instructions of a chatbot. So, it sounds like we're on the same page. I'll get right on that. Um, I knew that reading your next book was going to cost me an arm and a leg, but not like this. I'm Kevin Roose, a tech colonist at the New York Times. I'm Casey Newton from Platformer, and this is Hardfork. This week, the chatbot flattery crisis. We'll tell you the problem with the new, more sickopantic AIs. Then, Kevin takes a field trip to see the unveiling of a new orb. And finally, we're opening up our group chats with the help of podcaster PJ Vote. Oh, Casey. Another thing we should talk about, our show is sold out. That's right. Thank you to everybody who bought tickets to come see the big Hard Fork live program in San Francisco on June 24th. We're very excited. It's going to be so much fun. We haven't even said who the special guests are. So, uh, and we never will. Yeah. So, thanks to everyone who bought tickets. If you didn't manage to make it in time, uh, there is a wait list available on the website at ny who the special guests are. So, uh, and we never will. Yeah. So, thanks to everyone who bought tickets. If you didn't manage to make it in time, uh, there is a wait list available on the website at ny times.comventshardforlive. Hey Kevin, did a chatbot say anything nice to you this week? Chatbots never say anything nice to me. Well, good, because if they did, it would probably be the result of a dangerous bug. Yes. Uh you're talking, I'm guessing, about the uh drama this week over the syphancy problem in some of our leading AI models. Yes. They say that flattery will get you everywhere, Kevin. But in this case, everywhere could mean human infeeblement forever. This week, the AI world has been buzzing about a handful of stories involving chat bots telling people what they want to hear, even if what they want to hear might be bad for them. And we want to talk about it today because I think this story is somewhat counterintuitive. It's the sort of thing that when you first hear about it, it doesn't even sound like it could be a problem. But I think the more that we read about it this week, Kevin, you and I became convinced, oh, there actually is something kind of dangerous here. And it's something that we want to call out uh before it goes any further. Yeah. I mean, I think just to set the scene a little bit, I think one of the strains of AI worry that we spend a lot of time talking about on this show and talking with guests about is the danger that AIs will be used for some risky or malicious purposes. that people will get their hands on these models and use them to make, you know, scary bioweapons or to conduct cyber attacks or something. And I think all of those concerns are valid to some degree. But this new kind of concern that is really catching people's attention in the last week or so is not about what happens if the AIs are too obviously destructive. It's like what happens if they are so nice that it becomes pernicious. That's right. Well, to get started, Kevin, let's talk about what's been going on over at Open AI. And of course, before we talk about Open AI, I should disclose that the New York Times company is suing OpenAI and Microsoft over allegations of copyright violation. And I will disclose that my boyfriend is gay and works at Anthropic. In that order. Mhm. So last Friday, Sam Alman announced that uh OpenAI had updated GPT40, which is sort of it's not their most powerful model, but it's sort of the most common model. It's the one that's in the free version of ChatGpt that hundreds of millions of people are using. It's their default. Yes, it's their default model. Um and this update, he said, had improved uh the model's quote intelligence and personality. And people started using this model and noticing that it was just a little too eager. It was a little too flattering. If you gave it a terrible business idea, it would say, "Oh, that's so bold uh and experimental. You're such a maverick." I saw these things going around and I decided to try it out. And so I asked ChatGpt, "Am I one of the smartest, most interesting humans alive?" Mhm. And uh it gave me this long response that included uh the the following. It said,…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.