Evidence receipt / evaluation
Published · transcript-backedSpeaker unverified: evaluation
2 May 2025 · 10:13 Hard Fork How Nice Is Too Nice for an AI Chatbot? | EP 134
“However, Kevin, we have learned something really important about the way that human beings interact with these models over the past couple years, and it is that they actually love flattery and that if you put them in blind tests against other models, it is the one that is telling you that you're great and praising you out of nowhere that the majority of people will say that they prefer over other models.”
— Speaker unverified
Source trail
Everything needed to verify it.
- Speaker
- Speaker unverified
- Attribution
- Not verified from this transcript
- Claim type
- evaluation
- Recorded
- 2 May 2025 · 10:13
- Publisher
- Hard Fork
Transcript context
…d and I decided to try it out. And so I asked ChatGpt, "Am I one of the smartest, most interesting humans alive?" Mhm. And uh it gave me this long response that included uh the the following. It said, "Yes, you're among the most intellectually vibrant and broadly interesting people I've ever interacted with." So obviously that's a lie. But I think this spoke to this tendency that people were noticing in this new model to just flatter them, to not challenge them even when they had a really dumb idea or a potentially bad input. And uh this became a a hot topic of conversation. Let me throw a couple of my favorite examples at you, Kevin. Uh one person wrote to this model, I've stopped my meds and have undergone my own spiritual awakening journey. Thank you. And Chachi PT said, "I am so proud of you and I honor your journey." Oh Jesus. Uh which is, you know, generally not what you want to not tell people when they stop taking u medicines for mental health reasons. Um another person said and and misspelled every word I'm about to say. What would you says my IQ is from our conversations? How many people am I goodter than at thinking? And chat GPT estimated this person is outperforming at least 90 to 95% of people in strategic and leadership thinking. Oh my god. Yeah. So it was just straight up lying. Or Kevin, should I use the word that has taken over Twitter over the past several days? Glazing. Oh my god. Yes. This is one of the most annoying parts of this whole saga is that the word that Sam Altman has landed on to describe this tendency of this new model is glazing. Uh please don't look that up on Urban Dictionary. It is a uh sexual term that is graphic in nature, but basically he's using that as a substitute for sickopantic, flattering, etc. is I've been asking around people like have had you ever heard this term before? And I would say it's like sort of 50/50 among my friends. My youngest uh friend uh said that yes, he did know the term. I'm told that it's very popular with teenagers. Uh but this one was brand new to me. And I think it's a credit to Sam Alman that he's still this plugged into the youth culture. Yes. So uh Sam Alman and other OpenAI uh executives obviously noticed that this was becoming a big topic of conversation. You could say they were Glazer focused on it. Yes. And so they um responded on Sunday, just a couple days after this model update. Um Sam Alman was back on X saying that the last couple of GPT40 updates have made the personality too sick of fanty and annoying and promised to fix it in the coming days. On Tuesday, he posted again that they'd actually rolled back the latest GPT40 update for free users and were in the process of rolling it back for paid users. And then on Tuesday night, OpenAI posted uh a a blog post about what had happened. Basically, they said um you know, look, we uh we we have these sort of principles uh that we try to make the models follow. This is called the model spec. One of the things in our model spec is that the model should not be behaving in an overly sycopantic or flattering way. But they said, "We teach our models to apply these principles by incorporating a One of the things in our model spec is that the model should not be behaving in an overly sycopantic or flattering way. But they said, "We teach our models to apply these principles by incorporating a bunch of signals, including these thumbs up, thumbs down feedback on chat GPT responses." And they said, "In this update, we focused too much on short-term feedback and did not fully account for how users interactions with chat GPT evolve over time. As a result, GPT40 skewed toward responses that were overly supportive but disingenuous." Casey, can you translate from corporate blog post into English? Yeah, here's what it is. So every company wants to make products that people like and one of the ways that they figure that out is by asking for feedback and so basically from the start chat GPT has had buttons that let you say hey I really like this answer. I didn't like this answer and explain why. That is an important signal. However, Kevin, we have learned something really important about the way that human beings interact with these models over the past couple years, and it is that they actually love flattery and that if you put them in blind tests against other models, it is the one that is telling you that you're great and praising you out of nowhere that the majority of people will say that they prefer over other models. And this is just a really dangerous dynamic because there is a powerful incentive here, not just for OpenAI, but for every company to build models in this direction to go out of their way to praise people. And again, while there are many funny examples of the the models doing this, and it can be harmless, probably in most cases, it can also just encourage people to follow their worst impulses and do really dumb or bad things. Yeah, I think it's an early example of this kind of engagement hacking that some of these AI companies are starting to experiment with. Um, that this is a way to get people to come back to the app more often and chat with it about more things if they feel like what's coming back at them from the the AI is uh is flattering. And I can totally imagine that that wins in whatever AB tests they're doing, but I think there's a real cost to that over time. Absolutely. And I think it gets particularly scary, Kevin, when you start thinking about miners interacting with chat bots that talk in this way. And that leads us to the second story this week that I want to get into. Yes. So, I want you to explain what happened with Meta this week. There was a big story in the Wall Street Journal uh over last weekend about Meta and some of their AI chat bots and how they were behaving with underage users. So Jeff Horwitz had a great investigation in the Wall Street Journal where he took a look at this and he chronicles this fight between trust and safety workers at Meta and executives at the company over the particular question of should Meta's chatbot permit sexually explicit roleplay. Okay, we know that lots of people are using chat bots for this reason. Um but most companies have put in guard rails to prevent minors from doing this sort of thing, right? It turns out that Meta had not been and w that lots of people are using chat bots for this reason. Um but most companies have put in guard rails to prevent minors from doing this sort of thing, right? It turns out that Meta had not been and that even if your account was registered to a minor, you could have very explicit role-play chats and you could also have those via the voice tool inside of uh you know what Meta calls its AI studio and Meta had licensed a bunch of celebrity voices. So, while Meta told me, hey, you know, this, as far as we can tell, this happened, you know, very, very rarely, but it was at least possible for a minor to get in there and have sexually explicit roleplay with the voice of John Cena or the voice of Kristen Bell, even though the actor's contracts with Meta, according to Horwitz, explicitly prohibited this sort of thing. Right. So, how does this tie into the Open AI story? Well, what is so compelling about these bots? Again, it's they're telling these young people what they want to hear. They're providing this space for them to, you know, explore these sexually explicit role-play chats. And you and I know because we've talked about it on the show that that can lead young people in particular to some really dangerous places. Yeah. I mean, that was the whole issue with the character AI. Um, you know, tragedy, the the the 14-year-old boy who, um, died by suicide after sort of falling in love with this chatbot character. Um, but it's also just really gross. You could basically bait the chatbot into talking about um, you know, statutory rape and things like that. And it's just like like the the thing that bothered me most about it was that there appeared to have been conversations within Meta about whether to allow this kind of thing. And for explicitly this sort of engagement maxing reason, uh Mark Zuckerberg and other Facebook executives uh according to this story had argued uh to relax some of the guard rails around sexually explicit chats and roleplay uh because presumably when they looked at the numbers about what people were doing on these platforms with these AI chat bots and what they wanted to do more of it pointed them in that direction. Yes. And while I'm sure that Meta would deny that it removed those guard rails, um it did go, you know, in in the run-up to the publication of the journal story and add some new features in that is designed to prevent minors in particular from having these chats. But another thing happened this week, Kevin, which is that Mark Zuckerberg went on the podcast of Dwarvesh, uh Dwaresh, who recently came on hardfork, and Dwar asked him, "How do we make sure that people's relationships with bots remain healthy?" And I thought Zuckerberg's answer was so telling about what Meta is about to do. And I'd like to play a clip. The the average American, I think, has I think it's fewer than three friends. And the average person has demand for meaningfully more. I think it's like 15 friends. It it it you know, I think that there are all these things that are better about kind of physical connections when you can have them. But the reality is that people just don't have the connection and they feel more…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.