High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Edward Gibson: prediction

17 Apr 2024 Lex Fridman Podcast #426 – Edward Gibson: Human Language, Psycholinguistics, Syntax, Grammar & LLMs

“Should I trade for door number two, even though I know the good prize in door number one? And then the large language model say, “Yes, you should trade,” because it just goes through the forms that it’s seen before so many times on these cases where yes, you should trade because your odds have shifted from one in three now to two out of three to being that thing.”

— Edward Gibson

Source trail

Everything needed to verify it.

Speaker
Edward Gibson
Attribution
Verified speaker
Claim type
prediction
Recorded
17 Apr 2024
Publisher
Lex Fridman Podcast

Transcript context

…Do you think that large language models understand language? Are they mimicking language? I guess the deeper question there is, are they just understanding the surface form or do they understand something deeper about the meaning that then generates the form? I would argue they’re doing the form. They’re doing the form, they’re doing it really, really well. And are they doing the meaning? No, probably not. There’s lots of these examples from various groups showing that they can be tricked in all kinds of ways. They really don’t understand the meaning of what’s going on. And so there’s a lot of examples that he and other groups have given which show they don’t really understand what’s going on. So the Monty Hall problem is this silly problem. Let’s Make a Deal is this old game show, and there’s three doors and there’s a prize behind one, and there’s some junk prizes behind the other two and you’re trying to select one. And Monty, he knows where the target item is. The good thing, he knows everything is back there, and he gives you a choice. You choose one of the three and then he opens one of the doors and it’s some junk prize. And then the question is, should you trade to get the other one? And the answer is, yes, you should trade because he knew which ones you could turn around, and so now the odds are two-thirds. And then if you just change that a little bit to the large language model, the large language model has seen that explanation so many times. If you change the story, it’s a little bit, but you make it sound like it’s the Monty Hall problem, but it’s not. You just say, “Oh, there’s three doors and one behind them is a good prize and there’s two bad doors. I happen to know it’s behind door number one. The good prize, the car is behind door number one, so I’m going to choose door number one.” Monty Hall opens door number three and shows me nothing there. Should I trade for door number two, even though I know the good prize in door number one? And then the large language model say, “Yes, you should trade,” because it just goes through the forms that it’s seen before so many times on these cases where yes, you should trade because your odds have shifted from one in three now to two out of three to being that thing. It doesn’t have any way to remember that actually, you have 100% probability behind that door number one. You know that. That’s not part of the scheme that it’s seen hundreds and hundreds of times before. And so even if you try to explain to it that it’s wrong, that they can’t do that, it’ll just keep giving you back the problem. But it’s also possible the larger language model would be aware of the fact that there’s sometimes over-representation of a particular kind of formulation, and it’s easy to get tricked by that. And so you could see if they get larger and larger, models be a little bit more skeptical, so you see over-representation. It just feels like training on form can go really far in terms of being able to generate things that look like the thing understands deeply the underlying world model, of the kind of mathematical world, physical world, psychological world that would generate these kinds of sentences. It just feels like you’re creeping close to the meaning part, easily fooled, all this kind of stuff, but that’s humans too. So it just seems really impressive how often it seems like it understands concepts.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence