High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / preference

Published · transcript-backed

Yann LeCun: preference

7 Mar 2024 Lex Fridman Podcast #416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI

“There is another set of methods which are non-contrastive, and I prefer those, and those non-contrastive methods basically say, the energy function needs to have low energy on pairs of XYs that are compatible that come from your training set.”

— Yann LeCun

Source trail

Everything needed to verify it.

Speaker
Yann LeCun
Attribution
Verified speaker
Claim type
preference
Recorded
7 Mar 2024
Publisher
Lex Fridman Podcast

Transcript context

…You’re talking about the ability to think deeply or to reason deeply, how do you know what is an answer that’s better or worse based on deep reasoning? Then we are asking the question of, conceptually, how do you train an energy based model? Energy based model is a function with a scaler output, just a number, you give it two inputs, X and Y, and it tells you whether Y is compatible with X or not. X, you observe, let’s say it’s a prompt, an image, a video, whatever, and Y is a proposal for an answer, a continuation of video, whatever and it tells you whether Y is compatible with X. And the way it tells you that Y is compatible with X is that the output of that function would be zero if Y is compatible with X and would be a positive number, non-zero, if Y is not compatible with X. How do you train a system like this at a completely general level, is you show it pairs of X and Ys that are compatible, a question and the corresponding answer, and you train the parameters of the big neural net inside to produce zero. Now that doesn’t completely work because the system might decide, well, I’m just going to say zero for everything, so now you have to have a process to make sure that for a wrong Y, the energy would be larger than zero. And there you have two options, one is contrastive method, so contrastive method is, you show an X and a bad Y and you tell the system, well, give a high energy to this, push up the energy, change the weights in the neural net that confuse the energy so that it goes up. So that’s contrasting methods. The problem with this is, if the space of Y is large, the number of such contrasting samples are going to have to show is gigantic. But people do this, they do this when you train a system with RLHF, basically what you’re training is what’s called a reward model, which is basically an objective function that tells you whether an answer is good or bad, and that’s basically exactly what this is. So we already do this to some extent, we’re just not using it for inference, we’re just using it for training. There is another set of methods which are non-contrastive, and I prefer those, and those non-contrastive methods basically say, the energy function needs to have low energy on pairs of XYs that are compatible that come from your training set. How do you make sure that the energy is going to be higher everywhere else? And the way you do this is by having a regularizer, a criterion, a term in your cost function that basically minimizes the volume of space that can take low energy. And the precise way to do this is all kinds of different specific ways to do this depending on the architecture, but that’s the basic principle. So that if you push down the energy function for particular regions in the XY space, it will automatically go up in other places because there’s only a limited volume of space that can take low energy by the construction of the system or by the regularizing function. We’ve been talking very generally, but what is a good X and a good Y? What is a good representation of X and Y? Because we’ve been talking about language and if you just take language directly that presumably is not good, so there has to be some kind of abstract representation of ideas.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence