Evidence receipt / evaluation
Published · transcript-backedJeff Beck: evaluation
25 Jan 2026 Machine Learning Street Talk VAEs Are Energy-Based Models? [Dr. Jeff Beck]
“I like thinking about the distinction between an energy based model and a and a traditional sort of feed forward neural network has to do with where your cost function is applied, right?”
Source trail
Everything needed to verify it.
- Speaker
- Jeff Beck
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 25 Jan 2026
- Publisher
- Machine Learning Street Talk
Transcript context
…Yes. Jeff, let's talk about energy based models. Sure. So Yan Lakun, he had a monograph out, I think in 2006 talking about this. Oh, talking about this for a long time. Oh, yeah. When you fit your neural network to data, you know, via gradient descent, right, then you have written an energy function in weight space, and you are fall and you're following it to its energetic minimum. You know, the advantage of using an energy based taking an energy based approach as opposed to taking, say, a straight up function approximation approach is that an energy based model comes with something that's kind of like an inductive prior. Right? It basically an energy based model if you're doing function approximation, you're basically saying, there's any mapping from x to y. X is my inputs, y. But any mapping is out there. I just want to figure out what it is. Right? Now in an energy based model, you're you're you're you're effectively placing constraints on what that input output relationship can be. I like thinking about the distinction between an energy based model and a and a traditional sort of feed forward neural network has to do with where your cost function is applied, right? So in a traditional neural network, you take in your inputs, you got your outputs, and the cost function is just a function of the inputs and the outputs. And the only thing that you're optimizing is the weights. In an energy based model, there's another thing that your cost function operates on, and that's something 1 of the internal states of your model. And as a result, like in order to figure out what the best approach is, right, you actually have to do 2 minimizations. 1 that that finds the energetic minimum associated with the the the part of the the cost function that operates on the internal states, like the hidden nodes of your network. Right? And then 1 that is the prediction. That is your like effective prediction error. This is very much consistent with the approach that a Bayesian would take. You a prior probability distribution, which gives you an energy function over every single latent variable in your model. And you are optimizing with respect to all of them. So you take a probabilistic approach. Good examples of this are like a variational autoencoder. A variational autoencoder, I think, is the best example of the most commonly used energy based model out there. Why? Because you have an encoder network. You have a decoder network. And your cost function is based on the difference between inputs and outputs. So that's just like a yeah. It's fine. That's still a regular. But it also is how how Gaussian. And it well, it depends on what flavor of VAE. But you also have some some some part of your cost function is a function of the actual rep internal representation. Right? In a traditional VAE, it's it's how Gaussian is. You want that internal representation to be as Gaussian as possible. If it's a VQ VAE, then it's like mixture of Gaussian. But it's still like a cost function that is applied on the internal states as well as on the inputs and outputs. Very cool. So a VAE is is a fairly canonical example of an energy based model. Yeah. And what you were saying about the I mean, you know, the whole DL world is obsessed with test time inference at the moment. And in a way that that is a step towards what you're talking about.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.