High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / preference

Published · transcript-backed

Jeff Beck: preference

25 Jan 2026 Machine Learning Street Talk VAEs Are Energy-Based Models? [Dr. Jeff Beck]

“In fact, time someone hands me a new neural data set, the first thing I do I'm not ashamed to admit, I run PCA on and pass it through a VAE, and then sort of take a look, right?”

— Jeff Beck

Source trail

Everything needed to verify it.

Speaker
Jeff Beck
Attribution
Verified speaker
Claim type
preference
Recorded
25 Jan 2026
Publisher
Machine Learning Street Talk

Transcript context

…Lakun is a big fan of this non contrastive thing where you work in the latent space. There are many different algorithms that do this. Had a whole load of shows all about non contrastive learning. There's things like VCREG and BYOL and Barlow Twins. And there's an entire thread of research all around that. And in many different ways, what they're trying to do is avoid this motor collapse problem that you're talking about. And they use different forms of regularization. There's an old school way of accomplishing the same thing. And that is to do all of your, it's called pre processing, right? And this is something that a lot of people do. You take your data, in fact we do this all the time with like vision language models, Right? So we wanna do, we wanna use an LLM and we wanna predict images. So what do we do? Well the first thing we have to do is tokenize the image. Right? And so what do we do? We run a VA that we do the preprocessing. And we do it by the preprocessing step is completely independent, right, from the actual algorithm that's gonna be tasked with solving the problem of interest. And that's not something that we necessarily have to stick with, right? It would be very nice if there was a way of like, again, jointly, we're getting right back to JEP again. What we'd like to do is we'd like to choose our preprocessing algorithm in a manner that that, you know, not a priori, not do it first. We'd like to choose the preprocessor that works the best in in this space. Yep. And I think that that's the ultimate motivation for a lot of this work is that it's like, what's the right embedding? 1 of my favorite tricks of course, I pre process the VAs all the time. In fact, time someone hands me a new neural data set, the first thing I do I'm not ashamed to admit, I run PCA on and pass it through a VAE, and then sort of take a look, right? It's the first thing you do with your data, because it gives you a good idea of what the signal to noise ratio is in the data set itself. Yes. And then I yeah. And then what do I do? I subsequently do most of my analysis, right, in that discovered embedding space. And there's I I I I don't see a huge problem with that from a purely pragmatic perspective, But it's certainly cleaner, right, to have a single algorithm and approach and not just be stringing these sort of things together in an ad hoc way. There's, you know, when doing PCA, PCA is a really great example of this. There's a failure mode for principal component analysis, which is actually really common in neural data. Because principal component analysis basically says, well, where's the most variability? Okay, I'm worried about that. And then all the stuff that's not varying very much, I'm just going to throw it away. Just like dimensions in which there's low variability are not important. Well, turns out that in neural data, the dimensions in which there's very little variability are some of the most important dimensions. Yes. And so preprocessing with PCA runs a risk of throwing out the most valuable information in your data set. Yes. And so there's a lot of wisdom in jointly fitting your preprocessing model as well as your inference and prediction model. I mean, this subject of not throwing things away, JEPR and non contrastive learning, it's part of this bigger field of self supervised learning. And we want to learn representations that maintain fidelity and richness. And Lakun's hypothesis is that when you do something like supervised learning with some particular downstream task in mind, the neural network gets wise. And what it does is it kind of discards all of the long tail stuff that aren't relevant for that particular task. So when you train these models, you're trying to do is sort of maintain enough ambiguity so that it compresses the information, but it also maintains enough fidelity to work broadly for different things.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence