High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Speaker unverified: belief

2 Aug 2024 Lex Fridman Podcast #438 – Elon Musk: Neuralink and the Future of Humanity

“The direction I think of further improvement is primarily going to be in the dataset side of how do you construct the optimal labels for the model.”

— Speaker unverified

Source trail

Everything needed to verify it.

Speaker
Speaker unverified
Attribution
Not verified from this transcript
Claim type
belief
Recorded
2 Aug 2024
Publisher
Lex Fridman Podcast

Transcript context

…You mentioned NeuroDecodeR. How much machine learning is in the decoder, how much magic, how much science, how much art? How difficult is it to come up with a decoder that figures out what these sequence of spikes mean? Yeah, good question. There’s a couple of different ways to answer this, so maybe I’ll zoom out briefly first and then I’ll go down one of the rabbit holes. So the zoomed out view is that building the decoder is really the process of building the dataset plus compiling it into the weights, and each of those steps is important. The direction I think of further improvement is primarily going to be in the dataset side of how do you construct the optimal labels for the model. But there’s an entirely separate challenge of then how do you compile the best model? And so I’ll go briefly down the second rabbit hole. One of the main challenges with designing the optimal model for BCI is that offline metrics don’t necessarily correspond to online metrics. It’s fundamentally a control problem. The user is trying to control something on the screen and the exact user experience of how you output the intention impacts their ability to control. So for example, if you just look at validation loss as predicted by your model, there can be multiple ways to achieve the same validation loss. u output the intention impacts their ability to control. So for example, if you just look at validation loss as predicted by your model, there can be multiple ways to achieve the same validation loss. Not all of them are equally controllable by the end user. And so it might be as simple as saying, oh, you could just add auxiliary loss terms that help you capture the thing that actually matters. But this is a very complex nuanced process. So how you turn the labels into the model is more of a nuanced process than just a standard supervised learning problem. One very fascinating anecdote here, we’ve tried many different neural network architectures that translate brain data to velocity outputs, for example. And one example that’s stuck in my brain from a couple of years ago now is at one point, we were using just fully-connected networks to decode the brain activity. We tried A-B test where we were measuring the relative performance in online control sessions of one deconvolution over the input signal. So if you imagine per channel you have a sliding window that’s producing some convolved feature, for each of those input sequences for every single channel simultaneously, you can actually get better validation metrics, meaning you’re fitting the data better and it’s generalizing better in offline data if you use this convolutional architecture. You’re reducing parameters. It’s a standard procedure when you’re dealing with time series data. Now it turns out that when using that model online, the controllability was worse, was far worse, even though the offline metrics were better, and there can be many ways to interpret that. But what that taught me at least was that, hey, it’s at least the case right now that if you were to just throw a bunch of compute at this problem and you were trying to hyperparameter optimize or let some GPT model hard code or come up with or invent many different solutions, if you were just optimizing for loss, it would not be sufficient, which means that there’s still some inherent modeling gap here. There’s still some artistry left to be uncovered here of how to get your model to scale with more compute, and that may be fundamentally a labeling problem, but there may be other components to this as well.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence