Evidence receipt / observation
Published · transcript-backedLuke Darlow: observation
23 Nov 2025 Machine Learning Street Talk He Co-Invented the Transformer. Now: Continuous Thought Machines - Llion Jones and Luke Darlow [Sakana AI]
“I think in terms of stability, what's what we found is kind of fun, this was a sentiment that we had throughout the the experiments that we ran with this paper, was it tended no matter what we tried it on, it it just kind of worked with all spreads of hyperparameters.”
Source trail
Everything needed to verify it.
- Speaker
- Luke Darlow
- Attribution
- Verified speaker
- Claim type
- observation
- Recorded
- 23 Nov 2025
- Publisher
- Machine Learning Street Talk
Transcript context
…So on this on the scaling, I think the time complexity is quadratic in respect of the dimension of the synchronization matrix. Right? And in your paper, were talking about subsampling to improve the performance. But how how did that affect the the stability and the you know, like, there any things that that cost you doing that? Yeah. It's a neat question. I think in terms of stability, what's what we found is kind of fun, this was a sentiment that we had throughout the the experiments that we ran with this paper, was it tended no matter what we tried it on, it it just kind of worked with all spreads of hyperparameters. And this the problems that you have with backprop through time, typically with recurrence models like RNNs and LSTMs, it's a challenge. You run for many internal ticks with the RNNs or the LSTMs and the learning seems to break down. But the fact that we use synchronization in some sense touches all of the neurons through all of the time. So it really helps with gradient propagation. A nice interesting point that's maybe a bit oblique to what you asked about synchronization is we have a system of d neurons and like I said earlier, there are d over 2 squared possible combinations. This essentially means that our underlying state or underlying representation to the system is quite a lot larger than what you would get with just taking those de neurons. And as to what that means in terms of downstream computation and performance and the things that we can do with this is what we're actively exploring right now. You guys used an exponential decay rate?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.