High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Dwarkesh Patel: belief

29 Apr 2026 Dwarkesh Podcast Reiner Pope – The math behind how LLMs are trained and served

“Ok, so I was asking you, have neural networks actually been used for cryptography? And we realized it may be better to just do this on the blackboard.”

— Dwarkesh Patel

Source trail

Everything needed to verify it.

Speaker
Dwarkesh Patel
Attribution
Verified speaker
Claim type
belief
Recorded
29 Apr 2026
Publisher
Dwarkesh Podcast

Transcript context

…This is an adversarial context? This is actually a place where you get exactly the avalanche property that ciphers have as well. Adversarial attacks on image classification models are about finding a very small perturbation of the image that totally changes the classification, totally changes the output. That is the common case in ciphers, whereas that’s the undesired case in neural nets. Ok, so I was asking you, have neural networks actually been used for cryptography? And we realized it may be better to just do this on the blackboard. Are they actually being used for cryptography? Using neural nets for cryptography… In general, creating a new cipher is a very dangerous proposition. Almost all of them are broken. 99% of them are broken, so it’s probably a bad place to start. But the other direction has been, in at least one very clear case, quite productive. There’s a construction that exists in ciphers and then was imported into neural nets called a Feistel cipher, or Feistel network. The idea is that you may have some function f which is not invertible, but you like the function because it does interesting things, like it does an MLP, for example. Or it mixes it in an interesting way. You’d like to build something out of this that is invertible. The construction we’re going to make is going to be a two-input function rather than a one-input function. We’re going to apply f(x). We need to actually remember what x was, so we’re going to stick x over here so that we can work backwards, and then we also can’t drop y. We’re going to remember y, and we’re going to add them together to form this tuple. The way to invert this, if you think I have this output and I want to recover x and y, I can easily recover x. That’s right there, I just read it off. To recover y, if this thing was called z, I can recover y by z minus f(x), because I’ve already recovered x. That means this construction is invertible. This was used in ciphers a ton and still is used. It’s one of the main mechanisms of constructing ciphers. Often you want ciphers to be invertible, especially the layers of ciphers, because that has better cryptographic properties. This has actually been ported over into neural nets. There’s a 2017 paper called RevNets, reversible networks. What it does is make the entire network invertible. You can apply it to any network, like a transformer network. I do a forwards pass, but then I can run the entire pass backwards as well. The whole neural network is invertible with exactly this construction. This paper applied it to some layer, like a transformer layer, for example. We’ve got this function f, which is our transformer layer. Normally we would have just an input and then a residual connection coming out, and it gets added over here. Now, the variation of this is going to be we’ve got two inputs, x and y. x goes through the function, gets added to y, and then this becomes the new x, output x. Then this x becomes the output y. Really what this is doing, if you think of two layers back, is the thing you mentioned before. It’s doing the residual connection from two layers back. This y came from the previous layer and was the residual connection there. Because of this construction, the whole thing is invertible. Why do I care? What does invertible matter for? The big thing that it can be interesting for is training. If I think of a forward pass of training… Let’s say I have four layers and I run them in zero, one, two, three order. I have to write all of the activations to HBM.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence