Evidence receipt / evaluation
Published · transcript-backedPavan Kumar Reddy: evaluation
30 Mar 2026 Latent Space Mistral: Voxtral TTS, Forge, Leanstral, & what's next for Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
“That doesn’t work. At least that doesn’t work that well because audio has more entropy.”
Source trail
Everything needed to verify it.
- Speaker
- Pavan Kumar Reddy
- Attribution
- Verified speaker
- Claim type
- evaluation
- Recorded
- 30 Mar 2026
- Publisher
- Latent Space
Transcript context
…Yeah. But if you have K tokens, then the name thing would be to predict all of them in paddle. That doesn’t work. At least that doesn’t work that well because audio has more entropy. And the, one of the techniques people use is this depth transformer where you you almost have a small transformer, or it can be L-S-T-M-R in as well, but people use transformers and you predict the K tokens in auto aggressive fashion in that. So you have two auto reive things going on. So the thing we did differently is in, instead of having this auto aggressive K step prediction, we have a flow matching model. Instead of modeling this as a discrete token set we trained the codec to be both discrete and continuous to have this flexibility. So we did try the discrete stuff too, and which it works well, but the continuous stuff works just better. So yeah, we took this flow matching, so the, it’s a flow matching head, which takes the latent from the main transformer and like kind in fusion, it’s denoising, but in this flow matching itself, velocity estimate. So you go from this noise t all the way to there. Audio latent, which corresponds to the 80 millisecond audio and then, which is sent through the work order to get back the 80 millisecond audio frame. Yeah. Is this the first application of flow matching in audio? Because usually I come across this in the image.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.