High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Pavan Kumar Reddy: prediction

30 Mar 2026 Latent Space Mistral: Voxtral TTS, Forge, Leanstral, & what's next for Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample

“I think I could have a lot of details. But me I think the [00:27:00] summary of it, actually, some of the considerations in this paper were, because we started with the wipa encoder as the starting point, and now we have in-house encoders, like the bigger time model, for instance, which we released in January.”

— Pavan Kumar Reddy

Source trail

Everything needed to verify it.

Speaker
Pavan Kumar Reddy
Attribution
Verified speaker
Claim type
prediction
Recorded
30 Mar 2026
Publisher
Latent Space

Transcript context

…My, my basic example is you don’t want to call to customer services and have the same exact voice. It’s just, it’s gonna be weird. But also on the technical side of this, so there’s like a few things in TRO that I thought were pretty interesting. He’s a big fan of this paper. Oh, he said very good paper. He said this is the best SR paper he’s ever read. Yeah. I’ve hyped up this voice paper enough. We covered it. Somewhere, but a big thing. So Whisper is known for 32nd generation a 32nd processing. You extended this to 40 minutes. There was a lot of good detail in the paper about how this was done. Even little niches of how the padding is. So it’s very much needed. You need to have that padding in there, the synthetic data generation around this. I’m wondering if you can share the same about the new speech to text, right? Text to speech. So how do you. How do you generate long form, coherent? How do you generate, how do you do that? And then any gems? Is there gonna be a paper? Yeah. Yeah. They would be a technical report. Okay. Yeah. I think I could have a lot of details. But me I think the [00:27:00] summary of it, actually, some of the considerations in this paper were, because we started with the wipa encoder as the starting point, and now we have in-house encoders, like the bigger time model, for instance, which we released in January. Also release a technical report for that real time model as well, which is this dual stream architecture. It’s an interesting architecture. You should check it out. And there we have a causal encoder and I don’t think there’s any strong, multilingual causal encoder out in the community. So we thought it’s a good contribution. So that’s one nice encoder there. Other people want to adapt. That’s a good end code. And we train it from scratch. I think her. Post stack is now mature enough that we are able to train super strong ENC codes. And some of these considerations, like spatting and stuff, is a function of the Whisper ENC code. And now that we train encoders, inhouse the design concentrations are different. And for the question on text to speech, I think that’s also leans onto the original auto aggressive decoder backbone. I think, it says very, almost identical considerations. I think the long context in it’s not even long con, [00:28:00] so the model processes audio at 12.5 herds, so one second maps to like 12.5 tokens. So I think one minute is like 7.8 tokens. You can get like up to 10 minutes in eight K context window and get half an hour and 30 K context window. So that’s and 30 2K context is something that’s we are very comfortable training on. We can extend it even much longer. 1 48 K. Okay. You can naturally see how it can extend to even our long generations. Yeah. We need the. Like data recipe and the whole algorithm to work coherently enough through such long context. But the techniques are some way very similar to the text, long context modeling. And the key differences, it’s just doing flow matching order regressively instead of a text open prediction. Okay. I think that was most, most of the sort of voice questions that we had. But…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence