Evidence receipt / preference
Published · transcript-backedPavan Kumar Reddy: preference
30 Mar 2026 Latent Space Mistral: Voxtral TTS, Forge, Leanstral, & what's next for Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
“I think at least personally prefer the operations, which are the simplest, and so we try to see, can we just add audio as just another head to our regular transformer decode model because that kind of makes it easier for eventual end-to-end modeling of audio text native modeling.”
Source trail
Everything needed to verify it.
- Speaker
- Pavan Kumar Reddy
- Attribution
- Verified speaker
- Claim type
- preference
- Recorded
- 30 Mar 2026
- Publisher
- Latent Space
Transcript context
…What are some of the ways to look at it? There are ways where you can do diffusion for audio generation, but if you want like real time generation, that’s a big thing with the approach I’m assuming that you took. Yeah. And also like how do you go about evaluating different axes of what you care about, yeah, good point. I think we so you can do just flow matching diffusion for the whole audio. We didn’t even go down that path because one of the main applications is voice agents and we want real time streaming, and that’s the use case. That’s not the only use case, but that’s one of the primary use cases we want to get to. So we picked the auto aggressive approach for that. And within the auto aggressive space, again, you can do chunk by chunk or you can do so we picked the. I think at least personally prefer the operations, which are the simplest, and so we try to see, can we just add audio as just another head to our regular transformer decode model because that kind of makes it easier for eventual end-to-end modeling of audio text native modeling. Yeah. And it works pretty well. So I guess we went with that and we tried a little bit, but the flow matching head itself, like we had a discreet. Diffusion kind of approach, which also works well, but the flow matching work better. I was just curious about how you also think about this overall direction of research. Do you basically, when you work with the audio team, do you set some high level parameters and then let them explore whatever, or how does it work between you guys?…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.