High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / belief

Published · transcript-backed

Pavan Kumar Reddy: belief

30 Mar 2026 Latent Space Mistral: Voxtral TTS, Forge, Leanstral, & what's next for Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample

“Each enterprise want a customized, specialized something which is representative both their brand and also their, I guess safety considerations and the use case I think the kind of thing that you would deploy as a empathetic assistant in the context of a healthcare domain would be very different from the kind of thing that would be in a customer support bot and would be different from like more conversational aspects.”

— Pavan Kumar Reddy

Source trail

Everything needed to verify it.

Speaker
Pavan Kumar Reddy
Attribution
Verified speaker
Claim type
belief
Recorded
30 Mar 2026
Publisher
Latent Space

Transcript context

…This one I was gonna ask you, we never talked about cloning voice clothing here. How important is it, right? Like I can clone a famous person’s voice. Okay. But the main use case would be like for enterprise personalization, like enterprises need like a lot of customization. You don’t want the same. Voice for all the enterprises. Each enterprise want a customized, specialized something which is representative both their brand and also their, I guess safety considerations and the use case I think the kind of thing that you would deploy as a empathetic assistant in the context of a healthcare domain would be very different from the kind of thing that would be in a customer support bot and would be different from like more conversational aspects. I think those are the. Customizations you would expect from enterprise. And that’s the main use case, at least from our side. My, my basic example is you don’t want to call to customer services and have the same exact voice. It’s just, it’s gonna be weird. But also on the technical side of this, so there’s like a few things in TRO that I thought were pretty interesting. He’s a big fan of this paper. Oh, he said very good paper. He said this is the best SR paper he’s ever read. Yeah. I’ve hyped up this voice paper enough. We covered it. Somewhere, but a big thing. So Whisper is known for 32nd generation a 32nd processing. You extended this to 40 minutes. There was a lot of good detail in the paper about how this was done. Even little niches of how the padding is. So it’s very much needed. You need to have that padding in there, the synthetic data generation around this. I’m wondering if you can share the same about the new speech to text, right? Text to speech. So how do you. How do you generate long form, coherent? How do you generate, how do you do that? And then any gems? Is there gonna be a paper?…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence