Evidence receipt / prediction
Published · transcript-backedMati Staniszewski: prediction
14 Apr 2026 Cheeky Pint The world of voice AI, with Mati Staniszewski of ElevenLabs
“In speech-to-speech, as you think about maybe more of a companion version of the applications, that's where that will flourish because maybe the hallucinations aren't as important, but the latency is a little bit more, and maybe hallucinations are even a feature.”
Source trail
Everything needed to verify it.
- Speaker
- Mati Staniszewski
- Attribution
- Verified speaker
- Claim type
- prediction
- Recorded
- 14 Apr 2026
- Publisher
- Cheeky Pint
Transcript context
…I'm sorry, a cascaded approach is? Is the speech-to-text... Going through the text layer. As we work with a lot of the businesses and enterprises, they will need that visibility into what happens. They will want to execute certain tasks on top of that. They want a good visibility into each of the steps and great accuracy of all the models. But beyond that, they can abstract away what's the LLM layer, what's the intelligence layer, the integrations are easier in that system. That's where we are betting a lot of the research work of how you can make that great, and we think we can make that great. In speech-to-speech, as you think about maybe more of a companion version of the applications, that's where that will flourish because maybe the hallucinations aren't as important, but the latency is a little bit more, and maybe hallucinations are even a feature. Maybe in the future-future, just to finish that part, you will have some version of combination of the models. That for low complexity, easy models, you will have speech-to-speech, and for higher complexity we will have to cascade it. I was going to ask about this. You know the way there is research on how the invention of writing changed human brains and just changed the neural pathways in ways beyond the actual written language. Do you observe that speech-to-speech models think differently than cascaded models? It sounds like they're dumber.…
Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.