High Signal Podcasts Evidence ledger
Method
Browse
← Back to evidence

Evidence receipt / prediction

Published · transcript-backed

Mati Staniszewski: prediction

14 Apr 2026 Cheeky Pint The world of voice AI, with Mati Staniszewski of ElevenLabs

“In speech-to-speech, as you think about maybe more of a companion version of the applications, that's where that will flourish because maybe the hallucinations aren't as important, but the latency is a little bit more, and maybe hallucinations are even a feature.”

— Mati Staniszewski

Source trail

Everything needed to verify it.

Speaker
Mati Staniszewski
Attribution
Verified speaker
Claim type
prediction
Recorded
14 Apr 2026
Publisher
Cheeky Pint

Transcript context

…I'm sorry, a cascaded approach is? Is the speech-to-text... Going through the text layer. As we work with a lot of the businesses and enterprises, they will need that visibility into what happens. They will want to execute certain tasks on top of that. They want a good visibility into each of the steps and great accuracy of all the models. But beyond that, they can abstract away what's the LLM layer, what's the intelligence layer, the integrations are easier in that system. That's where we are betting a lot of the research work of how you can make that great, and we think we can make that great. In speech-to-speech, as you think about maybe more of a companion version of the applications, that's where that will flourish because maybe the hallucinations aren't as important, but the latency is a little bit more, and maybe hallucinations are even a feature. Maybe in the future-future, just to finish that part, you will have some version of combination of the models. That for low complexity, easy models, you will have speech-to-speech, and for higher complexity we will have to cascade it. I was going to ask about this. You know the way there is research on how the invention of writing changed human brains and just changed the neural pathways in ways beyond the actual written language. Do you observe that speech-to-speech models think differently than cascaded models? It sounds like they're dumber.…

Stored transcript either side of the excerpt. The highlighted words are the published quote; the surrounding text is unedited source, never generated.

Search evidence