01 / evaluation
That doesn’t work. At least that doesn’t work that well because audio has more entropy.
“That doesn’t work. At least that doesn’t work that well because audio has more entropy.”
- Speaker
- Pavan Kumar Reddy
- Publisher
- Latent Space
Public evidence record
Published podcast speaker
Claim ledger
6 transcript-backed records
01 / evaluation
“That doesn’t work. At least that doesn’t work that well because audio has more entropy.”
02 / belief
“I think that’s the great part and I feel like with even the existing stack, we should be able to get to this very natural speech conversational abilities soon enough I guess.”
03 / belief
“Each enterprise want a customized, specialized something which is representative both their brand and also their, I guess safety considerations and the use case I think the kind of thing that you would deploy as a empathetic assistant in the context of a healthcare domain would be very different from the kind of thing that would be in a customer support bot and would be different from like more conversational aspects.”
04 / belief
“The walkthrough chat that we released I think July last year, and the follow up transcription only, models family that we released in January, that would be one bucket, and the generation is another bucket.”
05 / preference
“I think at least personally prefer the operations, which are the simplest, and so we try to see, can we just add audio as just another head to our regular transformer decode model because that kind of makes it easier for eventual end-to-end modeling of audio text native modeling.”
06 / prediction
“I think I could have a lot of details. But me I think the [00:27:00] summary of it, actually, some of the considerations in this paper were, because we started with the wipa encoder as the starting point, and now we have in-house encoders, like the bigger time model, for instance, which we released in January.”