This is really cool work! I'm curious like what do you see as the biggest lever for speeding up TTS models or from a technical perspective that this was a promising direction in the first place to push on. If I were to guess, some distillation but I'm certain there are probably TTS model aware architectural changes that just make inference wayyyy faster?
Voice models are not winner take all market unlike LLM APIs
Coming here as Developer Relations at AssemblyAI
https://github.com/loudreader/loudkit
I think real time natural tts should be possible everywhere soon
For OP the clip name is nari-nina-01a0a12f-980a-765e-8029-fa56bd23210d.wav
Is the ASR inference engine open source as well?