10 comments

  • iharnoor 16 minutes ago
    By next month the competition for TTS will be even more!

    Voice models are not winner take all market unlike LLM APIs

    Coming here as Developer Relations at AssemblyAI

  • mowmiatlas 1 hour ago
    Cool, I’ve released something to the same beat of the dr this weekend as well

    https://github.com/loudreader/loudkit

    I think real time natural tts should be possible everywhere soon

  • rahimnathwani 1 hour ago
    For some reason it switched voices half way through a 33 second clip.

    For OP the clip name is nari-nina-01a0a12f-980a-765e-8029-fa56bd23210d.wav

  • yoloakki 44 minutes ago
    You definitely need independent evals by Datapoint AI or someone who can verify your claims about TTS quality
  • ipsum2 1 hour ago
    If you're going to announce a TTS model, service, or whatever, you really need demos.
  • asaiacai 2 hours ago
    This is really cool work! I'm curious like what do you see as the biggest lever for speeding up TTS models or from a technical perspective that this was a promising direction in the first place to push on. If I were to guess, some distillation but I'm certain there are probably TTS model aware architectural changes that just make inference wayyyy faster?
  • DylanMerigaud 23 minutes ago
    Rooting for you on this one.
  • meatmanek 1 hour ago
    > and Qwen3-ASR

    Is the ASR inference engine open source as well?

  • nthypes 1 hour ago
    [dead]