Gemini 3.8 text-to-speech

(blog.google)

68 points | by swolpers 1 hour ago

19 comments

  • seemaze 3 minutes ago
    My primary use case for TTS is converting written content (blogs, articles, etc.) in to clips I can listen to on the go.

    Is there a good browser extension that does this with a flexible TTS backend? I know Qwen, Kokoro, and VibeVoice all have decent quality..

  • simonw 41 minutes ago
    > Voice replication: Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent.

    I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.

    • Multicomp 40 minutes ago
      They probably do something similar to GPT-Live where they expect a given voice profile to send them a sample saying 'This is the owner of this voice and I consent for synthetic samples to be made of it'

      and/or local voice cloning is good enough as is so Google doesn't grant a uniquely liable ability?

    • miltonlost 25 minutes ago
      "Don't be evil... unless other companies are doing it first"
      • imjonse 12 minutes ago
        voice cloning is a tool, it is not necessarily evil, even though the scenarios it can be used for nefarious purposes outnumber the legitimate ones.
        • bakies 1 minute ago
          is the legit ones just like... putting carrie fisher in star wars?
  • nater5000 8 minutes ago
    It's giving me an error when I try to generate a voice with Voice Design in AI Studio. It also says voice replication isn't available in my region.

    Also weird that there are no "neutral gender" voices in the English language. There's also limited "use cases," like the "Gaming" use case is empty?

    And there's no pricing listed anywhere.

    I don't know, I guess their roll out is a bit sloppy. It's a bit of a shame, though, since the voices which are available all sound like generic Gemini voices to me. Nothing stands out is being particularly interesting or impressive about this.

  • Multicomp 36 minutes ago
    I direct my own extended daydream Star Trek fanfic (okay, I'm on season 2 episode 17) and recently I looked to see if I could have each scene file be read aloud a la an audiobook or radio drama.

    Getting GPT-Live to have unique enough voices and to be expressive with how I imagine the voices going in my head is hard to direct, there's not enough control there.

    So this Gemini 3.8 specific large voice library and ability to tightly control (if you are willing to write a script) is nice to find, and while I'm not sure which of the 5,286 Gemini products this is, nor how to onboard and get started feeding this my own text files, nor what training will happen to my data if I did somehow use it, I love that the state of the industry is such that Google can do this and release it publicly, because that means eventually an equivalent product can come from someone else and be used locally / confidently that the generated audio or inputs won't be retained and misused.

  • thangalin 42 minutes ago
    Of possible interest is my Emotive Audiobook Creator, KeenLore. Here's a video of the web app showing how it works:

    https://www.youtube.com/watch?v=WAeHgE94rVo

    Locally hosted, no cloud, no tokens to pay. Reads a book using a full cast of characters. Quotation attribution detection (for my novel) is at 97.2% accuracy (485/499 quotes identified and assigned correctly). The autofill of character voice descriptions uses the prose to determine how the character sounds.

    Uses Gemma 4[1] for the prose analysis (voice fills, quotation detection) and Qwen3 TTS Voice Design[2] for creating voice samples. Runs on an 8GB NVIDIA T1000 GPU card, 96 GB RAM, and a AMD Ryzen 5 7600.

    [1]: https://deepmind.google/models/gemma/gemma-4/

    [2]: https://huggingface.co/spaces/Qwen/Qwen3-TTS-Voice-Design

    • Multicomp 41 minutes ago
      <grumble grumble people putting in links they expect you to follow to arbitrary goatse youtube videos for all I know>

      The title of the video is 'KeenLore - Emotive Audiobook Creator Demo' and it appears to be a web UI and some local stack that reads text files.

    • loremm 36 minutes ago
      It's cool technology and I read a lot of audiobooks, even hundreds of hours of TTS. I feel like my brain can fill in the character voices from the text - on the page it's not like they're different fonts.

      I understand audiobook narrators often do it, and that's fun. But it's not so critical in my opinion

    • talon8635 34 minutes ago
      Our imaginations and minds continue to rot under the weight of endless and effortless entertainment
  • xmorse 2 minutes ago
    Based on this video I think this model was trained on p*rn

    https://storage.googleapis.com/gweb-uniblog-publish-prod/ori...

    • hmokiguess 0 minutes ago
      That argument says as much about the training as it says about your sexuality, you do realize that right.
    • droidjj 0 minutes ago
      It’s just a woman’s voice…
  • maelito 21 minutes ago
    Related, for embedding small models, this lib is incredible.

    Having a voice under 1Mo is crazy, even if it sounds robotic.

    https://tts.ampixa.com/sanoTTS/

  • xnx 29 minutes ago
    Would be great if this would power the Google Books app feature. The voice system there is pretty out of date.
    • laweijfmvo 21 minutes ago
      the ratio of new voice models i see on hackernews to the number actually deployed in any product i use is approximately infinity.
    • mamudo 22 minutes ago
      Yes, I am quite disappointed by seeing all this cool AI stuff and yet the same Play Books. Come on, it is the best place to apply AI, in my opinion.
  • burkaman 12 minutes ago
    Seems like voice actors are safe for now. This is technologically incredible, but the results are really not very good, and usually not particularly close to the prompt. In basically all of these examples some core part of the prompt is completely ignored.
    • avazhi 3 minutes ago
      > Seems like voice actors are safe for now. This is technologically incredible, but the results are really not very good,

      Um, what?

  • fullstackwife 2 minutes ago
    It would be nice to have sound effect generation (use case: games)
  • accountrequired 7 minutes ago
    "users must provide a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created"

    How long is this stored? What could go wrong? :P

  • 112233 38 minutes ago
    "Super tinny monotone robotic voice" does not sound neither tinny nor monotone. Compared to what TTS from 90s sounded like. Or even how actors impersonated robots in movies. Has the model been eating too much hype DJs?
    • burkaman 6 minutes ago
      None of these examples are really what the prompt asked for. It's just like image models, once you get over how unbelievable it is that a computer produced this you realize the result isn't actually what you want.
  • sgc 32 minutes ago
    Sorry if this is in that article, but I am on my phone and can't see it. How much would this cost to batch generate an audiobook? Right now I just listen to things in the 11 labs app which is free, but I would rather just generate audio files.
    • thevinter 29 minutes ago
      Roughly 5-10$ for 10h, assuming you few-shot it.

      Price per hour:

      - 3.8 Flash TTS, standard: $0.81

      - 3.8 Flash TTS, batch: $0.41

      - 3.8 Flash‑Lite TTS, standard: $0.54

      - 3.8 Flash‑Lite TTS, batch: $0.27

  • m3kw9 4 minutes ago
    Still sounds AI, you can tell they exaggerate all the tone and trailing "high scoring expressive sounds" like your job depends on it.
  • drewbitt 16 minutes ago
    Great price at least until December 31 too.
  • perrohunter 39 minutes ago
    Gemini 3.8 "Flash" says hello
  • andrewstuart 26 minutes ago
    I’ve never found a TYS that does convincing British accents.

    They all sound like Americans putting in their best fake British accent.

  • talon8635 36 minutes ago
    Great. Now in additional to AI email responses I will get AIs impersonating my contacts on the phone too. Lovely.