Gemini 3.8 text-to-speech
by swolpers on 9/23/2026, 3:29:23 PM
https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/
Comments
by: rcr-anti
Pet peeve on Google's AI rollouts: there's no alignment across the three platforms they have, consumer, prosumer, cloud. Scroll to the end of every release, including this one, and you'll see different availabilities. The fun part is the models don't even have the same capabilities across platforms! Omni Flash, last I tried and read the docs, is video and text out on consumer and prosumer but video out only on GCP. So if your org disables consumer and prosumer, like mine, it's a coin flip whether you can use the fancy new models or what they can do.
9/23/2026, 6:22:50 PM
by: simonw
> Voice replication: Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent.<p>I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.
9/23/2026, 4:19:57 PM
by: thangalin
Here's a video of my Emotive Audiobook Creator, KeenLore, a locally hosted web app:<p><a href="https://www.youtube.com/watch?v=WAeHgE94rVo" rel="nofollow">https://www.youtube.com/watch?v=WAeHgE94rVo</a><p>No cloud, no tokens to pay. Reads a book using a full cast of characters. Quotation attribution detection (for my novel) is at 97.2% accuracy (485/499 quotes identified and assigned correctly). The autofill of character voice descriptions uses the prose to determine how the character sounds.<p>Employs Gemma 4[1] for the prose analysis (voice fills, quotation detection) and Qwen3 TTS Voice Design[2] for creating voice samples. Runs on an 8GB NVIDIA T1000 GPU card, 96 GB RAM, and a AMD Ryzen 5 7600.<p>[1]: <a href="https://deepmind.google/models/gemma/gemma-4/" rel="nofollow">https://deepmind.google/models/gemma/gemma-4/</a><p>[2]: <a href="https://huggingface.co/spaces/Qwen/Qwen3-TTS-Voice-Design" rel="nofollow">https://huggingface.co/spaces/Qwen/Qwen3-TTS-Voice-Design</a>
9/23/2026, 4:18:47 PM
by: Multicomp
I direct my own extended daydream Star Trek fanfic (okay, I'm on season 2 episode 17) and recently I looked to see if I could have each scene file be read aloud a la an audiobook or radio drama.<p>Getting GPT-Live to have unique enough voices and to be expressive with how I imagine the voices going in my head is hard to direct, there's not enough control there.<p>So this Gemini 3.8 specific large voice library and ability to tightly control (if you are willing to write a script) is nice to find, and while I'm not sure which of the 5,286 Gemini products this is, nor how to onboard and get started feeding this my own text files, nor what training will happen to my data if I did somehow use it, I love that the state of the industry is such that Google can do this and release it publicly, because that means eventually an equivalent product can come from someone else and be used locally / confidently that the generated audio or inputs won't be retained and misused.
9/23/2026, 4:24:46 PM
by: simonw
I vibe coded a playground UI for trying this out. The conversation mode is neat, and it's very expensive - most of my experiments have cost less than a cent.<p><a href="https://tools.simonwillison.net/gemini-tts-playground#compose=%7B%22v%22%3A1%2C%22mode%22%3A%22single%22%2C%22model%22%3A%22gemini-3.8-flash-tts%22%2C%22single%22%3A%7B%22voice%22%3A%22Kore%22%2C%22style%22%3A%22cheerful%20and%20friendly%22%2C%22text%22%3A%22Have%20a%20wonderful%20day!%22%7D%2C%22speakers%22%3A%5B%7B%22id%22%3A1%2C%22name%22%3A%22Joe%22%2C%22voice%22%3A%22Puck%22%7D%2C%7B%22id%22%3A2%2C%22name%22%3A%22Jane%22%2C%22voice%22%3A%22Kore%22%7D%5D%2C%22turns%22%3A%5B%7B%22speaker%22%3A1%2C%22text%22%3A%22How's%20it%20going%20today%20Jane%3F%22%2C%22style%22%3A%22cheerful%20and%20friendly%22%7D%2C%7B%22speaker%22%3A2%2C%22text%22%3A%22Not%20too%20bad%2C%20how%20about%20you%3F%20Ready%20to%20test%20these%20new%20voices%3F%22%2C%22style%22%3A%22calm%20and%20relaxed%22%7D%5D%7D" rel="nofollow">https://tools.simonwillison.net/gemini-tts-playground#compos...</a>
9/23/2026, 5:19:54 PM
by: seemaze
My primary use case for TTS is converting written content (blogs, articles, etc.) in to clips I can listen to on the go.<p>Is there a good browser extension that does this with a flexible TTS backend? I know Qwen, Kokoro, and VibeVoice all have decent quality..
9/23/2026, 4:58:03 PM
by: accountrequired
"users must provide a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created"<p>How long is this stored? What could go wrong? :P
9/23/2026, 4:54:21 PM
by: nater5000
It's giving me an error when I try to generate a voice with Voice Design in AI Studio. It also says voice replication isn't available in my region.<p>Also weird that there are no "neutral gender" voices in the English language. There's also limited "use cases," like the "Gaming" use case is empty?<p>And there's no pricing listed anywhere.<p>I don't know, I guess their roll out is a bit sloppy. It's a bit of a shame, though, since the voices which are available all sound like generic Gemini voices to me. Nothing stands out is being particularly interesting or impressive about this.
9/23/2026, 4:53:27 PM
by: maelito
Related, for embedding small models, this lib is incredible.<p>Having a voice under 1Mo is crazy, even if it sounds robotic.<p><a href="https://tts.ampixa.com/sanoTTS/" rel="nofollow">https://tts.ampixa.com/sanoTTS/</a>
9/23/2026, 4:40:05 PM
by: Thaxll
What is the best open model / tool for text to speech running locally?
9/23/2026, 5:52:49 PM
by: 112233
"Super tinny monotone robotic voice" does not sound neither tinny nor monotone. Compared to what TTS from 90s sounded like. Or even how actors impersonated robots in movies. Has the model been eating too much hype DJs?
9/23/2026, 4:23:29 PM
by: xnx
Would be great if this would power the Google Books app feature. The voice system there is pretty out of date.
9/23/2026, 4:32:34 PM
by: hatingisok
My ROFLcopter goes: SOISOISOISOISOISOISOISOISOISOISOISOISOISOISOISOI
9/23/2026, 7:32:51 PM
by: dangoodmanUT
In their demo for "Monologue" (2nd video over, after the medeival video game example), it is clearly ignoring the vocal cues like `<chuckles>` and `<laughing>`...
9/23/2026, 9:03:36 PM
by: burkaman
Seems like voice actors are safe for now. This is technologically incredible, but the results are really not very good, and usually not particularly close to the prompt. In basically all of these examples some core part of the prompt is completely ignored.
9/23/2026, 4:48:45 PM
by: sgc
Sorry if this is in that article, but I am on my phone and can't see it. How much would this cost to batch generate an audiobook? Right now I just listen to things in the 11 labs app which is free, but I would rather just generate audio files.
9/23/2026, 4:29:02 PM
by: nitroedge
Couldn't see this in the article, does it support API calls for one-shot conversation type responses like ElevenLabs offers?<p>$0.50 per hour pricing could last a long time with back and forth conversation use.
9/23/2026, 5:04:05 PM
by: tantalor
Would pay any amount of money for a zombo.com voice.<p>The original is the best: <a href="https://youtu.be/qxWwEPeUuAg" rel="nofollow">https://youtu.be/qxWwEPeUuAg</a>
9/23/2026, 8:23:38 PM
by: yipinwong
Google is spreading too thin, as gemini isn't really that intelligent.<p>They are creating gemini SOTA (not really any more), flash versions, text-to-speech, video (omni), etc.<p>I can see they want to create an ecosystem, but I see no focus in any one area.
9/23/2026, 7:26:43 PM
by: AyanamiKaine
Still, all voices sound like they are missing something only real human speech can sound like. But many people will not notice the difference between AI and normal voices.
9/23/2026, 7:06:03 PM
by: LarsDu88
I've been trying to track the SOTA for years switching between wavenet on Gcloud, to Azure, to ElevenLabs, and now Fish.audio. This is damn good
9/23/2026, 5:30:22 PM
by: newhotelowner
Where can I get text to speech audio easily?
9/23/2026, 8:22:55 PM
by: cainxinth
They keep announcing new 3.8 variants. I wonder why they still haven't updated 3.1 Pro yet.
9/23/2026, 5:17:57 PM
by: perrohunter
Gemini 3.8 "Flash" says hello
9/23/2026, 4:22:41 PM
by: andrewstuart
I’ve never found a TYS that does convincing British accents.<p>They all sound like Americans putting in their best fake British accent.
9/23/2026, 4:35:05 PM
by: fullstackwife
It would be nice to have sound effect generation (use case: games)
9/23/2026, 4:59:15 PM
by: dadoum
here is a competitor, if someone wants to compare <a href="https://gradium.ai/" rel="nofollow">https://gradium.ai/</a>
9/23/2026, 5:24:21 PM
by: drewbitt
Great price at least until December 31 too.
9/23/2026, 4:44:55 PM
by: dainiusse
Where is "pro"?
9/23/2026, 5:48:48 PM
by: m3kw9
Still sounds AI, you can tell they exaggerate all the tone and trailing "high scoring expressive sounds" like your job depends on it.
9/23/2026, 4:57:21 PM
by: talon8635
Great. Now in additional to AI email responses I will get AIs impersonating my contacts on the phone too. Lovely.
9/23/2026, 4:25:38 PM
by: WarmWash
I hope this spills over into more voices for android/android auto default assistant voices. The current choices are all so meh
9/23/2026, 6:10:46 PM
by: OutOfHere
Pricing isn't noted.
9/23/2026, 5:18:25 PM
by:
9/23/2026, 4:14:33 PM
by: UmYeahNo
HN: Lament devs losing jobs to AI<p>Also HN: Fuck those voice actors and their careers.
9/23/2026, 7:27:22 PM
by: xmorse
Based on this video I think this model was trained on p*rn<p><a href="https://storage.googleapis.com/gweb-uniblog-publish-prod/original_videos/AudioWaveform-Blue_SingleSpeakerCES_Teleprompter.mp4#t=5" rel="nofollow">https://storage.googleapis.com/gweb-uniblog-publish-prod/ori...</a>
9/23/2026, 4:59:07 PM