
Image to help understand the article
New voice tools arrive across Google’s AI services
Google unveiled Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on Sept. 23, 2026, highlighting voice customization and accent reproduction as key features. The models are available through the Gemini API and Google AI Studio under the IDs gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts. Access through the Gemini Enterprise API is expected later.
Gemini 3.8 Flash TTS is also available to all Gemini Notebook users, while Flash-Lite TTS is available to all users in Google Vids. The Next Web reported that the models support more than 100 languages and dialects and offer more than 2,000 premade voices.
Flash TTS can clone a voice from a 30-second sample after a consent-verification process, according to The Next Web. Voice cloning through Google AI Studio is unavailable in Illinois, Texas, the European Economic Area, Britain, Switzerland and India.
Audio produced by Gemini Audio models includes Google’s SynthID watermark to identify AI-generated speech. Cloned voices also carry C2PA content credentials. RuntimeWire reported, however, that Google had not disclosed specific details about its consent checks, latency or usage figures at launch, and that the claims had not been independently verified.
Google’s benchmark claims face independent scrutiny
Google said Gemini 3.8 Flash TTS ranked first in Hume AI’s Voice Design Benchmark, scoring 71.4 overall and 60.8 for accent modeling. The company also said Flash TTS and Flash-Lite TTS placed first and second, respectively, in Hume AI’s Overall Quality Index.
In blind preference testing by Voice Arena, the models ranked among the leading competitors for Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi.
RuntimeWire cautioned that Google announced the Hume benchmark results and that the claims alone do not settle comparisons among products. The publication also reported that Alan Cowen, Hume AI’s former CEO and now a research science director at Google DeepMind, and Leland Rechis participated in the launch work.
An independent Artificial Analysis leaderboard produced a different ranking. Gemini 3.8 Flash TTS placed second with an Elo score of 1,260 and a 95% confidence interval of plus or minus 17. Flash-Lite TTS ranked sixth with an Elo score of 1,235 and a confidence interval of plus or minus 16.
Cartesia Sonic 3.6 led the ranking with 1,757 samples and an Elo score of 1,273. Alibaba’s Qwen-Audio-3.0-TTS-Plus ranked third at 1,259, followed by Inworld Realtime TTS-2 at 1,245. ElevenLabs v3 Conversational ranked 12th with an Elo score of 1,196 and a confidence interval of plus or minus 15.
Gemini 3.8 Flash TTS led Artificial Analysis’ Pronunciation Robustness Benchmark with a score of 89.5%, compared with 88.2% for the earlier Gemini 3.1 Flash TTS.
OrcaRouter noted that Flash TTS and Qwen-Audio-3.0-TTS-Plus were separated by just one Elo point in the Provider Voice Arena. It said the broad confidence intervals placed the top three models within the same statistical range.
Free access through 2026, followed by tiered pricing
Google is offering both models free under its standard tier through Dec. 31, 2026. Beginning Jan. 1, 2027, Gemini 3.8 Flash TTS will cost $1 per 1 million input tokens, $18 per 1 million audio-output tokens and 25 cents for context caching. Flash-Lite TTS will cost $1 per 1 million input tokens and $12 per 1 million audio-output tokens.
Batch and Flex pricing will range from 25 cents to 50 cents per 1 million input tokens for both models. Output will cost $4.50 to $9 for Flash TTS and $3 to $6 for Flash-Lite TTS.
The Priority tier is also free through 2026. After that, Flash TTS will cost $1.80 per 1 million input tokens and $32.40 per 1 million output tokens. Flash-Lite TTS will carry the same input price and a $21.60 output price.
Comments
Post a Comment