Google released Gemini 3.1 Flash TTS on April 15, its latest text-to-speech model and the newest addition to the Gemini 3.1 family. It’s built for developers, enterprises, and Workspace users who want more control over how AI-generated speech sounds.

Google isn’t the only company pushing hard on voice AI right now. ElevenLabs, OpenAI, and Inworld have all shipped major updates in recent months, and the leaderboard rankings have been shifting. Gemini 3.1 Flash TTS enters that race with a strong benchmark score, new speech controls, and wide language support.

Here are a few things you should know about the model:

/1. It Ranks Number Two on the Global AI Voice Leaderboard

The Artificial Analysis TTS leaderboard, which ranks models based on blind human listening tests across thousands of comparisons, currently places Gemini 3.1 Flash TTS at number two globally with an Elo score of 1,211. Inworld TTS 1.5 Max holds the top spot at 1,215, and ElevenLabs Eleven v3 sits third at 1,179.

Artificial Analysis also placed Gemini 3.1 Flash TTS in its "most attractive quadrant," a designation for models that balance high speech quality with low generation cost.

Subscribe for free to continue reading this article

Subscribe Subscribe

Already Have an Account? Log In