If you have been waiting for voice AI to grow up, this was the day. Google's Gemini 3.8 Flash TTS lets you design a voice from a written description, ChatGPT Voice can now book meetings, search your inbox and hand hard jobs to GPT-6, and Alibaba's Qwen-Audio-3.1 cuts speech prices by up to 95 per cent. NVIDIA added an open model that finally untangles who said what in a messy recording. Away from audio, Alibaba put three agents inside Android phones, Anthropic pointed hundreds of Claude agents at a biology problem and found something new, and Black Forest Labs and Apple both released open models worth a look. The pattern is familiar from text models a year ago: quality climbs, prices fall, and the tools start doing real work rather than demos.
Voice AI launches at a glance: Gemini TTS, ChatGPT Voice, Qwen-Audio-3.1 and Nemotron
| Launch | What it does | Where to get it |
|---|---|---|
| Gemini 3.8 Flash TTS and Flash-Lite TTS | Design a voice from a text prompt, 100+ languages, voice replication with consent | Gemini API and Google AI Studio now |
| ChatGPT Voice | Uses plugins (calendar, email, Slack), runs on GPT-6, builds files in ChatGPT Work | Web, iOS and Android |
| Qwen-Audio-3.1 | Five-model ASR, TTS and Realtime stack; prices cut 70 to 95 per cent | QwenCloud API |
| Nemotron 3 Diarization | Open 100M model that labels up to eight speakers, even overlapping | Free on Hugging Face |
Gemini 3.8 Flash TTS and Flash-Lite TTS let you design a voice from a text prompt
If you make podcasts, audiobooks, explainer videos or voice agents, this is the release to try today. You can now describe a voice in plain English and get it, fine-tune delivery line by line (pacing, emotion, even laughs and pauses) and, with consent, replicate a speaker from a 30-second sample. The two models extend the Gemini 3.8 family into speech, cover more than 100 languages and took the top two places on Hume AI's quality index, which makes them a serious challenge to ElevenLabs on price and quality.
Everything generated carries a SynthID watermark and replicated voices get C2PA credentials, which matters if you publish commercially. Developers can use both now in the Gemini API and Google AI Studio, and they are coming to Gemini Notebook for everyone else.
ChatGPT Voice can now use plugins, run on GPT-6 and work inside ChatGPT Work
Voice mode has gone from a chat novelty to something you can get work done with. You can now ask it to check your calendar for a free slot and book it, search your email, or look something up in Slack, all by speaking. It can hand harder jobs to GPT-6 Astra or the new GPT-6 Sol and Luna in the background (depending on your plan), and inside ChatGPT Work it can build documents, decks, sites and spreadsheets on request.
The rollout covers web, iOS and Android. If you have avoided voice because it could not touch your real tools, this is the update that changes that.
MentalHealthBench: an open benchmark for how chatbots handle mental health conversations
Plenty of people already talk to chatbots about how they are feeling, so an open yardstick for how well models handle that is useful for anyone building or choosing an assistant. MentalHealthBench was written with more than 80 clinicians across 22 countries and 19 languages, and covers everyday support as well as crisis situations.
It contains 1,215 synthetic conversations and is released openly so other labs can test against it. In the first results GPT-6 Astra scored highest, with Claude Opus 5.5 close behind GPT-6 Sol.
Airbnb widens GPT-6 Astra access and ships 80 per cent more features
A useful real-world data point on what frontier models do inside a big company. Airbnb says its development teams are shipping roughly 80 per cent more features than a year ago, with GPT-6 Astra a core part of its developer tooling.
Access runs through both OpenAI's own API and Amazon Bedrock, a sign that OpenAI models are now a normal option inside AWS shops. If you are making the case for AI coding tools at work, this is a case study worth quoting.
Qwen-Audio-3.1 arrives as a five-model audio stack with price cuts of up to 95 per cent
If you pay for speech recognition or text-to-speech by the minute, check your bill against this. Alibaba has upgraded its ASR, TTS and Realtime voice models and added two new ones: TTS-Next, which generates voice, sound effects and background audio in one pass, and ASR-Next, which returns timestamped, speaker-labelled transcripts and can answer questions about emotion and background sounds.
Alibaba claims list-price cuts of around 70 per cent for TTS, 85 per cent for Realtime and up to 95 per cent for ASR. The Realtime-Plus model supports interruptions, tool use and a 262K context window, which makes it a credible budget engine for voice agents.
Qwen Intelligence brings three specialised AI agents to Android smartphones
This is Alibaba's play to become the brain inside Android phones, following the roadmap it set out at Apsara yesterday. Qwen Intelligence splits phone automation across three agents: a Planner that breaks multi-step jobs into tasks, a Mobile-Use agent that operates apps (API first, tapping the screen as a fallback), and a Creative agent that turns a sentence into an image in about three seconds.
Alibaba reports a 90 per cent end-to-end task completion rate and leads on its own MobileWorld benchmarks, which it has also open-sourced. HONOR is the first partner, with the Magic9 series launching on 28 September, so UK buyers of Chinese-brand phones may meet this sooner than expected.
Nemotron 3 Diarization, an open model that works out who spoke when
If you record meetings, interviews or podcasts, messy transcripts with overlapping speakers are a familiar pain. NVIDIA's new open model tracks up to eight speakers, even when they talk over each other, and at 100M parameters it is small enough to run locally.
It ranked first of 12 systems on Voice Arena's diarisation benchmark, with an error rate about 24 per cent lower than the runner-up. It is free on Hugging Face now, with day-one support in Transformers, so expect meeting-notes apps to pick it up quickly.
Claude agents find a previously unknown enzyme system with CRISPR-like repeats
This is not a product launch, but it shows where Claude's agentic abilities are heading. In the first result from Anthropic's new molecular biology lab, roughly 950 Claude agents sifted more than 200,000 reverse transcriptases over 21 hours and flagged a new system sitting next to a CRISPR-like DNA array.
Nobody yet knows what it does, and Anthropic has published a preprint inviting outside researchers to help. For everyday users the takeaway is that long-running, multi-agent research is becoming a real use case, not a demo.
Black Forest Labs releases FLUX 3 Action, an open-weight model for robot control
The team behind the FLUX image models has moved into robotics. FLUX 3 Action is a 7B open-weight model that predicts a robot's next actions and the video frames that follow, and Black Forest Labs reports it beats NVIDIA's Cosmos 3 Nano on the RoboLab benchmark with far fewer parameters.
The weights are on Hugging Face and it can be fine-tuned through LeRobot, which puts it within reach of hobbyists with a cheap robot arm.
Apple publishes LensVLM-9B on Hugging Face for cheaper long-document reading
Apple rarely releases models, so this one is worth noting. LensVLM-9B compresses long documents into page images and only zooms into the pages it needs to answer a question, a smart way to handle long PDFs cheaply.
It is released under Apple's research licence, so it is for experimentation rather than commercial products.
Industry themes
Voice was the story of the day. Google, OpenAI, Alibaba and NVIDIA all shipped speech products within about ten hours of each other, and Alibaba's steep price cuts suggest voice is heading the same way text models went: fast gains and falling prices. If you pay a specialist voice provider today, it is worth re-benchmarking this month.
The second pattern is agents moving onto the devices and tools you already use, from ChatGPT Voice working with your calendar and inbox to Qwen Intelligence operating phone apps. The question is no longer whether an assistant can talk, but whether it can act on what you say.
Several of these, notably Qwen-Audio-3.1 and Nemotron 3 Diarization, appeared first on the companies' own X accounts with little mainstream press coverage. Open weights also had a good day, with NVIDIA, Black Forest Labs and even Apple putting models on Hugging Face.