Fish Audio released S2.1 Pro, a conversational speech model built for live back-and-forth dialogue rather than pre-scripted narration, and named it the company’s default production voice going forward. Developers can call it today through Fish Audio’s API, and a free tier runs the identical model under fair-use limits rather than a stripped-down variant.

The headline number is speed. Fish Audio clocks the model’s response gap, the delay before the first sound plays, at roughly 90 milliseconds, fast enough to sustain natural turn-taking in a live call. The model holds one consistent voice identity across 83 languages, so a product does not need a different voice per market.

Delivery cues run through bracket tags typed directly into the script instead of a dropdown of preset emotions. A line can be marked to sound hushed and secretive, or shaky with anxiety, and the model reads the cue inline rather than switching voice presets. S2.1 Pro also handles multi-speaker dialogue and clones a voice from a 10 to 30 second reference clip without extra fine-tuning. Fish Audio says the new model beats its predecessor, S2-Pro, on cleanliness of output, response speed, and how many concurrent calls it can carry.

The company is also courting the agent-building crowd directly. S2.1 Pro connects to agent workflows through MCP support and agent-skill integrations, positioning it for developers wiring voice into phone systems, live assistants, and long-running audio pipelines rather than one-off narration jobs.

Real-time conversational voice is no longer a gap in the market; it is where most serious voice vendors now compete, and 83-language coverage reads as a distribution claim more than a quality one. A single voice holding its identity across dozens of languages says nothing about how natural it sounds in any one of them, and Fish Audio’s own post does not publish per-language error rates or blind listening comparisons alongside the launch.

Fish Audio built its following on the open-source Fish Speech project, which has passed 20,000 GitHub stars, and moved through OpenAudio S1 into the open-weight S2 family earlier this year before layering S2.1 Pro on top as the hosted production tier.

Teams evaluating this against ElevenLabs, Cartesia, or OpenAI’s realtime voice stack should benchmark three things before switching: latency under real network conditions rather than the vendor’s own test harness, per-language quality in the specific markets they ship to, and whether the free tier’s fair-use ceiling survives production traffic. A voice model that clones convincingly in English can still degrade in the 60-plus languages nobody at the company is fluent enough to QA.

TestingCatalog reported the S2.1 Pro launch on July 28, 2026.