AI Briefing
KO

Fish Audio Launches S2.1 Pro, Supporting 83 Languages

·2026.07.29 09:00

Key point

Fish Audio has launched S2.1 Pro, a high-performance voice model optimized for real-time conversation.

Details

Fish Audio has launched S2.1 Pro, a production-grade voice model optimized for real-time conversational speech, moving away from the traditional script-reading approach. This model is available through the Fish Audio API, and a free tier for development and testing is also supported.

The core strengths of S2.1 Pro are its low latency and broad language support. Its key features are as follows:

  • Ultra-low latency performance: It records a time-to-first-audio latency of approximately 90ms on standard calls, enabling natural conversation.
  • Multilingual support: It supports 83 languages while maintaining a single voice identity.
  • Fine-grained emotion control: By inputting bracket tags directly within the text, instructions such as whispering or nervous laughter can be executed without separate configuration.
  • High-performance cloning: With just a short sample of 10-30 seconds, it captures the speaker's tone and style without additional fine-tuning.

This model can be easily integrated into agent workflows through MCP and Agent-skill support, making it useful for developers building voice agents, phone systems, or long-form audio pipelines. It was also built on the technical achievements of the open-source Fish Speech lineup, which has recorded over 20,000 stars on GitHub.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.