Fish Audio Raises $50M Seed to Build AI Voice Models for Creators and Enterprises

ai voice modelscreator toolsenterprise voice aifish audioopen-source voice generationseed fundingvoice cloning

The market for AI-generated voice models continues to expand rapidly in 2026, driven by demand from both creative professionals and enterprises. Creatives require expressive, nuanced voice synthesis for storytelling and gaming, while businesses seek steerable, low-latency voice solutions for customer support and sales automation.

Palo Alto-based Fish Audio aims to serve both segments with its library of over 15,000 natural language controls. Since its launch in 2024, the startup has attracted more than 8 million users across its open-source and hosted model versions. As of early 2026, Fish Audio generates $21 million in annual recurring revenue (ARR), reflecting strong commercial traction.

To accelerate growth, Fish Audio announced a $50 million seed round led by Coreline Ventures and Capital Today, with participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0. The funding will support model development, platform scaling, and enterprise go-to-market efforts.

Founder Shijia Liao, a former NVIDIA researcher, started Fish Audio as a side project. Frustrated by the lack of expressive synthetic voices, he trained a voice generation model on a single GPU and open-sourced it. The resulting Fish Speech repository on GitHub has earned over 31,000 stars, gaining a community of indie developers, video game designers, and content creators.

Fish Audio has released five models to date: four for speech generation and one for speech-to-text. Three speech generation models are open-source, while its latest S2.1 Pro model is available exclusively through a paid API. The startup offers monthly subscription plans for creators and teams, providing minutes of generation and voice cloning capabilities. Its enterprise API platform is used by organizations such as HeyGen, Sanas, and Plaud.

“Every enterprise has different use cases and preferences,” said Rissa Cao, CEO and co-founder of Fish Audio. “HeyGen uses our voices to power AI avatars and needs realism; gaming studios want expressive voices for characters; and voice agent companies like LiveKit require natural-sounding, low-latency voices that are expressive enough for live calls.”

To build its voice library, Fish Audio invited users to submit their own voices for training, compensating those whose voices were used. However, this approach attracted criticism in 2025 when some creators alleged their voices were uploaded without consent. Although Fish Audio had a DMCA takedown process, removals were slow.

In response, Cao said the company has automated the takedown process. Creators can now submit a short voice sample or contract to prove ownership, and their voice will be removed from the platform in under three minutes. Nevertheless, the system still relies on creators discovering unauthorized uploads before filing a takedown request.

“Until the artist finds out, their voice will continue to be used on the platform,” Cao acknowledged, highlighting an ongoing challenge for the industry.

Oskue Honda, partner at Coreline Ventures, noted that Fish Audio’s combination of open-source innovation and enterprise-grade control positions it well for the next wave of voice AI adoption in 2026 and beyond.

via TechCrunch

Related