Founded as Hanabi AI Inc., Fish Audio traces its origins to a side project created by co-founder and Chief Scientist Shijia Liao, a former NVIDIA video researcher. Frustrated by the limited expressiveness of early synthetic voices, Liao trained his first voice generation models using a single GPU. That work evolved into Fish Speech, an open-source project that has accumulated more than 31,000 GitHub stars and attracted developers, game studios, and content creators building AI-powered voice applications.
Since launching in 2023, Fish Audio has expanded into a platform for text-to-speech, voice cloning, and voice agents. The company says its models allow developers to control tone, pacing, and emotional expression using more than 15,000 natural-language prompts, with support for 83 languages.
Its flagship S2.1 Pro model can clone a voice from a five-second audio sample in under 15 seconds, according to the company. Fish Audio also said blind listening tests found that 67% of participants preferred its voice outputs over competing models. The startup reports more than 8 million users across its open-source and hosted products and says it has surpassed $21 million in annual recurring revenue.
While the platform initially gained traction among creators and game developers, Fish Audio has expanded into enterprise deployments for organizations in regulated industries, including healthcare and financial services. The company offers on-premises deployments with zero-data retention and HIPAA compliance, and says enterprise customers include companies such as HeyGen and Sanas.
Chief Executive and co-founder Rissa Cao said the company was founded with the goal of making realistic AI voices broadly accessible. "We make high-quality, human-sounding voices available to every user, from beginner creatives to million-dollar enterprises, so communication is not only more efficient, but more trustworthy," Cao said. "We've always believed that if we kept making the models better, people would notice. Eight million of them did."
Cao also told TechCrunch that enterprise demand varies widely depending on customer needs. "Every enterprise has different use cases and different preferences. For example, companies like HeyGen, which use our voices to power AI avatars, want realism in voices; a gaming studio would want expressive voices for their characters; LiveKit-style voice agent companies want natural sound and low latency for phone calls."
The company plans to use the new funding to develop additional voice-native large language models, speech-to-speech translation technology, and audio understanding models. It also intends to grow its enterprise sales organization and expand API integrations with partners including Retell AI and LiveKit.
As part of its developer strategy, Fish Audio said it will make its flagship S2.1 Pro model available free through its official API beginning at the end of August.
Fish Audio has also updated its approach to content moderation after creators alleged their voices had been uploaded without permission. Cao said the company has automated its DMCA takedown process so creators can submit a voice sample or proof of ownership and have unauthorized voices removed in under three minutes.
Coreline Ventures Managing Partner Osuke Honda said he believes voice is becoming the primary way people interact with AI systems. "In its short history, Fish Audio has built an unbeatable track record of pushing the envelope on performance, multilingual support, emotional expression and cost," Honda said. "All factors that have quickly made Fish Audio the default choice for creators, developers, and now enterprises globally, and we expect them to continue to lead the way."
This analysis is based on reporting from TechCrunch & Silicon Angle.
Image courtesy of Fish Audio.
This article was generated with AI assistance and reviewed for accuracy and quality.