Former Nvidia researcher open-sourced a voice model out of personal frustration – one year later: $21M ARR, 8M users, $52M raised.
ENTRY ANGLES
Open-source a domain-specific voice model to acquire developer distribution before charging for the API · Build a voice model optimized for one vertical where generic models underperform – medical dictation, legal narration, or regional call center accents · Enter the non-English voice AI market where ElevenLabs and current leaders have thin coverage
VERTICALS
CAPABILITIES
ML research (voice models), developer relations, open-source distribution strategy, enterprise API sales
Voice synthesis has been commoditized before it was solved. The tools are everywhere, and almost all of them sound like tools.
Shijia Liao noticed this while working as a researcher at Nvidia. The synthetic voices available in 2024 – for creative projects, content production, enterprise automation – shared a common flaw: technically impressive, expressively flat. Liao trained a voice generation model on a single GPU, open-sourced it, and moved on.
What happened next was not planned. The model spread. Developers building creative tools integrated it. HeyGen – which produces AI video with synthetic presenters – made it core to its voice output. Sanas, a voice-clarity platform used in call centers, built on top of it. Plaud, an AI voice recorder, adopted it. GitHub stars accumulated faster than most open-source projects reach in years. Liao formalized the project into a company: Fish Audio.
Within twelve months the company had $21M in annual recurring revenue and 8 million users across creators, developers, and enterprises. The flagship model family – Fish-Speech and S2 – supports more than 80 languages. On July 28, 2026, one year in, the company raised a $52M seed round led by Coreline Ventures and Capital Today, with participation from 645 Ventures, HF0, and a handful of others. It was the first institutional money Fish Audio had ever taken.
The capital is earmarked for expanding from text-to-speech into what the company calls the full audio-native stack: voice-native large language models, real-time speech-to-speech conversion, and deeper developer integrations with platforms like LiveKit and Retell. The text-to-speech product is being treated as the base layer for something considerably larger.
ElevenLabs, the voice AI market leader, reached $80M in annual recurring revenue by end of 2024 after raising $180M. Fish Audio reached $21M ARR after raising nothing – which means voice synthesis is a market where execution speed and product quality can beat capital, at least until the capital catches up.
Fish Audio's position is shaped by two structural advantages its English-first competitors mostly ignored. The first is multilingual depth. The 80-language capability isn't a feature list item – it's a distribution moat. Most of the 8 million users on the platform are not American. The demand for expressive, high-quality synthetic voice in non-English languages is enormous relative to what the market has produced. Standard voice AI in Mandarin, Arabic, or Portuguese is noticeably worse than in English; Fish Audio's model, trained without the English-first bias, closes that gap.
The second is the open-source feedback loop. Fish Audio understood its users with unusual depth before the company existed, because thousands of developers were building with the raw Fish-Speech model and reporting back what didn't work. By the time Liao started charging for the API, the product had already been shaped by a user base most early-stage companies would pay millions to access.
The "audio-native stack" roadmap reflects where AI interfaces are actually going. Text was the default interface for AI because text was the primary training data format. Voice is the natural interface for human communication. Building the full infrastructure layer for voice-native AI – models that process and respond in audio rather than converting from text – is a substantial bet that has few credible competitors.
Fish Audio's path from hobby project to $21M ARR is a clean version of a pattern worth understanding: open-source as distribution, API as conversion. Release the model to get genuine developer adoption, then build the hosted platform business on the distribution you already created. The pattern is not new – Hugging Face and MongoDB built significant businesses on variants of it – but Fish Audio demonstrates it works in voice AI specifically.
What most companies underestimate is how good the open-source product needs to be. Fish-Speech spread because it was genuinely better, not just free. A mediocre model released as open source doesn't get this kind of adoption.
ElevenLabs occupies the premium English market. Fish Audio is moving fast on multilingual and developer-facing use cases. The underserved gap is domain-specific voice: models trained specifically for a vertical's terminology, cadence, and persona requirements. Healthcare wants voices that sound clinically credible. Legal needs precision without performative warmth. Call center deployments need regional accents, natural code-switching between languages, and fast cold-start on unfamiliar names.
A concrete entry: train an open-source voice model specifically for one domain – medical dictation, for example, where the existing voice-to-text options consistently fail on clinical terminology. Distribute through the channels Fish Audio proved work. The barrier to the first step is dramatically lower than it was two years ago; the barrier to the second step is distribution, and a genuinely good open model solves that faster than any sales motion.