Fish Audio's Shijia Liao built a voice model on one GPU, open-sourced it, watched 8 million users arrive — and then raised $52M on $21M ARR.
ENTRY ANGLES
Character-as-a-service API for persistent AI voice identities in long-form creative projects · Voice localization with prosody controls trained on conversational data for major non-English markets · Developer tooling for voice-native application builders on top of the open-source base
VERTICALS
CAPABILITIES
Expressive TTS model training at scale, Natural language control surface design, Developer API infrastructure and SDKs, Enterprise reliability and compliance guarantees for voice output
The standard path to $52 million in seed funding begins with a pitch deck, a market map, and a theory of defensibility. Shijia Liao's path began with a single GPU and a personal frustration with synthetic voice quality. Working at Nvidia, he found that the text-to-speech options available to indie developers were flat without exception: the intonation was mechanical, the emphasis was wrong, the emotional range was absent. He trained a voice generation model on his own hardware and posted it to GitHub.
The Fish Speech repository now has more than 31,000 stars. Liao turned the open-source project into a hosted product, built enterprise tooling around it, and watched $21 million in annual recurring revenue accumulate before he raised a cent of institutional capital. The $52 million seed round, led by Coreline Ventures and Capital Today with 359 Capital, HF0, and Parable participating, closes with 8 million users already on the platform across both open-source and hosted versions.
The technical differentiator is the control surface. Standard voice AI models offer a small parameter set: speed, pitch, an emotion category or two. Fish Audio exposes more than 15,000 natural language controls — specific, composable instructions governing how a voice should behave in a given passage. A video game designer can specify "gravelly, exhausted, slightly menacing" without engineering custom fine-tuning. A healthcare enterprise can specify "clear, patient, clinical" and get deterministic output. The same underlying model serves both because the vocabulary of control is large enough to reach both.
The voice AI market is splitting into two segments that have almost nothing in common except the underlying technology. The creator segment needs expressiveness — a voice that carries narrative tension, performs comedy, portrays a specific character across an extended piece of content. Expressiveness requires large control surfaces and models trained on diverse emotional range. The enterprise segment needs steerability — consistent, compliant output that doesn't deviate from approved language and can be adjusted by non-engineers. Steerability requires deterministic controls and enterprise-grade reliability guarantees.
Most voice AI companies have picked one side. ElevenLabs built its reputation on creator use cases; Deepgram and Rime are targeting regulated enterprise. Fish Audio's bet is that 15,000 natural language controls span both without requiring different models — that the same expressive capacity a video game designer needs is what makes an enterprise voice feel trustworthy rather than robotic.
The open-source distribution model generates a dynamic that closed enterprise sales does not. A developer who encounters Fish Audio's model while building a personal project will, if they join a company that needs voice infrastructure, already have an opinion about the vendor. 31,000 GitHub stars represent a pre-qualified sales pipeline that required no marketing spend to build. The developer who open-sourced their way to $21M ARR understands this intuitively; it is why the hosted product and the open-source repository are maintained in parallel rather than the open-source version being deprecated once revenue materialized.
Character-as-a-service is the most direct extension the control surface enables. Gaming studios, audiobook publishers, and animation companies need consistent character voices that can be generated at scale without booking studio time. A Fish Audio API that assigns a persistent character identity — a fixed set of control parameters defining how a specific character sounds across any input, in any session — and maintains that identity over a long-running project eliminates the primary friction in AI-generated narrative content: voice inconsistency between scenes or chapters. The buyer is a creative production house that currently solves this problem with expensive session recording.
Voice localization is the second opening. An English-language product demo requires a different expressiveness profile for Japanese, French, and Brazilian Portuguese audiences — not just different phonemes, but different prosodic conventions, different norms for formality and pacing. The 15,000 controls Fish Audio has built are currently calibrated to English conversational norms. Extending that control vocabulary to other languages, trained on actual conversational recordings rather than studio audio, addresses a gap that no voice AI vendor has fully solved and that every multinational enterprise with AI-generated customer-facing audio encounters when localizing for new markets.