Fish Audio banks $52M seed as 8M users fuel its quest for perfect AI voice
Curated by the Inblix editorial team
The race to build an AI voice that doesn’t sound like a robot just got more intense. Fish Audio, a Palo Alto startup founded by former Nvidia researcher Shijia Liao, has raised a $52 million seed round led by Coreline Ventures and Capital Today. The company already generates $21 million in annual recurring revenue and counts 8 million users across its open-source and hosted models, a staggering adoption rate for a project that began as a one-person, single-GPU experiment born out of frustration with synthetic voice quality.
What makes Fish Audio’s approach distinct is its library of over 15,000 natural language controls, a granular steering system built to serve two very different masters simultaneously. Creative users—indie developers, game designers, video creators—want expressiveness and character. Enterprises deploying AI avatars or voice agents want low latency, realism, and clinical steerability for customer support and sales automation. CEO and co-founder Rissa Cao frames the challenge bluntly: HeyGen wants realism, gaming studios want expressiveness, and LiveKit wants natural-sounding calls that don’t lag. Serving all three from the same underlying tech is ambitious, if not slightly delusional.
The startup’s growth hasn’t been frictionless. A few months ago, creators alleged their voices were uploaded to the platform without consent, exposing the messier side of community-sourced training data. Fish Audio says it’s since automated its DMCA takedown process to under three minutes and now requires a voice sample or contract for validation. Coreline Ventures partner Osuke Honda insists that trust requires consent, transparency, and attribution built into the product, not bolted on later. That’s the right sentiment, but an automated takedown is still a reactive band-aid—it doesn’t prevent unauthorized uploads in the first place.
With the new capital, Fish Audio plans to release an audio understanding model and a speech-to-speech model this year, escalating its fight against well-funded rivals like ElevenLabs, Cartesia, and Speechify. Investor Rico Mallozzi from 359 Capital points to the team’s technical efficiency as a differentiator against labs with deeper pockets. That’s the bullish take. The skeptical view is that a crowded market and unresolved consent questions could make those 15,000 controls feel less like a moat and more like a distraction.
💡 Key Takeaways
- Fish Audio's library of 15,000+ natural language controls is designed to serve both creative users needing expressiveness and enterprises demanding steerable, low-latency voices.
- The startup's automated takedown process now resolves voice ownership claims in under three minutes, but it does not proactively stop unauthorized uploads of a creator's voice.
- Generating $21M in ARR from 8 million users while competing against ElevenLabs signals genuine product-market fit, but the crowded field makes technical efficiency alone a fragile advantage.
- Upcoming audio understanding and speech-to-speech models suggest Fish Audio is pivoting from a pure generation tool toward a broader audio infrastructure play for enterprises.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.