
Fish Audio
Fish Audio is an AI voice platform for creators and developers, generating speech, cloned voices and transcripts in 80+ languages with inline emotion tags, a 2,000,000-voice community library, a pay-as-you-go API and a permanent free tier.
What is Fish Audio?
Fish Audio is an AI voice platform published by Hanabi AI Inc., a Delaware company based in Dover. It gathers under one roof what usually takes several vendors: text-to-speech, voice cloning, speech-to-text, real-time voice conversion, audio separation, audio translation, sound effects and a chapter-based audiobook tool called Story Studio.
The engine is S2.1 Pro, the latest of five audio models shipped in a single year across the Fish Speech, S1, S2 and S2.1 Pro lineage — four for synthesis, one for recognition. It rests on a dual autoregressive architecture that the chief scientist originally bet could run on a single 4090 GPU. What sets it apart day to day is expressive control written straight into the script: emotion tags such as [angry], [sad], [whispering] or [excited], special sounds like [laughing], [sighing] or [pause], and more than 15,000 natural-language direction tags. The company reports an Audio Turing Test score of 0.515 across 581 head-to-head comparisons, and publishes both the methodology and the raw audio.
Coverage is broad: 80+ languages claimed on the enterprise pages (83+ in the July 2026 blog post), zero-shot cloning in 13 of them, and code-switching between English, Mandarin, Japanese, Spanish and Arabic. Cloning needs 10 seconds of reference audio, 15 for a ready voice and 30 for the best result, while the community Voice Library holds more than 2,000,000 voices. Latency is quoted at under 150 ms to first audio in the cloud and sub-300 ms end to end.
Distribution runs on two tracks. The hosted side offers a web app, REST and WebSocket APIs, Python, TypeScript and Go SDKs, an MCP server and a documented voice-agent platform with knowledge bases, webhooks, phone numbers, call transfers and Web and React SDKs. The open-weights side lets teams self-host fish-speech, S1 and S2 in a VPC, on-premise, air-gapped or sovereign cloud under a paid commercial licence.
Production users named on the site include HeyGen, Retell, Sanas, LiveKit, Dubbing AI and Pictoria. As of 27 July 2026 the company reported 8M+ users, 21M USD in ARR and 22 employees. A permanent free tier opens the door with 8,000 credits a month, up to seven minutes of generation and three public voice slots, no card required.
What it does
- Turn written text into voice-over audio, with inline tags controlling the delivery
- Clone a voice from 10 to 30 seconds of reference audio
- Transcribe audio with speaker separation, emotion tags and natural-language description
- Translate and dub audio while keeping the original voice
- Connect a real-time voice agent through the REST or WebSocket API
- Design a voice from a written description, or convert a voice in real time
- Produce an audiobook chapter by chapter in Story Studio, generate sound effects and split audio into stems
When to use Fish Audio / When not to
A quick filter to help you decide if Fish Audio is the right fit.
When to use Fish Audio
- Video creators and YouTubers who need voice-overs for explainers, tutorials or documentaries without booking a studio
- Audiobook and publishing teams producing chapter-by-chapter narration that has to meet ACX and Audible delivery specifications
- Developers building real-time voice agents, who depend on WebSocket streaming, first audio under 150 ms in the cloud and Python, TypeScript or Go SDKs
- Game studios and character-app makers looking for distinctive character voices, instant cloning and 15,000+ natural-language direction tags
- Multilingual teams working across Asia-Pacific, thanks to native-quality Mandarin, Japanese, Korean and Cantonese with code-switching
When not to use Fish Audio
- Businesses planning to monetise output from the free tier, which is restricted to personal, non-commercial use
- Procurement teams that require a signed standard DPA or a completed SOC 2 report, since no DPA is published and the Type II audit is still underway
- Organisations that must keep voice data inside the European Union, as cloud data stays in the United States by default and per-country residency runs through the self-hosted tier
- Cost-sensitive buyers with heavy monthly volume, where the Max plan reaches 749 USD per month billed annually and Enterprise starts at 999 USD per month
- Mobile-first users looking for an official app: the terms mention an iOS app, but no App Store or Google Play link is published on the site
How to use Fish Audio
A typical end-to-end flow, from setup to results.
- Try the demo widget on the home page, the text-to-speech page or the voice-cloning page, without creating an account
- Create an account when you are ready, using Google or GitHub sign-in if you prefer
- Open the web app, browse the Voice Library and pick a voice, or upload 10 to 30 seconds of reference audio to clone one
- Paste your script, add inline emotion and direction tags, then generate and download the audio
- For programmatic use, generate an API key from the developer section of the app
- Call the POST /v1/tts endpoint on api.fish.audio with an Authorization Bearer header and a JSON body carrying text, reference_id and format
- Set the model header to pick the engine, for example s2.1-pro-free
- Or install the Python or TypeScript SDK and stream chunks back from a TTS request
- Switch to the WebSocket endpoint when a voice agent needs real-time streaming
- Apply to the student or startup programmes, or contact sales for enterprise terms and self-hosting
Pros & Cons
Pros
- Fine-grained expressive control through inline tags, with no separate parameter schema to learn or migrate
- Flat, legible API pricing at 15 USD per million UTF-8 bytes from the first call to the billionth, with no per-seat fee
- A free tier that is permanent rather than a countdown, backed by an s2.1-pro-free API model billed at 0 USD
- Unusual depth in Asian languages, with native-quality Mandarin, Japanese, Korean and Cantonese and code-switching
- Cloning in seconds from 10 seconds of audio with no training queue, on top of a 2,000,000-voice library that often removes the need to clone at all
- Open weights that can be self-hosted in a VPC, on-premise, air-gapped or sovereign cloud, which keeps an exit route open
- Competitive latency for real-time voice agents, backed by named production references such as HeyGen, Retell, Sanas and Dubbing AI and by a published evaluation methodology
Cons
- The free tier excludes commercial use, so any monetised output requires a paid plan
- No DPA is published, and the SOC 2 Type II audit was still underway as of 23 August 2026
- No subprocessor list: the privacy policy names only Stripe and Google, and the enterprise FAQ adds Google Cloud and Cloudflare R2
- The privacy policy never states whether customer content is used to train the models, and no training opt-out is documented; Zero Data Retention is enterprise-contract only
- Data residency outside the United States is limited to the self-hosted premium offer, which starts at 10,000 USD in setup fees plus 10,000 USD per month
- No European legal presence and no Article 27 representative, and monthly credits expire instead of rolling over
- Figures shift from page to page — 80+ languages on the enterprise pages, 83+ on the blog, 8 in the home FAQ, 13 for zero-shot cloning — and the terms mention an iOS app with no published store link
Pricing & Plans
A permanent free tier is available at no cost, offering 8,000 credits per month, up to seven minutes of generation and three public voice slots, with no credit card required. The cheapest paid entry point is the Plus plan at 5.50 USD per month billed annually (66 USD per year), or 15 USD per month on monthly billing. The API is sold separately on a pay-as-you-go basis, with no subscription and no monthly minimum: 15.00 USD per million UTF-8 bytes on s2.1-pro, s2-pro and s1, 0.00 USD on s2.1-pro-free, and 0.36 USD per hour of audio for transcribe-1 speech recognition. All prices are quoted in USD.
- 8
- 000 credits per month
- up to 7 minutes of generation
- 500 characters per generation
- 3 public voices
- standard generation speed
- enhanced cloning
- 250
- 000 credits per month
- up to 200 minutes
- 15
- 000 characters per generation
- unlimited public voices plus 10 private ones
- priority generation
- Voice Design
- 2
- 000
- 000 credits per month
- up to 1
- 620 minutes
- 3 seats
- 30
- 000 characters per generation
- 25
- 000
- 000 credits per month
- up to 6
- 250 minutes
- 10 seats
- 15 professional voices
- pay-as-you-go with organisation-level controls
- Zero Data Retention
- on-premise deployment
- SOC 2 compliance work
- additional volume discounts and custom SSO announced
- from 10
- 000 USD in setup fees plus 10
- 000 USD per month on a 12-month commitment
- for VPC
- on-premise
- air-gapped or sovereign cloud deployment
- Billing can be switched between monthly and annual at any time
- taking effect at the next cycle
- annual billing is advertised as a 33% saving
Data, GDPR & hosting
A consolidated view of how Fish Audio handles your data.
GDPR overview
The acronym GDPR appears nowhere on the site and no compliance claim is made in so many words. The privacy policy does, however, carry a dedicated EEA, Switzerland and UK section that delivers in substance: a table maps each purpose to a legal basis (contract performance, legitimate interest, consent, legal obligation), and transfers outside the European Economic Area rely on appropriate safeguards such as standard data protection clauses adopted by the European Commission. The rights listed are the expected ones — confirmation and access, rectification, erasure, restriction, portability, withdrawal of consent, objection to legitimate-interest processing and to direct marketing, freedom from automated decisions, and complaint to a regulator — exercised by email to support@fish.audio, with possible identity verification. What is missing is structural: no Article 27 representative, no named DPO, no European entity or address.
Who owns the data?
Fish Audio treats everything you post or upload as a User Submission, and the terms state that the licence you grant is a licence only, leaving your ownership untouched. Hanabi AI Inc. receives the right to reproduce, translate and technically modify that content in order to operate the service. Scope depends on visibility: Personal, Limited Audience or Public. Public submissions widen the licence to other users and to Fish Audio's own marketing and promotion, and all these licences are described as royalty-free, perpetual, sublicensable, irrevocable and worldwide. Commercial rights over generated audio belong to paying subscribers only, and only for verified voices they own.
Reuse rights
The privacy policy sets out the purposes: providing, administering and analysing the service; improving it through testing, research, internal analytics and product development; personalisation; communications; legal obligations; and fraud or abuse detection. The terms add that automated systems may analyse User Submissions to detect infringement and abuse such as spam, malware and illegal content. Recipients are described by category — hosting and technology, analytics, support, payment processing, advertising and business partners — with only Stripe, Inc. and Google LLC named outright. Fish Audio states that it does not sell personal data and has not done so in the past twelve months, and it asks for consent before personalised advertising in the European Economic Area. Aggregated and anonymised data may be reused and shared where it no longer identifies anyone. One question is left open: the policy never addresses whether customer content is used to train the models, while the July 2026 funding post says post-training runs on real user preferences. Enterprise contracts can switch on Zero Data Retention, so that request text and audio are never written to disk.
Data retention & training
Hosting summary
By default, Fish Audio keeps customer data in the United States. The hosted platform runs on Google Cloud with Cloudflare R2 for object storage, and inference is served from edge regions in the United States and Asia-Pacific (Tokyo) through a Rust edge gateway fronting several GPU regions. First-audio latency under 150 ms is reported across US, EU and APAC regions, but that describes a serving footprint, not data residency. Country- or region-level residency is available only through the enterprise self-hosted tier, where the model runs inside the customer's own infrastructure: VPC, on-premise, air-gapped or sovereign cloud. For users in the European Economic Area, the privacy policy states that international transfers rely on standard data protection clauses adopted by the European Commission, and it acknowledges that affiliates, cloud storage providers, IT vendors and data centres involved in processing are in some cases located abroad. Users may object to such transfers by writing to support@fish.audio, unless the transfer is necessary to deliver the service. No subprocessor list is published.
Where Fish Audio works
Country-level availability.
Not available in
Things to keep in mind
Risks and trade-offs to weigh before adopting Fish Audio.
- Monetising a video made on the free tier falls outside the licence, which covers personal, non-commercial use only
- Public User Submissions may remain available after an account is deleted, under a licence described as perpetual and irrevocable and extending to the publisher's own marketing
- Compliance cannot yet be evidenced by a report: the SOC 2 Type II audit was unfinished as of 23 August 2026, and no DPA is published for GDPR-bound buyers to sign
- The company has never written down whether customer content is used to train its models, while the funding post mentions post-training on real user preferences
- Data stays in the United States by default, so any country-level residency requirement pushes you towards the premium self-hosted offer
- Convincing synthetic voices invite misuse: consent from the person being cloned is the user's responsibility, and impersonation or misleading audio is easier than it looks, so treat a cloned narration as a recording someone will one day be asked to account for
- The terms mention an iOS app while no store link is published and unrelated third-party apps share the name, and the open-weight models separate free non-commercial research use from paid commercial licensing, so check the scope before deploying
Setup & Integrations
Technical difficulty
Low on the interface side: nothing to install, the demo widget runs without an account, and cloning takes a 10 to 30 second upload with no training queue. Low to moderate on integration, since Fish Audio advertises signup to first audio in five minutes and a single POST to the text-to-speech endpoint with a bearer token is enough. Official Python and TypeScript SDKs, REST and WebSocket endpoints, an MCP server and ready-made integrations with Pipecat, LiveKit, n8n, Vapi, Twilio and Retell shorten the work. Self-hosting is the exception, being an engineering-assisted premium engagement.
Deployment
Integrations
Supported languages
Behind Fish Audio
Fundraising
Social
Resources
All the official URLs gathered for verification and reference.
Alternatives
Tools that compete with or complement Fish Audio.
Frequently asked questions
Which languages does Fish Audio support?
How much audio is needed to clone a voice?
Is there a free plan, and can it be used commercially?
How much does the API cost?
Do unused credits roll over to the next month?
Where is my data hosted?
What security certifications does Fish Audio hold?
Is there a mode where request data is not stored?
Can the models be self-hosted?
Are there free programmes for students or startups?
Should you pick Fish Audio?
Fish Audio makes a clear technical proposition: expressive, controllable AI voice that suits real-time production as readily as content creation. The inline tag system is its most distinctive trait, since direction is written into the script rather than dialled in through separate parameters. Around it sit a flat API price of 15 USD per million UTF-8 bytes, unusual depth in Mandarin, Japanese, Korean and Cantonese, open weights that can be self-hosted, and a free tier that is permanent rather than a countdown.
The reservations are not technical. Data governance trails the engineering: no DPA is published, the SOC 2 Type II audit is still underway, the privacy policy never says whether customer content trains the models, and European data residency is reachable only through a self-hosted contract starting at 10,000 USD in setup fees. Buyers working from a procurement checklist will have to negotiate rather than tick boxes.
Context helps in reading those gaps. As of 27 July 2026, Hanabi AI Inc. was a one-year-old company with 22 employees, 52M USD raised and 21M USD in ARR: young enough that the paperwork is visibly still catching up with the product, and funded enough that it plausibly will.
The practical split is simple. Creators, developers and voice-agent builders get a great deal for very little, between a large voice library, ten-second cloning, first audio under 150 ms and pricing with no seat tax. Regulated teams and enterprise buyers should treat it case by case, ask directly about model training and data processing agreements, and budget for the self-hosted route if residency is non-negotiable. The free tier is the fastest way to settle the quality question, and it costs nothing to run the test.
- Choosing a selection results in a full page refresh.
- Opens in a new window.