
FlowSpeech
FlowSpeech is a context-aware text-to-speech studio that reads a script's meaning before voicing it. Inline bracket tags direct emotion, accent and timing, and the same account also covers sound effects, voice changing and video dubbing.
What is FlowSpeech?
FlowSpeech is a browser-based text-to-speech studio built around one claim: the engine reads for meaning before it reads aloud. It analyses the sentiment, timing and nuance of a script, then voices it with prosody, breaths and pacing meant to pass for a person rather than a reader.
Three generation modes cover most needs. Single Speaker handles monologues and can auto-tag a whole document after analysing its tone. Multi Speaker detects who says what, splits the script, and keeps one voice attached to each recurring character across up to ten roles, returning the exchange as a single continuous track instead of clips to stitch together. Instant Speech is the quick path.
Direction happens inside the text. Typing an opening bracket opens a command palette, and the cue that follows steers the line: [whisper], [excited], [calmly and clearly], [strong British accent], or something written from scratch such as [relieved after a long wait]. There is no fixed preset list. Pause tags such as [1.0s] place a beat exactly where it belongs, which removes a round trip through a digital audio workstation for simple pacing work. An Auto Emotion switch will tag a long script on its own, leaving every suggestion editable before generation.
The catalogue runs to 30 voices in four registers — serious news, energetic marketing, warm narration and expressive character — with 70+ languages claimed, though the site never lists them. A single render accepts up to 200,000 characters, so a chapter need not be cut up. Input can be pasted or uploaded as PDF, DOC, DOCX, PPT, PPTX, TXT, RTF, EPUB or an image.
Three companion tools share the same account: a sound-effect generator driven by a written description with a duration you set, a voice changer offering 20 target voices plus background-noise removal, and an AI dubbing tool that carries a video into another language with lip-sync while preserving the speaker's identity.
Everything runs in the browser. There is no mobile app, no extension and no documented public API. The publisher, Zhidian Jump Technology Co., Ltd., registered the domain in December 2025 and the first web archive capture dates from January 2026, so this is a young service.
What it does
- Turn a script, document or image into human-sounding speech across 30 voices and 70+ languages
- Direct emotion, accent and delivery with plain-language tags typed straight into the text
- Time every beat with pause tags measured to a tenth of a second
- Split a dialogue across up to 10 speakers and render it as one continuous track
- Generate sound effects from a written description, with a duration you choose
- Swap the voice identity in an existing recording while keeping its emotion and pacing
- Dub a video into another language with lip-sync and the original speaker's identity preserved
When to use FlowSpeech / When not to
A quick filter to help you decide if FlowSpeech is the right fit.
When to use FlowSpeech
- Audiobook producers turning novels, textbooks and long articles into narrated audio, with up to 200,000 characters handled in a single render
- Podcast and audio-drama makers who need several recognisable voices in one continuous track rather than a pile of separate clips
- Video creators and YouTubers who want voiceovers with directed emotion, and who also need to dub the same upload into another language
- Instructional designers and e-learning developers scripting courses, onboarding modules and role-play scenarios for training
- Game and animation teams prototyping character lines, quest dialogue and cutscenes before committing to a final voice cast
When not to use FlowSpeech
- Anyone who needs a clone of a specific person's voice: custom voice cloning is not available and appears only on the roadmap
- Developers wanting to call the engine from their own product, since no public API documentation exists on the site
- Teams handling regulated data, as the terms explicitly exclude HIPAA, FISMA and GLBA use cases
- Procurement and compliance functions needing a processing agreement, a named subprocessor list or an Article 27 EU representative, none of which are published
- Users after live voice changing during calls or gaming chat: the voice changer works on uploaded files, not on a live stream
How to use FlowSpeech
A typical end-to-end flow, from setup to results.
- Open flowspeech.io: a first generation runs without an account, on a smaller quota
- Pick a generation mode, Single Speaker for a monologue, Multi Speaker for a conversation or Instant Speech for a quick result
- Paste your script, or upload a PDF, DOC, DOCX, PPT, PPTX, TXT, RTF, EPUB or image file and let the text be extracted
- In Multi Speaker, check the proposed split, then rename roles, add a missing speaker or reorder the exchange
- Type an opening bracket to call up the command palette and insert emotion, accent or delivery cues
- Add pause tags such as [1.0s] wherever the pacing needs a beat
- Turn on Auto Emotion for a long script if you would rather review suggested tags than write them
- Choose a voice among the 30 available, or cast one voice per role in a dialogue
- Generate, listen to the result, and refine any line that lands too strong or too flat
- Download the finished audio, and find earlier renders again under My Creations
Pros & Cons
Pros
- Fine control over delivery without leaving the editor: emotion, accent, pacing and pauses all live in the script
- Cues are written in ordinary language rather than picked from a closed list of presets
- 200,000 characters per render means long-form work does not have to be chopped into pieces
- Multi-voice dialogue comes back as one track, sparing the assembly a clip-by-clip workflow demands
- A permanent free tier that runs even without an account, rather than a countdown trial
- Broad file ingestion, images included, so source material rarely needs converting first
- Ownership and commercial use of the generated audio are granted explicitly
Cons
- No custom voice cloning: the FAQ places it on the roadmap, not in the product
- No documented public API, despite 'API Keys' labels sitting in the account panel's interface strings
- The 70+ languages are never listed, and the interface itself exists in only three of them
- The publisher is identified by a single header line on the policies, with no postal address anywhere on the site
- One mailbox answers support, legal and privacy alike, and no processing agreement, subprocessor list or Article 27 representative is published
- The third-party AI platform that processes submitted content is acknowledged but never named
- Template leftovers undercut the polish: the dubbing page is headed 'Video Translator', the copyright reads 2024, and the social links in the markup were never configured
Pricing & Plans
A permanent free plan is available and requires no account: guests receive 5,000 credits per month with a 5,000-character ceiling per request, while signed-in users receive 10,000 credits with a 10,000-character ceiling. The cheapest paid entry point is the Basic plan at USD 12.00 per month on an annual commitment, or USD 15.00 per month billed monthly. Annual billing is advertised at a 33% saving across the range, and all payments are stated to be in US dollars.
- Free — USD 0 per month
- guests get 5
- 000 credits monthly and up to 5
- 000 characters per request
- signed-in users get 10
- 000 credits and up to 10
- 000 characters per request
- Basic — USD 15 per month billed monthly
- or USD 12 per month on an annual commitment
- 200
- 000 credits monthly
- up to 200
- 000 characters per request
- 30+ voices
- Pro — USD 45 per month billed monthly
- or USD 39 per month on an annual commitment
- 1
- 000
- 000 credits monthly
- up to 200
- 000 characters per request
- 30+ voices
- Scale — USD 159 per month billed monthly
- or USD 129 per month on an annual commitment
- 4
- 000
- 000 credits monthly
- up to 200
- 000 characters per request
- 30+ voices
Data, GDPR & hosting
A consolidated view of how FlowSpeech handles your data.
GDPR overview
The privacy notice, last updated 15 December 2025, addresses the GDPR substantively without ever claiming compliance in so many words. It sets out the legal bases relied on under the GDPR and UK GDPR — consent, legal obligations and vital interests — and lists the rights available in the EEA, the UK and Canada: access, a copy, rectification, erasure, restriction, portability and objection. Consent can be withdrawn at any time. Complaints may be taken to a member-state authority, the UK ICO or the Swiss Federal Data Protection and Information Commissioner, and requests go through a Termly data subject access form or by email. What is missing is equally clear: no Article 27 EU representative is designated, no data protection officer is named, no processing agreement is offered, and one mailbox serves support, legal and privacy alike.
Who owns the data?
The terms are explicit about outputs: you keep full ownership of every audio file you generate, and the publisher states that it asserts no ownership over your contributions. Two carve-outs sit alongside that. Anything sent as a suggestion, comment or piece of feedback is assigned to the publisher, which may use and circulate it freely without credit or payment. Separately, the site's own content, software and marks remain with the publisher and are licensed to you for personal, non-commercial use only. Submitted scripts are not retained: the privacy notice says they are passed to an unnamed third-party AI platform purely for processing.
Reuse rights
Generated audio can be reused without asking permission. FlowSpeech states that you retain full ownership of the files and may use them commercially, including in monetised videos, podcasts, social advertising and other professional projects, provided the use complies with local law. No attribution is required and no separate licence has to be requested. The restriction sits elsewhere: the terms reserve the site's own content, software, designs and trademarks to the publisher and grant only a personal, non-commercial licence over them. In short, what you make is yours; what the platform itself is made of is not.
Data retention & training
Hosting summary
The terms state plainly that the Services are hosted in the United States, and that continuing to use them from any other region amounts to consenting to a transfer of personal data there for processing. No further detail is published: no region, no data-centre location, no named cloud provider and no option to choose where data resides. The domain resolves to a Cloudflare anycast address, which identifies the content delivery network in front of the site rather than the place anything is stored, so it should not be read as evidence of a hosting location. For a user in the European Economic Area this matters: the privacy notice sets out GDPR rights and legal bases, but names no Article 27 representative and offers no processing agreement, so the transfer rests on consent given through use rather than on documented safeguards. The publisher does state that submitted content is not saved or stored, which limits what sits in the United States to account data — names and email addresses — rather than the scripts themselves.
Things to keep in mind
Risks and trade-offs to weigh before adopting FlowSpeech.
- Submitted text is handed to a third-party AI platform the privacy notice never names, and that same notice asks you not to enter private information: treat any confidential script as exposed
- The publisher is barely identifiable, with one company name on the policies, no address, no named officer and a domain registrant hidden behind an Icelandic privacy proxy
- The domain was registered in December 2025 and expires in December 2026; a service this young may not still be there when a project needs regenerating
- Convincing synthetic emotion invites misuse, whether voicing words a real person never said or passing generated narration off as a recorded performance
- Compliance is thin for business use: no processing agreement, no subprocessor list, no Article 27 representative, and terms that exclude HIPAA, FISMA and GLBA scenarios
- Disputes run to arbitration in Brussels under Californian law, a combination that is expensive and impractical for an individual to pursue
- Leaning on auto-tagged delivery can dull the judgement that decides how a line should actually land, and the site's own guidance warns against over-directing
Setup & Integrations
Technical difficulty
Essentially none. FlowSpeech runs in a browser with nothing to install, no integration to configure and no API key to obtain; a first generation works without even creating an account, on a reduced quota. The only thing to learn is the bracket syntax for emotion and pause cues, and a command palette opens as soon as you type an opening bracket, so it is discovered rather than memorised. The companion FAQs state that neither the dubbing tool nor the sound-effect generator requires audio production experience. Expect a usable result within minutes of arriving.
Deployment
Supported languages
Behind FlowSpeech
Resources
All the official URLs gathered for verification and reference.
Frequently asked questions
How many voices and languages does FlowSpeech offer?
How do I add emotion, an accent or a pause?
Can I clone my own voice?
Can I use the generated audio commercially?
Is there a free plan?
How long can a single script be?
How many speakers can a dialogue contain?
What happens to the text I submit?
Is there an API or a mobile app?
Can I get a refund?
Should you pick FlowSpeech?
FlowSpeech makes a specific bet: that the useful part of text to speech is no longer the voice itself but the direction given to it. Typing a cue in plain language before a sentence, timing a beat to a tenth of a second, casting ten voices across one dialogue and getting a single finished track back — that is a workflow aimed at people who would otherwise be reassembling clips in an editor.
On that ground it delivers. Long-form work fits in one render, source files rarely need converting, the free tier runs without an account, and ownership of what you produce is granted without ambiguity. Audiobook narration, podcast dialogue, course modules and game prototyping all sit comfortably within reach.
The reservations concern the company, not the craft. Zhidian Jump Technology Co., Ltd. appears as a single header line on the policies and nowhere else: no postal address, no named officer, one mailbox for support, legal and privacy alike. Content you submit goes to a third-party AI platform the notice never identifies, and the same notice asks you not to type anything private. There is no processing agreement, no subprocessor list and no Article 27 representative. Loose ends elsewhere reinforce the impression of a service assembled quickly: a dubbing page still headed Video Translator, a 2024 copyright under policies dated December 2025, unconfigured social links in the markup, and terms naming PayPal while the refund policy follows Paddle.
The domain was registered in December 2025 and expires in December 2026. Treat FlowSpeech as a capable production tool for material you would be comfortable publishing anyway, try it on the free tier first, and hold back anything confidential or contractually sensitive until the publisher becomes easier to identify.
- Choosing a selection results in a full page refresh.
- Opens in a new window.