
ai-coustics
A real-time speech enhancement SDK that cleans, isolates and scores audio before it reaches a voice AI stack. It runs on the customer's own CPU infrastructure, cutting word error rates and false turn-taking triggers.
What is ai-coustics?
ai-coustics builds what it calls an audio intelligence layer: a set of models that sit between real-world sound and machine understanding, ahead of speech-to-text, the language model and text-to-speech. The company, ai-coustics GmbH, was founded in 2021 at the Technische Universität Berlin by Corvin Jaedicke and Fabian Seipel, both audio and machine learning engineers. The product ships as an SDK meant to run on the customer's infrastructure, with bindings for Python, Rust, Node.js, C, C++ and WebAssembly, on top of AirTen, an in-house CPU-first inference engine that needs neither a GPU nor an ONNX dependency.
Four model families cover distinct jobs. Quail Multi Speaker enhances all speech in far-field, multi-speaker rooms. Quail Voice Focus isolates the foreground speaker and suppresses everything else, shipping as a small 2.2 S variant and a higher-quality 2.2 L. VAD Multi Speaker and VAD Voice Focus handle voice activity detection without a separate denoiser. Tyto is the diagnostic piece, returning a single risk score per five-second window broken down across noise, reverberation, loudness, interfering speech, packet loss and codec degradation, with Good, Warn and Bad bands at 0.35 and 0.60. Rook Multi Speaker targets human listening rather than machines.
The headline claims are a relative word error rate reduction of up to 43 percent, up to 80 percent fewer false voice-activity triggers than a standard detector, and real-time inference under 30 milliseconds, though the site quotes latency inconsistently across pages. Training is said to span more than 500 noise types, over a million acoustic environments and data in more than 65 languages, the models themselves being described as language-agnostic.
Published customer cases include PolyAI, telli, Synthesia, Elgato and Phonely, and an independent benchmark run by SLNG compares the models against Krisp and NVIDIA Maxine. The company previously ran a consumer-facing web app for creators before refocusing on voice AI infrastructure.
What it does
- Strip background noise, competing voices and reverberation from live speech in real time
- Isolate the primary speaker and suppress everyone else in the room
- Lower the word error rate of downstream speech-to-text engines
- Detect voice activity robustly enough to keep turn-taking from breaking
- Score the audio risk of every call and name the dimension driving it
- Raise perceived audio quality for human listening with a dedicated model
- Run entirely on CPU inside the customer's own infrastructure, with no GPU
When to use ai-coustics / When not to
A quick filter to help you decide if ai-coustics is the right fit.
When to use ai-coustics
- Engineering teams running voice agents in production, where telephony compression, background chatter and packet loss break the pipeline long before the language model sees anything
- Developers already building on LiveKit or Pipecat, who can drop the enhancement in as a native plugin or a filter class rather than rebuilding an audio stage
- Companies with strict confidentiality constraints, since the SDK processes audio inside their own infrastructure and the vendor states it never receives it
- Embedded and hardware integrators shipping audio on CPU only, with no GPU and no ONNX dependency, as in the Elgato VST3 case published on the site
- High-volume operations weighing per-minute economics, with published tiers at 100,000, 300,000 and 500,000 minutes a month
When not to use ai-coustics
- Anyone looking for a transcription engine: the models lower word error rates upstream but produce no transcript of their own, the playground transcript view relying on a third party
- Teams after voice cloning or speech synthesis, which the vendor explicitly does not do, its stated position being to preserve the speaker's original timbre
- Non-technical users hoping for a ready-made app: there is no consumer desktop or mobile product, only a software development kit to be integrated
- Small projects on a tight budget, since the cheapest paid tier starts at 135 USD a month and the enterprise tier is announced from 2,000 USD a month
- Anyone chasing the cleanest possible audio for human ears through the Quail models, which are deliberately tuned for machine understanding and leave some ambience behind
How to use ai-coustics
A typical end-to-end flow, from setup to results.
- Create an account on the Developer Platform at developers.ai-coustics.com, which needs no credit card
- Generate an SDK key from the dashboard
- Try the models in the browser playground, comparing transcripts and word error rates with and without Voice Focus
- Read the benchmarks or the Hugging Face space if you want evidence before writing any code
- Install one of the SDK bindings, choosing between Python, Rust, Node.js, C, C++ and WebAssembly
- Let the application download the model files once from the vendor's artifact host, and verify the published hash on your side
- For LiveKit, add the native livekit-plugins-ai-coustics plugin, authenticating through LiveKit Cloud or with your own SDK key when self-hosting
- For Pipecat, pip install and drop the AICFilter class into the existing pipeline
- Optionally set AIC_SDK_OTEL_ENABLE=1 to export operational metrics over OpenTelemetry to your own endpoint
- Run in production free for 30 days from account creation, then subscribe to a monthly or annual plan through Stripe
Pros & Cons
Pros
- Audio is processed on the customer's own infrastructure and, per the privacy notice, never reaches the vendor
- Runs on CPU with no GPU and no ONNX dependency, which makes edge and embedded integration realistic
- Performance claims are backed by published figures, including a benchmark run by an independent third party
- Native drop-in integrations exist for LiveKit and Pipecat, shortening the path from evaluation to production
- The 30-day trial requires no credit card and allows unrestricted production use
- GDPR posture is unusually well documented: dated sub-processor list, retention table, DPA on request, primary hosting in Frankfurt
- No minimum contract period, with cancellation at any time on both monthly and annual billing
Cons
- There is no permanent free plan: once the 30 days are up, a subscription is mandatory
- The entry ticket is steep for a small project at 135 USD a month, with enterprise starting from 2,000 USD
- Unused minutes do not roll over from one period to the next
- Only technical teams can use it, since nothing ships as a finished application for an end user
- Latency figures contradict each other across the site, quoted as under 10 ms, 30 ms and sub-40 ms
- No security certification such as SOC 2 or ISO 27001 is claimed anywhere on the site
- Baseline support runs through a community Discord, with priority email support starting only at the Pro tier
Pricing & Plans
There is no permanent free plan. A 30-day free trial is available from account creation, with no credit card required and production use permitted. The lowest paid entry point is the Startup plan at 135 USD per month billed annually, that is 1,620 USD per year, which includes 100,000 processed minutes per month. Annual billing carries a 10 percent discount over monthly billing, and usage is metered on the duration of audio actually processed.
- 100
- 000 minutes a month
- all core SDK models
- real-time processing under 30 ms
- 100+ languages
- on-prem SDK deployment
- community Discord support
- 300
- 000 minutes a month
- same inclusions
- priority email support
- 500
- 000 minutes a month
- custom benchmarks
- priority email support and a dedicated Slack channel
- everything in Business plus custom SLAs
- dedicated engineering support
- custom audio evaluations
- white glove onboarding
- procurement and compliance support
- and offline or air-gapped licence options
Data, GDPR & hosting
A consolidated view of how ai-coustics handles your data.
GDPR overview
Implementation is detailed and specific. The privacy notice, version 4.0 effective 1 September 2026, names a legal basis article by article, distinguishes the vendor's controller and processor roles, and lists every data subject right with its article number: access (15), rectification (16), erasure (17), restriction (18), portability (20), objection (21) and withdrawal of consent (7(3)), with a one-month response time. A dated sub-processor list is published separately, transfers outside the EEA rely on Standard Contractual Clauses with transfer impact assessments on file, and a data processing agreement is offered on request. No data protection officer has been appointed, which the vendor states plainly. The supervisory authority named is the Berlin commissioner. No Article 27 representative is designated, and none is required for a company established in Berlin.
Who owns the data?
Customers keep their audio. The vendor states that anything the SDK processes stays on the customer's own infrastructure and that ai-coustics does not receive it, store it or have access to it. It positions itself as a controller only for account, developer portal, billing and SDK telemetry data, and as a processor for any personal data carried in audio a customer application handles with the SDK. Two exceptions are named: the playground transcript view sends audio to Soniox, and the call analysis demo sends uploads to Modal, both processed in transit and, according to the notice, never written to disk.
Reuse rights
The terms grant customers no reuse rights over anyone else's material, and the vendor claims none over customer audio. What ai-coustics does use is licensing and metering data: the key identifier or short-lived token, SDK version and wrapper type, model identifier, operating system, CPU architecture, a session identifier tied to the account, per-session processing durations and network metadata including the source IP. The stated purposes are license authorisation, usage metering for billing and reliability diagnostics, resting on Article 6(1)(b) and 6(1)(f). Account, onboarding and marketing attribution data serve the account relationship. Product analytics and session replay run through PostHog under consent. The notice states that no personal data is sold and no legally significant automated decision-making is used. Nothing describes customer audio being used to train models.
Data retention & training
Hosting summary
The vendor's own infrastructure runs on AWS in eu-central-1, Frankfurt, Germany, and application logs and metrics stay inside that environment. Several sub-processors are also within the EEA: Usercentrics in Germany, PostHog on its Frankfurt EU Cloud, and HubSpot in its EU data region. Others sit in the United States under Standard Contractual Clauses, namely Clerk, Metronome, Soniox, Modal, Mintlify, Framer and Vimeo, while Stripe operates from the United States and Ireland and Cloudflare from a global edge network. YouTube loads from the United States only after the visitor consents. Transfer impact assessments are said to be on file and a copy of the clauses can be requested. The essential point for most buyers is that audio handled by the SDK is hosted nowhere on the vendor's side, since it never leaves the customer's own infrastructure.
Things to keep in mind
Risks and trade-offs to weigh before adopting ai-coustics.
- The free trial is capped at 30 days from account creation rather than by usage, so an evaluation that stalls internally can quietly burn the whole window
- Unused minutes expire at the end of each period, which punishes teams whose volumes fluctuate
- Latency claims are inconsistent across the site, quoted as under 10 ms, 30 ms and sub-40 ms, so verify the figure for the model you actually deploy
- Language coverage is claimed at 65, 100 and 150 languages depending on the page, and no list of languages is ever published
- The public Tyto demos send uploaded audio to a US provider, unlike the SDK: never upload a call recording you do not have the right to share
- SDK telemetry still reports your source IP and session identifiers even though the audio itself stays put, which is worth stating plainly in your own privacy documentation
- Cleaner input makes a voice agent look more reliable than it is, so treat the enhancement as an audio fix and keep measuring the failures it does not solve
Setup & Integrations
Technical difficulty
Moderate, and strictly for developers. Integration means installing an SDK binding and modifying an audio pipeline, so there is no no-code path to production. The shortest routes are genuinely short: a native plugin for LiveKit, or a pip install and a filter class for Pipecat. No GPU or ONNX runtime is needed. Model files download once and verifying their published hash is the customer's responsibility. Anyone can evaluate the models beforehand without writing code, using the browser playground or the public Hugging Face space.
Deployment
Integrations
Supported languages
Behind ai-coustics
Fundraising
Social
Resources
All the official URLs gathered for verification and reference.
Alternatives
Tools that compete with or complement ai-coustics.
Frequently asked questions
Can I test ai-coustics before paying?
Does my audio leave my own infrastructure?
How is usage counted?
What happens if I exceed my monthly minutes?
Which programming languages does the SDK support?
Does it work with LiveKit and Pipecat?
How does the Tyto risk score work?
Which languages do the models handle?
Is a data processing agreement available?
Is there a minimum age or a minimum commitment?
Should you pick ai-coustics?
ai-coustics is infrastructure, not a product you open and use. It addresses a narrow and genuinely awkward problem: voice agents that behave impeccably in testing and fall apart on real calls, because the audio reaching them is nothing like the audio they were demonstrated on. The answer offered is a CPU-only SDK that cleans, isolates and scores that audio before anything else in the pipeline sees it.
Two things stand out. The first is the deployment model: audio is processed on the customer's own machines and, according to the privacy notice, never reaches the vendor, which turns a compliance conversation that usually drags on into a short one. The second is the evidence. The company publishes word error rate figures across eight commercial speech-to-text providers, names the competitors it measures itself against, and points to a benchmark run by a third party rather than by itself. That is more transparency than this category usually offers.
The reservations are commercial rather than technical. There is no permanent free tier, and at 135 USD a month the entry point rules out hobby projects and early prototypes. Unused minutes expire. Nothing here is usable without an engineer. And for a vendor selling to enterprises on a confidentiality argument, the absence of any claimed security certification is a gap a procurement team will notice.
For an engineering team already losing calls to bad audio, the free 30-day production trial makes the evaluation cheap and the decision empirical: measure your own word error rate with and without it. For everyone else, this is simply not the layer of the stack you are shopping in.
- Choosing a selection results in a full page refresh.
- Opens in a new window.