Deepgram
Deepgram is a voice AI platform whose APIs handle speech-to-text, text-to-speech, voice agents and audio intelligence. Built for developers and enterprises, it runs in real time or batch, in the cloud or self-hosted.
What is Deepgram?
Deepgram is a voice AI platform built by Deepgram, Inc., a San Francisco company founded in 2015. It sells infrastructure rather than a finished application: four families of APIs that developers call directly from their own products. Speech-to-text is the historical core, served by in-house models. Flux handles real-time conversation, with built-in turn detection and interruption handling across ten languages, while Nova-3 covers production transcription in more than fifty. Industry-tuned variants target healthcare, legal and finance vocabulary, and fully custom models can be trained on proprietary datasets.
Text-to-speech runs on Aura-2, which offers more than forty English voices and advertises latency below 200 milliseconds. The Voice Agent API is the piece that ties the platform together. Rather than stitching a transcription service, a language model and a speech synthesiser into a pipeline, a single endpoint handles all three, plus barge-in, turn-taking prediction, function calling and mid-session control. Teams that already have a preferred language model or voice provider can bring their own and keep Deepgram's orchestration, at a reduced rate. A fourth family, Audio Intelligence, pulls summaries, topics, sentiment and intent out of conversations.
The platform is unusually flexible about where it runs. Beyond the shared cloud there is a single-tenant Dedicated runtime, deployment inside a customer VPC, and full self-hosting on AWS, GCP or a customer's own data centre, the last aimed explicitly at teams that already have DevOps resources. A dedicated EU endpoint exists for organisations that need processing to stay inside the European Union. Deepgram also ships an installable voice assistant, Saga, named in its terms of service.
The commercial posture matches the technical one. Per-minute rates are published rather than negotiated, there is no minimum commitment, and new accounts receive USD 200 of credit without a card. Compliance is documented, with SOC 2 Type I and II, HIPAA business associate agreements, PCI, CCPA and GDPR, and a full subprocessor list is public. The clause to read closely is the content licence: by default customer audio may be used to improve Deepgram's models, with an opt-out set per API request.
What it does
- Transcribe speech in real time or from recordings, in fifty-two languages with automatic language detection
- Build a complete voice agent through one API that unifies recognition, language-model orchestration and synthesis
- Generate spoken audio from text with more than forty English voices and sub-200 ms latency
- Extract meaning from conversations: summaries, topics, sentiment and intent
- Label who spoke when, and format punctuation, casing, dates and currency automatically
- Redact sensitive personal data such as card numbers, phone numbers and social security numbers
- Deploy in shared cloud, single-tenant, inside a VPC or fully self-hosted on your own infrastructure
When to use Deepgram / When not to
A quick filter to help you decide if Deepgram is the right fit.
When to use Deepgram
- Engineering teams adding speech to an existing product, who want one API instead of a stitched-together transcription, language-model and synthesis pipeline
- Contact centre and CPaaS platforms that embed voice recognition in their own offering, alongside the Five9, Genesys, Twilio, Vonage and AudioCodes integrations the site lists
- Builders of real-time voice agents, where sub-300 ms transcription, barge-in and turn-taking prediction decide whether a conversation feels natural
- Regulated organisations in healthcare, finance or the public sector that need single-tenant, VPC or fully self-hosted deployment, HIPAA business associate agreements and a published subprocessor list
- Early-stage startups burning through audio volume, who can apply for up to USD 100,000 of credits over twelve months through the startup programme
When not to use Deepgram
- Anyone looking for a mobile app: Deepgram publishes no iOS or Android application, and no store badge appears anywhere on the site
- Non-technical users who want a ready-made meeting recorder or a drag-and-drop transcription tool, since every feature is reached by writing code against an API
- Occasional or hobby use, because there is no permanent free tier and the cheapest committed plan opens at USD 4,000 per year
- Teams that need translation rather than recognition: the fifty-two languages cover speech-to-text, and nothing on the site offers translated output
- Buyers who require a named support mailbox and a contractual response time on a self-serve plan, as support runs through documentation, the community forum and Discord
How to use Deepgram
A typical end-to-end flow, from setup to results.
- Create a free account on the Deepgram console; no credit card is required and USD 200 of credit is granted at signup
- Or skip signup entirely and try the endpoints in the browser playground, which covers transcription, synthesis and the voice agent
- Generate an API key from the console
- Pick a model for the job: Flux for a real-time agent, Nova-3 for production transcription, an industry-tuned variant for specialised vocabulary
- Call the REST endpoint for recorded audio, or the WebSocket endpoint for streaming
- Switch on the options you need: diarisation, smart formatting, keyterm prompting, entity detection or PII redaction, each billed separately
- For a conversational agent, connect to the single Voice Agent endpoint instead of assembling three services, and bring your own language model or voice if you prefer
- Follow the developer documentation and changelog for SDKs, self-hosting guides and model updates
- Request access to the EU endpoint through the dedicated form if data must stay inside the Union
- Move to the Growth plan through the console checkout, or contact sales for Enterprise, custom models and dedicated deployment
Pros & Cons
Pros
- One API covers the whole voice chain, which removes the latency and failure modes that come from wiring three vendors together
- Latency figures are stated rather than implied: under 300 ms for transcription, under 200 ms for synthesis
- Usage rates are published per model and per option, so a cost estimate can be built without talking to a salesperson
- USD 200 of credit at signup, with no card and no expiry, makes serious evaluation possible before any commitment
- The deployment range is genuinely rare, from shared cloud through single-tenant and VPC to self-hosting on your own hardware
- Compliance is documented and third-party audited: SOC 2 Type I and II, HIPAA with business associate agreements, PCI, CCPA and GDPR
- The training opt-out is explicitly documented and actionable, and the subprocessor list names each vendor and the data it touches
Cons
- The default content licence is very broad: irrevocable, perpetual, sublicensable through multiple tiers and surviving termination
- Published rates assume enrolment in the Model Improvement Program, so opting out of training may change the economics
- The API training opt-out is set request by request rather than once at account level, which shifts the burden onto every integration
- The privacy notice is dated 26 October 2021 while the terms were refreshed on 6 August 2026, a gap of more than four years
- No permanent free plan, and the first committed tier opens at USD 4,000 per year
- No support email address is published; help runs through documentation, the community forum and Discord
- Data is stored on servers in the United States, and the EU endpoint is described as generally available on the pricing page but as early access on the enterprise page
Pricing & Plans
There is no permanent free plan. New accounts receive USD 200 of credit at signup, with no credit card required and no stated expiry, after which usage is billed as it is consumed. The lowest published entry point for speech-to-text is USD 0.0042 per minute with Nova-3 monolingual on the Growth plan, and USD 0.0048 per minute on pay-as-you-go. Text-to-speech starts at USD 0.0150 per 1,000 characters with Aura-1, and the Voice Agent API at USD 0.050 per minute when both the language model and the voice are supplied by the customer. The cheapest committed plan, Growth, starts at USD 4,000 per year in prepaid credits. Enterprise pricing is quoted on request.
- Pay As You Go — USD 200 of credit at signup then billing on consumption
- with no minimum
- no expiry and no card required
- aimed at developers and startups
- concurrency up to 50 REST and 150 WebSocket for speech-to-text
- 45 for text-to-speech
- 45 for the Voice Agent API and 10 for Audio Intelligence
- community and Discord support
- Growth — from USD 4
- 000 per year in prepaid credits
- saving up to 20 percent
- aimed at growing applications
- concurrency up to 50 REST and 225 WebSocket for speech-to-text
- 60 for text-to-speech
- 60 for the Voice Agent API and 10 for Audio Intelligence
- community and Discord support
- Enterprise — quoted on request
- for large volumes
- specific data or deployment requirements and dedicated support needs
- opens access to custom models
- HIPAA business associate agreements and Dedicated or self-hosted deployment
- Startup programme — up to USD 100
- 000 in credits over twelve months
- granted on application
Data, GDPR & hosting
A consolidated view of how Deepgram handles your data.
GDPR overview
The GDPR is addressed explicitly rather than in passing. The privacy notice places Deepgram in the role of processor under both the EU and UK GDPR, with customers acting as controllers, states that data processing agreements are signed with customers, and confirms that the European Commission's Standard Contractual Clauses are frequently used to legitimise transfers of customer data to the United States. Requests from individuals exercising their rights are routed back to the customer controller. The pricing page goes further, claiming full GDPR compliance and advertising a dedicated EU endpoint so that processing can stay inside the Union. Two gaps are worth noting: no Article 27 EU representative and no data protection officer is named on any page reviewed, and the privacy notice is dated 26 October 2021 while the terms of service were refreshed on 6 August 2026.
Who owns the data?
Deepgram's terms are explicit on the point: the company does not claim ownership of customer content. Inputs submitted to the service and the outputs it returns are together defined as Your Content, and the customer keeps whatever right, title and interest they already held. Two limits qualify that. Outputs are not guaranteed to be unique, because other customers may generate similar results from similar inputs. And nothing in an output conveys any right over the underlying models, algorithms, neural networks, weights or parameters, which remain Deepgram's property. Under the GDPR the company positions itself as a processor acting on the customer's instructions, the customer being the controller.
Reuse rights
Customers may reuse outputs without asking permission: they keep their rights and no further authorisation is required. In exchange they carry the responsibility, both for evaluating whether an output is accurate and appropriate for their use case, and for holding the rights needed over whatever they submit. The reverse direction is where the terms bite. Section 3.2 grants Deepgram an irrevocable, perpetual, transferable, sublicensable, worldwide and royalty-free licence over customer content, covering operation of the service, its improvement, the development of other products, and the training and testing of Deepgram's models. That licence survives expiry or termination of the agreement. An opt-out exists and is documented: on the APIs it is set per request through a parameter, so it applies to that request alone rather than to the account; Saga users find the choice in their settings. One consequence deserves attention before any budget is signed off, because the published rates assume enrolment in the Model Improvement Program.
Data retention & training
Hosting summary
The privacy notice is direct on jurisdiction: data is stored on servers in the United States. That places customer content under US law by default, which is why the notice also states that Standard Contractual Clauses are frequently used to legitimise transfers out of the European Economic Area, the United Kingdom and Switzerland. For organisations that cannot accept a transfer, Deepgram advertises a dedicated EU endpoint so processing stays inside the Union, although the enterprise page still describes the EU-hosted transcription endpoint as early access, available on request through a form. Beyond the shared cloud, three alternatives exist: a single-tenant Dedicated runtime with region-specific deployment, deployment inside a customer VPC, and full self-hosting on AWS, GCP or a customer's own data centre, which puts the jurisdiction question back in the customer's hands. Amazon Web Services is the principal infrastructure subprocessor, with Cloudflare listed for audio, signup information and interactive text.
Where Deepgram works
Country-level availability.
Not available in
Things to keep in mind
Risks and trade-offs to weigh before adopting Deepgram.
- The content licence granted in section 3.2 is irrevocable, perpetual, transferable, sublicensable through multiple tiers and survives termination; anyone processing confidential or client-owned audio should read it before integrating
- Published rates assume enrolment in the Model Improvement Program, so the cheapest path is also the one that feeds your audio into model training
- The API training opt-out applies per request, not per account, which means a single forgotten parameter in one code path quietly re-enables training
- Voice synthesis at this quality invites misuse: synthetic speech can impersonate, and responsibility for what an agent says to a customer sits with the integrator, not the vendor
- Automating conversations erodes the human feedback loop: teams that stop listening to their own calls lose the intuition that told them what customers actually struggle with
- Transcription accuracy is never uniform across accents, dialects and noisy environments, and treating a transcript as ground truth in a medical, legal or disciplinary context can cause real harm
- Arbitration is mandatory with a class-action waiver, and the thirty-day window to opt out by email closes quickly after you first accept the terms
Setup & Integrations
Technical difficulty
Moderate, and strictly for people who write code. Signup is self-serve, an API key is issued from the console, and the browser playground lets you hear the models work before writing anything. From there a first call is a straightforward REST or WebSocket request, and the voice agent needs one endpoint rather than three services wired together. Difficulty rises sharply at the other end of the range: self-hosting is described by Deepgram itself as built for teams that already have DevOps resources, and custom models and dedicated deployment go through sales.
Deployment
Integrations
Supported languages
Behind Deepgram
Fundraising
Social
Resources
All the official URLs gathered for verification and reference.
Frequently asked questions
What does Deepgram actually sell?
Is there a free plan?
How much does transcription cost?
How many languages are supported?
Can I run Deepgram on my own infrastructure?
Is my audio used to train Deepgram's models?
Where is the data hosted?
Which compliance frameworks does Deepgram meet?
Is a data processing agreement available?
What is the minimum age to use the service?
Should you pick Deepgram?
Deepgram is one of the more coherent voice AI platforms on the market, and its coherence is the point. Where most teams assemble transcription, a language model and speech synthesis from three vendors, Deepgram offers the whole chain behind a single endpoint, with the latency figures such a stack is supposed to deliver: under 300 milliseconds for recognition, under 200 for synthesis. Ten years of work on its own models shows in the range, from Flux for live conversation to Nova-3 across more than fifty languages and industry-tuned variants for healthcare, legal and finance.
The commercial posture is refreshingly plain. Rates are published per model and per option, there is no minimum commitment, and USD 200 of credit lets a team evaluate seriously before committing. The deployment range is the real differentiator for regulated buyers: shared cloud, single-tenant, VPC or full self-hosting on your own hardware, backed by SOC 2 Type II, HIPAA business associate agreements and a published subprocessor list.
Two reservations deserve weight. The default content licence is broad — irrevocable, perpetual, sublicensable and surviving termination — and the published prices assume enrolment in the model improvement programme, so opting out of training is a commercial decision as much as a technical one. That opt-out is also set request by request on the APIs rather than once at account level. The privacy notice, dated October 2021 against terms refreshed in August 2026, has not kept pace with the rest.
For an engineering team building voice into a product, this is a strong default choice. For anyone wanting a finished transcription tool, a mobile app or a permanent free tier, it is the wrong shelf entirely. Read section 3 of the terms before signing.
- Choosing a selection results in a full page refresh.
- Opens in a new window.