Deepgram
A developer voice AI platform with speech-to-text, text-to-speech and voice agents through unified, real-time APIs.
- Starting price
- Usage-based, pay-as-you-go
- Free trial
- $200 in free credits to start, no credit card required
- Commercial use
- Yes, usage-based
Overview
Deepgram is a voice AI platform for developers and product teams building speech features and conversational applications. It offers speech-to-text through its Nova models, text-to-speech through Aura, and a Voice Agent API that orchestrates listening, an LLM and speaking for real-time voice bots, all reached through unified APIs. It supports both real-time streaming and batch processing, speech-to-text across 45 or more languages with automatic language detection, and add-ons such as redaction, speaker diarization, keyterm prompting and audio intelligence like summarization and sentiment. It runs in the cloud or self-hosted for in-region and in-house data control, and carries SOC 2 Type 1 and 2, HIPAA, GDPR, CCPA and PCI compliance. Billing is usage-based with no minimums and $200 in free credits to start.
Best for developers building real-time speech-to-text, text-to-speech and voice agents through an API.
Deepgram suits teams building voice into their own product rather than producing audio in a studio: every capability is an endpoint, and the usage-based model rewards steady programmatic volume over occasional creative work. Prototyping costs nothing up front, and self-hosting is what makes it viable where recordings are not allowed to leave the building. Anyone who wants to clone a voice or dub finished video will need a second tool alongside it.
Key Facts
| Category | |
|---|---|
| Starting price | Usage-based, pay-as-you-go No minimums and no expiry. Speech-to-text from $0.0048 / minute on Nova-3 monolingual streaming, text-to-speech from $0.0150 / 1,000 characters on Aura-1, and the Voice Agent API from $0.056 / minute. |
| Free trial | $200 in free credits to start, no credit card required |
| Commercial use | Yes, usage-based |
| Output | Speech-to-text, text-to-speech and voice agents |
| API | Yes |
| Hosting | Cloud, or self-hosted for in-region and in-house data control |
Screenshots
Placeholders. They are replaced as soon as the vendor sends material, and both open the same short form.
Features
- Voice cloning
- No, the catalogue is fixed: 40+ prebuilt Aura-2 voices across English, Spanish, Dutch, French, German, Italian and Japanese, with no cloning or custom-voice product on the price list
- Real-time use
- Yes, real-time streaming speech-to-text and text-to-speech, plus batch processing
- Languages
- 45+ languages for speech-to-text, with automatic language detection
- Speech-to-speech and dubbing
- No dubbing or translation product. Speech-to-speech runs as a pipeline through the Voice Agent API, which handles listening, thinking and speaking with configurable speech-to-text models, LLM providers and voices, plus function calling mid-conversation
- Editing and controls
- Voice selection, plus encoding, bit rate, container and sample rate on the output, streamed audio and callbacks. Aura-2 is built to read naturally without SSML markup, and speed and style can be changed mid-conversation without restarting the session
- Consent and trust
- SOC 2 Type 1 and 2, HIPAA, GDPR, CCPA and PCI compliant; self-hosting for in-house data control
- API
- Yes, unified APIs for speech-to-text, text-to-speech and voice agents
Integrations
Integrations coming soon.
Deployment and Support
- Deployment
-
- Cloud (SaaS)
- Browser-based
- On-premise
- Support
-
- Community forum
- Onboarding
-
- Documentation and knowledge base
- Company size
-
- Small and mid-sized
- Enterprise
Suitability
| Industry | Freelancers | Small and mid-sized | Enterprise |
|---|---|---|---|
| Agencies | Not stated | Not stated | Not stated |
| Architecture & Engineering | Not stated | Not stated | Not stated |
| Consulting | Not stated | Not stated | Not stated |
| Media & Creative | Not stated | Stated as a fit | Stated as a fit |
| Technology | Not stated | Stated as a fit | Stated as a fit |
The table combines the industries the vendor addresses with the company sizes the tool is aimed at. An empty cell means the combination is not stated, not that it is ruled out.
Pros and Cons
Strengths
- Unified speech-to-text, text-to-speech and voice-agent APIs
- Real-time streaming with low latency for live applications
- Usage-based with $200 free credit and no card required
- Runs self-hosted for data control and compliance
Limitations
- Developer platform, not a consumer voice studio
- Several headline per-minute rates are promotional and sit below the list price, so the long-run cost depends on which rate applies
Pricing
-
Pay As You Go
$0.0048
Per minute of Nova-3 monolingual streaming speech-to-text at the current promotional rate, $0.0077 at list. Nova-3 multilingual $0.0058, Flux English $0.0065. Text-to-speech $0.0150 / 1,000 characters on Aura-1 and $0.030 on Aura-2. Voice Agent API from $0.056 / minute. Add-ons are priced separately: redaction and diarization $0.0020 / minute each, entity detection $0.0017, keyterm prompting $0.0013. No minimums, no expiry, $200 in credits to start without a card.
-
Growth
From $4,000
Per year in prepaid credits, redeemed against usage at lower rates: Nova-3 monolingual streaming $0.0042 / minute, Aura-1 $0.0135 and Aura-2 $0.027 / 1,000 characters.
-
Enterprise
Not stated
Custom pricing, volume commitments, self-hosting and dedicated support.
Best For
Alternatives and Comparisons
What users say
Reviews
-
0 reviews
- Ease of use
- Not rated yet
- User interface
- Not rated yet
- Onboarding
- Not rated yet
- Support and service
- Not rated yet
- Value for money
- Not rated yet
- Features
- Not rated yet
No reviews yet. Be the first to review this tool. Every review is checked for fairness before it appears, and the vendor cannot have one removed.
Share your experience
Write a Review
Last verified:
Ready to take a closer look at Deepgram
The link goes straight to the vendor. Some links on this site are affiliate links.
Visit Deepgram