CataleoSoftware Get Your Software Listed Get Listed

AI Voice & Speech Comparison

Deepgram vs. Speechmatics

At a glance

At a glance

Deepgram

Best for

Best for developers building real-time speech-to-text, text-to-speech and voice agents through an API.

  • Small and mid-sized
  • Enterprise

Speechmatics

Best for

Best for teams transcribing real conversations across accents and languages, including where the audio may not leave their own infrastructure.

  • Small and mid-sized
  • Enterprise

Cataleo does not name a winner. Both statements come from the vendors themselves.

Full Comparison

Criterion Deepgram Speechmatics
Starting price Usage-based, pay-as-you-go No minimums and no expiry. Speech-to-text from $0.0048 / minute on Nova-3 monolingual streaming, text-to-speech from $0.0150 / 1,000 characters on Aura-1, and the Voice Agent API from $0.056 / minute. From $0.129 / hour of audio Pro tier rate, billed to the second with no commitment. Volume discounts of 20% apply automatically above 500 hours per month for each speech-to-text type, with further discounts above 24,000 hours a year. One credit equals one dollar.
Free trial $200 in free credits to start, no credit card required Free tier with $100 in credit, no card required
Commercial use Yes, usage-based Yes, usage-based
Output Speech-to-text, text-to-speech and voice agents Transcripts, translations and synthesised speech
API Yes Yes, batch and real-time speech-to-text, text-to-speech, translation and a management API
Hosting Cloud, or self-hosted for in-region and in-house data control Cloud, CPU or GPU containers, Kubernetes, on-premises virtual appliance, or on-device
Made in Not stated United Kingdom
Voice cloning No, the catalogue is fixed: 40+ prebuilt Aura-2 voices across English, Spanish, Dutch, French, German, Italian and Japanese, with no cloning or custom-voice product on the price list No, the text-to-speech API ships four fixed voices, sarah and theo in British English and megan and jack in American English, and the documentation describes no custom or cloned voice; speed, pitch and emphasis are not adjustable either
Real-time use Yes, real-time streaming speech-to-text and text-to-speech, plus batch processing Yes, streaming transcription returned in under a second, alongside batch transcription of recorded files
Languages 45+ languages for speech-to-text, with automatic language detection 55+ languages and dialects with accent, dialect and code-switching coverage; translation across 69 language pairs
Speech-to-speech and dubbing No dubbing or translation product. Speech-to-speech runs as a pipeline through the Voice Agent API, which handles listening, thinking and speaking with configurable speech-to-text models, LLM providers and voices, plus function calling mid-conversation Audio translation across 69 language pairs, returned as translated transcript rather than dubbed audio; a Voice Agent API is in early access
Editing and controls Voice selection, plus encoding, bit rate, container and sample rate on the output, streamed audio and callbacks. Aura-2 is built to read naturally without SSML markup, and speed and style can be changed mid-conversation without restarting the session Speaker diarization and identification, custom dictionary, language identification, background speech filtering, and speech intelligence for summaries, sentiment, topics and auto-chapters
Consent and trust SOC 2 Type 1 and 2, HIPAA, GDPR, CCPA and PCI compliant; self-hosting for in-house data control Data is not logged as standard; container, Kubernetes, virtual appliance and on-device deployment keep audio inside your own infrastructure
MCP-ready stated Yes No
Deployment Cloud (SaaS), Browser-based, On-premise Cloud (SaaS), Browser-based, On-premise
Support Community forum Email helpdesk
Onboarding Documentation and knowledge base Documentation and knowledge base
Company size Small and mid-sized, Enterprise Small and mid-sized, Enterprise
Integrations Not stated LiveKit, Pipecat, Vapi, Zapier

Compare more tools

Row where the values differ The small markers point at the lower number, the longer list or the stated feature. They describe the data, they are not a verdict.

The short version

Key Differences

  • Starting price Deepgram Usage-based, pay-as-you-go Speechmatics From $0.129 / hour of audio
  • Free trial Deepgram $200 in free credits to start, no credit card required Speechmatics Free tier with $100 in credit, no card required
  • Hosting Deepgram Cloud, or self-hosted for in-region and in-house data control Speechmatics Cloud, CPU or GPU containers, Kubernetes, on-premises virtual appliance, or on-device
  • Support Deepgram Community forum Speechmatics Email helpdesk
  • MCP-ready Deepgram Yes Speechmatics No

Every line is one datapoint from the table above, picked automatically. Nothing here is written text.

The trade-offs

Strengths and Limitations

Deepgram

Strengths

  • Unified speech-to-text, text-to-speech and voice-agent APIs
  • Real-time streaming with low latency for live applications
  • Usage-based with $200 free credit and no card required
  • Runs self-hosted for data control and compliance

Limitations

  • Developer platform, not a consumer voice studio
  • Several headline per-minute rates are promotional and sit below the list price, so the long-run cost depends on which rate applies

Speechmatics

Strengths

  • 55+ languages with explicit accent, dialect and code-switching coverage
  • Runs in the cloud, in containers, on-premises or on-device
  • $100 in credit to start, with no card required
  • Data is not logged as standard

Limitations

  • A developer API rather than a ready-made voice studio
  • Text-to-speech on the free tier covers English only
  • The Voice Agent API is still in early access
  • Volume discounts start at 500 hours a month, so small workloads pay the full rate

Plans

Pricing

Deepgram

  • Pay As You Go $0.0048 Per minute of Nova-3 monolingual streaming speech-to-text at the current promotional rate, $0.0077 at list. Nova-3 multilingual $0.0058, Flux English $0.0065. Text-to-speech $0.0150 / 1,000 characters on Aura-1 and $0.030 on Aura-2. Voice Agent API from $0.056 / minute. Add-ons are priced separately: redaction and diarization $0.0020 / minute each, entity detection $0.0017, keyterm prompting $0.0013. No minimums, no expiry, $200 in credits to start without a card.
  • Growth From $4,000 Per year in prepaid credits, redeemed against usage at lower rates: Nova-3 monolingual streaming $0.0042 / minute, Aura-1 $0.0135 and Aura-2 $0.027 / 1,000 characters.
  • Enterprise Not stated Custom pricing, volume commitments, self-hosting and dedicated support.
Visit site

Speechmatics

  • Free $0 $100 in credit, no card required. 55+ languages for speech-to-text, 2 concurrent real-time sessions and low-latency English text-to-speech.
  • Pro $0.129 Per hour of audio, billed to the second with no commitment. 50 concurrent real-time sessions, 10 file jobs per second and email support.
  • Enterprise Not stated Volume-based pricing that scales with usage, unlimited concurrency and no rate limits, custom models, custom vocabularies and formatting rules, and SaaS or on-premises deployment.
Visit site

What users say

Review Scores · opens after launch

Deepgram

Ease of use
Not rated yet
User interface
Not rated yet
Onboarding
Not rated yet
Support and service
Not rated yet
Value for money
Not rated yet
Features
Not rated yet

No reviews yet. Be the first to review this tool. Every review is checked for fairness before it appears, and the vendor cannot have one removed.

Speechmatics

Ease of use
Not rated yet
User interface
Not rated yet
Onboarding
Not rated yet
Support and service
Not rated yet
Value for money
Not rated yet
Features
Not rated yet

No reviews yet. Be the first to review this tool. Every review is checked for fairness before it appears, and the vendor cannot have one removed.

Where this comes from

Where this comes from

Deepgram https://deepgram.com · checked against the official source on 2026-09-01

Speechmatics https://www.speechmatics.com · checked against the official source on 2026-08-30

This table lists factual criteria taken from information the vendors publish themselves. For how we put it together, see the methodology.