Best for
Best for teams putting AI voice agents on real phone lines and testing them before they meet customers.
- Small and mid-sized
- Enterprise
No tools match
Press / to search, arrow keys to move, Enter to open
AI Voice & Speech Comparison
At a glance
Best for
Best for teams putting AI voice agents on real phone lines and testing them before they meet customers.
Best for
Best for teams transcribing real conversations across accents and languages, including where the audio may not leave their own infrastructure.
Cataleo does not name a winner. Both statements come from the vendors themselves.
| Criterion | Retell AI | Speechmatics |
|---|---|---|
| Starting price | lower value From $0.07 / minute of calls Pay as you go, quoted by the vendor as $0.07 to $0.31 per minute for voice agents depending on the models, telephony and add-ons chosen, plus from $0.002 per message for chat agents. Enterprise is quoted by sales. | From $0.129 / hour of audio Pro tier rate, billed to the second with no commitment. Volume discounts of 20% apply automatically above 500 hours per month for each speech-to-text type, with further discounts above 24,000 hours a year. One credit equals one dollar. |
| Free trial | $10 in free credits, no contract and no commitment; 20 concurrent calls are included | Free tier with $100 in credit, no card required |
| Commercial use | Yes, usage-based | Yes, usage-based |
| Output | Voice agents for inbound and outbound phone calls and web calls, plus chat agents over SMS and a web widget | Transcripts, translations and synthesised speech |
| API | Yes, REST API with Node.js and Python libraries, webhooks and an MCP server | Yes, batch and real-time speech-to-text, text-to-speech, translation and a management API |
| Languages | 57 languages listed for text-to-speech and 55 for speech recognition, each with the providers that cover it | 55+ languages and dialects for speech-to-text; translation across 69 language pairs |
| Hosting | Cloud on AWS; the vendor states it does not currently operate services within the European Union | Cloud, CPU or GPU containers, Kubernetes, on-premises virtual appliance, or on-device |
| Made in | Not stated | United Kingdom |
| Voice cloning | Yes, custom voices can be added by searching ElevenLabs community voices, importing an existing clone or training one from uploaded recordings, with Inworld, Cartesia, MiniMax, ElevenLabs and Retell's own platform clones as providers; up to 100 custom voices per account | No, the text-to-speech API ships four fixed voices, sarah and theo in British English and megan and jack in American English, and the documentation describes no custom or cloned voice; speed, pitch and emphasis are not adjustable either |
| Real-time use | Yes, live inbound and outbound phone calls and browser web calls, with a proprietary turn-taking model, batch calling for campaigns and live monitoring that lets a human listen in or step in | Yes, streaming transcription returned in under a second, alongside batch transcription of recorded files |
| Speech-to-speech and dubbing | A call runs as spoken conversation by chaining speech recognition, a language model and text-to-speech; dubbing of existing recordings is not part of the product | Audio translation across 69 language pairs, returned as translated transcript rather than dubbed audio; a Voice Agent API is in early access |
| Editing and controls | Single and multi-prompt agents or node-based conversation flows, knowledge bases, expressive mode for emotion and emphasis, custom pronunciation with IPA, CMU, Pinyin or Jyutping, denoising and interruption sensitivity, voicemail and IVR handling, plus a playground, simulation tests, A/B testing and post-call analysis | Speaker diarization and identification, custom dictionary, language identification, background speech filtering, and speech intelligence for summaries, sentiment, topics and auto-chapters |
| Consent and trust | HIPAA compliant, SOC 2 Type 1 and Type 2 certified and GDPR compliant, with a BAA and a DPA including EU standard contractual clauses available for self-signing at no extra fee. Per-agent data retention runs from 1 day to 2 years, PII can be excluded from what is stored, and recording URLs can be signed. Services are not currently operated inside the EU | Data is not logged as standard; container, Kubernetes, virtual appliance and on-device deployment keep audio inside your own infrastructure |
| MCP-ready | stated Yes | No |
| Deployment | Cloud (SaaS), Browser-based | more entries Cloud (SaaS), Browser-based, On-premise |
| Support | Not stated | Email helpdesk |
| Onboarding | Documentation and knowledge base | Documentation and knowledge base |
| Company size | Small and mid-sized, Enterprise | Small and mid-sized, Enterprise |
| Integrations | more entries Twilio, Telnyx, Vonage, Genesys Cloud, Five9, Avaya, Amazon Connect, HubSpot, Make, n8n, GoHighLevel | LiveKit, Pipecat, Vapi, Zapier |
Adds a column to the table, from the tools in this category.
Row where the values differ The small markers point at the lower number, the longer list or the stated feature. They describe the data, they are not a verdict.
The short version
Every line is one datapoint from the table above, picked automatically. Nothing here is written text.
The trade-offs
Retell AI
Strengths
Limitations
Speechmatics
Strengths
Limitations
Plans
Retell AI
Speechmatics
What users say
Retell AI
No reviews yet. Be the first to review this tool. Every review is checked for fairness before it appears, and the vendor cannot have one removed.
Speechmatics
No reviews yet. Be the first to review this tool. Every review is checked for fairness before it appears, and the vendor cannot have one removed.
Where this comes from
Retell AI https://www.retellai.com · checked against the official source on 2026-09-13
Speechmatics https://www.speechmatics.com · checked against the official source on 2026-08-30
This table lists factual criteria taken from information the vendors publish themselves. For how we put it together, see the methodology.