Best for
Best for teams transcribing real conversations across accents and languages, including where the audio may not leave their own infrastructure.
- Small and mid-sized
- Enterprise
No tools match
Press / to search, arrow keys to move, Enter to open
AI Voice & Speech Comparison
At a glance
Best for
Best for teams transcribing real conversations across accents and languages, including where the audio may not leave their own infrastructure.
Best for
Best for developers putting voice agents on real phone lines for support, lead qualification or scheduling.
Cataleo does not name a winner. Both statements come from the vendors themselves.
| Criterion | Speechmatics | Vapi |
|---|---|---|
| Starting price | From $0.129 / hour of audio Pro tier rate, billed to the second with no commitment. Volume discounts of 20% apply automatically above 500 hours per month for each speech-to-text type, with further discounts above 24,000 hours a year. One credit equals one dollar. | lower value From $0.05 / minute of calls Build plan, usage-based, plus $0.005 per SMS or chat message. Ten concurrent lines are included and further lines cost $10 / month each. Model provider costs for speech-to-text, the LLM and text-to-speech are passed through at cost, or nothing if you bring your own API keys. |
| Free trial | Free tier with $100 in credit, no card required | Not stated |
| Commercial use | Yes, usage-based | Yes, usage-based |
| Output | Transcripts, translations and synthesised speech | Voice agents for inbound and outbound phone calls, web voice, and SMS or chat |
| API | Yes, batch and real-time speech-to-text, text-to-speech, translation and a management API | Yes, REST API, CLI and SDKs |
| Hosting | Cloud, CPU or GPU containers, Kubernetes, on-premises virtual appliance, or on-device | Not stated |
| Made in | United Kingdom | Not stated |
| Voice cloning | No, the text-to-speech API ships four fixed voices, sarah and theo in British English and megan and jack in American English, and the documentation describes no custom or cloned voice; speed, pitch and emphasis are not adjustable either | Not part of Vapi itself; voices come from connected providers such as ElevenLabs and Deepgram |
| Real-time use | Yes, streaming transcription returned in under a second, alongside batch transcription of recorded files | Yes, live inbound and outbound phone and web conversations at an average latency under 500 ms |
| Languages | 55+ languages and dialects with accent, dialect and code-switching coverage; translation across 69 language pairs | Depends on the providers chosen for a given assistant: automatic language detection and code-switching with Deepgram (100+ languages), Google STT (125+) or Gladia (110+), while Azure, OpenAI Whisper, Speechmatics and Talkscriber run one language at a time; on the voice side Vapi's own voices cover 40+ languages, Azure 140+ and ElevenLabs 30+. The languages an assistant may use have to be listed in its system prompt |
| Speech-to-speech and dubbing | Audio translation across 69 language pairs, returned as translated transcript rather than dubbed audio; a Voice Agent API is in early access | Yes, a call runs as spoken conversation by chaining speech-to-text, a language model and text-to-speech, with natural turn-taking |
| Editing and controls | Speaker diarization and identification, custom dictionary, language identification, background speech filtering, and speech intelligence for summaries, sentiment, topics and auto-chapters | Assistants for a single prompt with tools, Squads for multi-assistant orchestration, tool calls into APIs and databases mid-conversation, plus testing and call observability |
| Consent and trust | Data is not logged as standard; container, Kubernetes, virtual appliance and on-device deployment keep audio inside your own infrastructure | SOC 2, HIPAA and PCI compliance with SSO, OAuth and role-based access control; HIPAA is a $2,000 / month add-on and zero data retention $1,000 / month on both plans |
| MCP-ready | No | stated Yes |
| Deployment | more entries Cloud (SaaS), Browser-based, On-premise | Cloud (SaaS), Browser-based |
| Support | Email helpdesk | more entries Email helpdesk, Community forum |
| Onboarding | Documentation and knowledge base | Documentation and knowledge base |
| Company size | Small and mid-sized, Enterprise | Small and mid-sized, Enterprise |
| Integrations | LiveKit, Pipecat, Vapi, Zapier | more entries OpenAI, Anthropic, Groq, Deepgram, Gladia, ElevenLabs, Salesforce, HubSpot, Zapier, Slack, Google Workspace |
Adds a column to the table, from the tools in this category.
Row where the values differ The small markers point at the lower number, the longer list or the stated feature. They describe the data, they are not a verdict.
The short version
Every line is one datapoint from the table above, picked automatically. Nothing here is written text.
The trade-offs
Speechmatics
Strengths
Limitations
Vapi
Strengths
Limitations
Plans
Speechmatics
Vapi
What users say
Speechmatics
No reviews yet. Be the first to review this tool. Every review is checked for fairness before it appears, and the vendor cannot have one removed.
Vapi
No reviews yet. Be the first to review this tool. Every review is checked for fairness before it appears, and the vendor cannot have one removed.
Where this comes from
Speechmatics https://www.speechmatics.com · checked against the official source on 2026-08-30
Vapi https://vapi.ai · checked against the official source on 2026-08-30
This table lists factual criteria taken from information the vendors publish themselves. For how we put it together, see the methodology.