Best for
Best for teams putting AI voice agents on real phone lines and testing them before they meet customers.
- Small and mid-sized
- Enterprise
No tools match
Press / to search, arrow keys to move, Enter to open
AI Voice & Speech Comparison
At a glance
Best for
Best for teams putting AI voice agents on real phone lines and testing them before they meet customers.
Best for
Best for developers putting voice agents on real phone lines for support, lead qualification or scheduling.
Cataleo does not name a winner. Both statements come from the vendors themselves.
| Criterion | Retell AI | Vapi |
|---|---|---|
| Starting price | From $0.07 / minute of calls Pay as you go, quoted by the vendor as $0.07 to $0.31 per minute for voice agents depending on the models, telephony and add-ons chosen, plus from $0.002 per message for chat agents. Enterprise is quoted by sales. | lower value From $0.05 / minute of calls Build plan, usage-based, plus $0.005 per SMS or chat message. Ten concurrent lines are included and further lines cost $10 / month each. Model provider costs for speech-to-text, the LLM and text-to-speech are passed through at cost, or nothing if you bring your own API keys. |
| Free trial | $10 in free credits, no contract and no commitment; 20 concurrent calls are included | Not stated |
| Commercial use | Yes, usage-based | Yes, usage-based |
| Output | Voice agents for inbound and outbound phone calls and web calls, plus chat agents over SMS and a web widget | Voice agents for inbound and outbound phone calls, web voice, and SMS or chat |
| API | Yes, REST API with Node.js and Python libraries, webhooks and an MCP server | Yes, REST API, CLI and SDKs |
| Hosting | Cloud on AWS; the vendor states it does not currently operate services within the European Union | Not stated |
| Voice cloning | Yes, custom voices can be added by searching ElevenLabs community voices, importing an existing clone or training one from uploaded recordings, with Inworld, Cartesia, MiniMax, ElevenLabs and Retell's own platform clones as providers; up to 100 custom voices per account | Not part of Vapi itself; voices come from connected providers such as ElevenLabs and Deepgram |
| Real-time use | Yes, live inbound and outbound phone calls and browser web calls, with a proprietary turn-taking model, batch calling for campaigns and live monitoring that lets a human listen in or step in | Yes, live inbound and outbound phone and web conversations at an average latency under 500 ms |
| Languages | 57 languages are listed for text-to-speech and 55 for speech recognition, and a language needs a provider on both sides to be usable. An agent runs either in a single language or as a multilingual agent across a chosen set | Depends on the providers chosen for a given assistant: automatic language detection and code-switching with Deepgram (100+ languages), Google STT (125+) or Gladia (110+), while Azure, OpenAI Whisper, Speechmatics and Talkscriber run one language at a time; on the voice side Vapi's own voices cover 40+ languages, Azure 140+ and ElevenLabs 30+. The languages an assistant may use have to be listed in its system prompt |
| Speech-to-speech and dubbing | A call runs as spoken conversation by chaining speech recognition, a language model and text-to-speech; dubbing of existing recordings is not part of the product | Yes, a call runs as spoken conversation by chaining speech-to-text, a language model and text-to-speech, with natural turn-taking |
| Editing and controls | Single and multi-prompt agents or node-based conversation flows, knowledge bases, expressive mode for emotion and emphasis, custom pronunciation with IPA, CMU, Pinyin or Jyutping, denoising and interruption sensitivity, voicemail and IVR handling, plus a playground, simulation tests, A/B testing and post-call analysis | Assistants for a single prompt with tools, Squads for multi-assistant orchestration, tool calls into APIs and databases mid-conversation, plus testing and call observability |
| Consent and trust | HIPAA compliant, SOC 2 Type 1 and Type 2 certified and GDPR compliant, with a BAA and a DPA including EU standard contractual clauses available for self-signing at no extra fee. Per-agent data retention runs from 1 day to 2 years, PII can be excluded from what is stored, and recording URLs can be signed. Services are not currently operated inside the EU | SOC 2, HIPAA and PCI compliance with SSO, OAuth and role-based access control; HIPAA is a $2,000 / month add-on and zero data retention $1,000 / month on both plans |
| MCP-ready | Yes | Yes |
| Deployment | Cloud (SaaS), Browser-based | Cloud (SaaS), Browser-based |
| Support | Not stated | Email helpdesk, Community forum |
| Onboarding | Documentation and knowledge base | Documentation and knowledge base |
| Company size | Small and mid-sized, Enterprise | Small and mid-sized, Enterprise |
| Integrations | Twilio, Telnyx, Vonage, Genesys Cloud, Five9, Avaya, Amazon Connect, HubSpot, Make, n8n, GoHighLevel | OpenAI, Anthropic, Groq, Deepgram, Gladia, ElevenLabs, Salesforce, HubSpot, Zapier, Slack, Google Workspace |
Adds a column to the table, from the tools in this category.
Row where the values differ The small markers point at the lower number, the longer list or the stated feature. They describe the data, they are not a verdict.
The short version
Every line is one datapoint from the table above, picked automatically. Nothing here is written text.
The trade-offs
Retell AI
Strengths
Limitations
Vapi
Strengths
Limitations
Plans
Retell AI
Vapi
What users say
Retell AI
No reviews yet. Be the first to review this tool. Every review is checked for fairness before it appears, and the vendor cannot have one removed.
Vapi
No reviews yet. Be the first to review this tool. Every review is checked for fairness before it appears, and the vendor cannot have one removed.
Where this comes from
Retell AI https://www.retellai.com · checked against the official source on 2026-09-13
Vapi https://vapi.ai · checked against the official source on 2026-08-30
This table lists factual criteria taken from information the vendors publish themselves. For how we put it together, see the methodology.