CataleoSoftware Get Your Software Listed Get Listed

AI

AI Voice & Speech

Text-to-speech, voice cloning and AI voice tools. Every tool below is measured against the same criteria, in the same order. Note that this field moves fast, so treat every value as a snapshot and check it against the vendor before you rely on it.

Filters

Also Covers
AI
Deployment
Company Size
Industry

32tools

  • Price Free tier; Plus from $15 / month

    Not rated yet

    Voice platform for expressive text-to-speech and fast voice cloning, with open-source models and a large community voice library.

    Strengths

    • Fast voice cloning and expressive emotion control
    • Open-source models (Fish Speech and Fish Audio S2)
    • Large community voice library and a developer API

    Limitations

    • API access and commercial use require a paid plan
    • Cloned and community voices raise consent questions
  • Read AI

    MCP-ready

    Price Free plan; Pro $19.75 / user / month

    Not rated yet

    Meeting assistant that reports on calls, email and messages, with live notes and coaching metrics.

    Strengths

    • Live notes and meeting metrics while the call is still running
    • Meeting coach and trend reporting on how meetings themselves are run
    • Search spans meetings, email and messages rather than calls alone
    • MCP server with Claude and ChatGPT connectors

    Limitations

    • Free plan is capped at five meeting transcripts a month
    • Audio and video playback only from the Enterprise plan up
    • Meeting length is capped by plan, from one hour on Free to eight on Enterprise+
    • SSO, SCIM, HIPAA support and data retention policies need Enterprise+ and five or more licences
  • Price Free tier; paid from $3 / month

    Not rated yet

    Expressive, emotionally intelligent text-to-speech (Octave) and an empathic voice interface (EVI), with voice cloning and voice design.

    Strengths

    • Expressive, emotionally aware speech
    • Voice cloning from about 15 seconds and voice design from text prompts
    • Low-latency streaming for real-time voice agents

    Limitations

    • Voice design and acting instructions are English only or still in preview
    • Character and EVI-minute caps scale the monthly cost
  • Price Text-to-speech API from $2 / hour

    Not rated yet

    Enterprise voice cloning and speech-to-speech used in film and games, with a real-time TTS API.

    Strengths

    • Broadcast-quality cloning trusted by major film and game studios
    • Covers speech-to-speech, accent correction and voice de-aging, not just text-to-speech
    • Marketplace and API are self-serve with pay-as-you-go pricing

    Limitations

    • Built for studios; the strongest features run through custom enterprise projects
    • Public pricing is limited to the marketplace and API
  • Price Free tier; Basic from $5 / month

    Not rated yet

    AI voice generator with 700+ voices and emotion control, plus voice cloning and an API.

    Strengths

    • 700+ voices with fine emotion and delivery control
    • Context-aware emotion model adapts tone to the script
    • Affordable entry with a commercial license from $5 per month

    Limitations

    • Free tier is limited to trial voices and requires attribution
    • Credit allowances are modest on the lower tiers
  • ElevenLabs

    MCP-ready

    Price Free tier; Starter from ~$5 / month

    Not rated yet

    The leading AI voice platform for realistic narration, voice cloning and dubbing.

    Strengths

    • Most realistic voice and cloning
    • 70+ languages on the v3 models
    • Speech-to-speech and dubbing
    • Strong API

    Limitations

    • Commercial use only from the paid Starter tier
    • API pricey at high volume
    • Credits are shared across every product and roll over for at most two months
  • Price Free; paid tiers above

    Not rated yet

    AI text-to-speech reader that turns documents and articles into listenable audio.

    Strengths

    • Best for turning reading material into audio
    • Accessibility focus
    • Popular with creators

    Limitations

    • Free tiers lack commercial rights
    • More consumer reader than production studio
  • Price Free tier (Flex, pay as you go); detection plans from $280 / month

    Not rated yet

    Voice cloning and synthetic-speech platform with speech-to-speech, voice design and deepfake detection.

    Strengths

    • Professional cloning with consent verification
    • Speech-to-speech and Voice Design
    • Trust and security tools

    Limitations

    • Public generative text-to-speech pricing is opaque, verify with a quote
    • More developer and security oriented
  • Price Free plan; Pro from $10 / seat / month

    Not rated yet

    An AI meeting assistant that records calls across Zoom, Teams and Meet, transcribes them and pushes the summaries into the CRM.

    Strengths

    • Covers Zoom, Microsoft Teams, Google Meet, Webex, GoToMeeting, Dialpad and Lifesize, plus dialers and uploaded audio files
    • More than 100 languages with automatic detection
    • Over 100 integrations carry summaries into CRM, ATS, chat and docs
    • SOC 2 Type II, GDPR and HIPAA compliance with zero data retention

    Limitations

    • Transcription only, no text-to-speech or voice cloning
    • API access, analytics and private channels all start at the Business plan
    • The vendor marks the wider language set as beta
  • Rev

    Price Free tier, then from $25.49 / seat / month

    Not rated yet

    AI and human transcription, captions and subtitles, plus a speech-to-text API, from a company that has been doing this since 2010.

    Strengths

    • Human transcription at a stated 99%+ accuracy alongside the machine tier
    • Developer API priced per hour of audio, from $0.10 on Reverb Turbo
    • Runs on-premises as well as in the cloud on the API side
    • Free credits worth 5 hours of transcription, with no card required

    Limitations

    • Speech-to-text only, no text-to-speech, dubbing or voice cloning
    • The subscription is built around legal and investigative work, which is dead weight for other uses
    • The free plan covers 45 minutes a month and English only
    • Human services are charged on top of the subscription rather than included
  • Retell AI

    MCP-ready

    Price From $0.07 / minute of calls

    Not rated yet

    Build, test and monitor AI voice and chat agents that handle phone calls, billed by the minute.

    Strengths

    • Usage-based per-minute pricing with no contract and $10 in free credits to start
    • Testing tooling built into the product: playground, simulation runs, batch tests and A/B splits of live traffic
    • Custom voices can be cloned or imported across several voice providers
    • HIPAA, SOC 2 Type 1 and Type 2 and GDPR, with BAA and DPA available for self-signing at no extra fee

    Limitations

    • The per-minute rate varies widely with the models, telephony and add-ons chosen, so the entry rate is not the rate to budget with
    • The vendor states it does not currently operate services within the European Union
    • Retell-managed phone numbers cover the US and Canada only; elsewhere a number has to be brought in
    • Building and running an agent assumes developer work; there is no packaged seat-based plan
  • Bland AI

    MCP-ready

    Price From $0.14 / minute of calls

    Not rated yet

    Enterprise AI phone agents on a stack the vendor owns end to end, priced as one per-minute rate.

    Strengths

    • One per-minute rate covers the language model, speech-to-text, text-to-speech and telephony, with no token charges or provider pass-throughs
    • Self-hosted and on-premises deployments, plus US, EU and APAC data residency
    • SOC 2 Type II, HIPAA, GDPR and PCI DSS v4.0, with FedRAMP 20X Class A available
    • The same pathways, personas and memory run across voice, SMS, RCS, iMessage and web chat

    Limitations

    • Owning the stack means no choice of speech or language model provider
    • The lower $0.12 per minute rate carries a $299 / month platform fee, so it only pays off above a certain volume
    • Transfer time to a human is billed on top of the per-minute rate
    • Positioned around regulated enterprise workloads, which is more platform than a small team needs for a single use case
  • Price From $9 / month

    Not rated yet

    AI video dubbing, subtitles and text-to-speech in 32 languages, with a low-latency speech API.

    Strengths

    • Commercial rights on both self-serve plans, with unlimited downloads and no project expiry
    • Deep coverage of Indian languages alongside European and Asian ones
    • Low-latency text-to-speech API with automatic language detection and streaming
    • Entry pricing is low compared with other dubbing platforms

    Limitations

    • Lip sync, multi-speaker support and human review are Enterprise only
    • Voice cloning starts on the Supreme plan
    • Both paid plans include only 50 credits per month, so extra credits are bought on top
    • 32 languages is a narrower set than the largest localization platforms offer
  • Price Free trial; paid from $33 / month

    Not rated yet

    Video and audio localization: transcription, translation, voice cloning and lip sync in 130+ languages.

    Strengths

    • Transcription, translation, dubbing and lip sync in one pass
    • Voice cloning carries the original speaker into 32 languages
    • Multi-speaker detection keeps panels and interviews on separate voices
    • SOC 2 Type II certified, with team spaces and shared voice presets

    Limitations

    • Lip sync only from the Creator Pro plan upwards
    • The entry plan carries only a limited API; a production API with webhooks starts on Creator Pro
    • Every plan caps processed minutes, so heavy localization moves up the tiers quickly
    • Voice cloning covers 32 of the 130+ translation languages
  • Otter.ai

    MCP-ready

    Price Free tier, then from $8.33 / user / month

    Not rated yet

    An AI meeting assistant that joins calls, transcribes them live and turns the recording into summaries, action items and searchable notes.

    Strengths

    • Records without a bot in the call from the desktop app
    • Summaries and action items sync into CRM, chat and project tools
    • MCP server exposes meeting data to Claude and ChatGPT
    • Free plan with 300 transcription minutes a month

    Limitations

    • Transcription only, no text-to-speech or voice cloning
    • A small number of languages, and only one per conversation
    • Per-meeting length caps on the lower plans: 30 minutes on Basic and 90 on Pro
    • API access is reserved for Enterprise
  • Synthflow

    MCP-ready

    Price From $30,000 / year

    Not rated yet

    A no-code builder for AI phone agents, sold as an enterprise contract with implementation included.

    Strengths

    • No-code Flow Designer with a multi-agent system, so a non-developer can build and change call flows
    • Connects to existing telephony and contact centre stacks over SIP, including Genesys, Avaya, Five9, Cisco and RingCentral
    • White-labelled deployments for partners reselling voice agents under their own brand
    • SOC 2, HIPAA, PCI DSS, GDPR and ISO 27001, with agent-level retention and PII redaction controls

    Limitations

    • No self-serve or trial tier is published; the only listed plan is an enterprise contract from $30,000 per year
    • Pricing is scoped by sales, so the cost of a given call volume is not knowable from the website
    • Latency figures differ between the vendor's own pages, so the published number is not a specification to rely on
    • A cloned voice has to be created in a separate ElevenLabs account and imported, so it means a second subscription alongside Synthflow
  • Not rated yet

    A real-time voice changer and soundboard that sits between your microphone and Discord, OBS or the game you are in.

    Strengths

    • Works live in more than 30 games and applications through a virtual microphone
    • Voicelab builds custom voices from over 140 effects and publishes them to the community
    • Control API drives the app from Unity, JavaScript, C/C++ or Java
    • AI voices are trained with professional voice actors under a Fairly Trained certification
    • Free version for Windows and Mac

    Limitations

    • No voice cloning: Voicelab shapes effects rather than reproducing a specific person's voice
    • Plan amounts are not published on the site
    • Built around gaming, streaming and chat applications rather than rendered audio files
  • Price Free tier, then from $4 per 1 million characters

    Not rated yet

    Google Cloud's speech synthesis API, from cheap legacy voices to Chirp 3 HD and prompt-controlled Gemini-TTS.

    Strengths

    • One API spans cheap legacy voices at $4 per million characters and current HD voices at $30
    • Gemini-TTS takes a written prompt for the delivery instead of hand-written SSML markup
    • Free monthly character allowance on every voice type except Gemini-TTS and custom voices
    • Bidirectional streaming, single request and long-form synthesis from the same service
    • Voice cloning requires a recorded consent statement from the speaker, in their own language

    Limitations

    • Text to speech only: recognition and translation are separate Google Cloud services
    • Instant Custom Voice is restricted to allow-listed customers and goes through sales
    • Studio voices cost $160 per million characters, forty times the Standard rate
    • Gemini-TTS is billed per token rather than per character, so the two price models do not compare directly
    • The overlapping generations of voice model take some reading before the right one is obvious
  • Price Free tier, then pay-as-you-go

    Not rated yet

    Microsoft's speech platform: transcription, neural text to speech, custom voices, translation and talking avatars, in the cloud or in containers.

    Strengths

    • One service covers transcription, synthesis, translation, speaker recognition and avatars
    • Runs in containers at the edge, embedded on-device and in sovereign clouds
    • Custom neural voice and custom speech models for brand voices and domain vocabulary
    • Free F0 tier with 5 audio hours and 0.5 million characters a month

    Limitations

    • Pay-as-you-go rates are not published as fixed figures and depend on region and currency
    • Custom neural voice requires a limited-access application before it can be used
    • A platform assembled from many services rather than a single endpoint, so setup takes longer
    • The LLM speech model is still in preview
  • AssemblyAI

    MCP-ready

    Price From $0.15 / hour of audio

    Not rated yet

    Speech-to-text, speech understanding and voice agent APIs for developers, billed per hour of audio.

    Strengths

    • Pay-as-you-go with no subscription and $50 in free credits to start
    • 99 languages for pre-recorded transcription on Universal-2
    • Compliance certifications and an EU region at no price premium
    • Official MCP server exposes transcription and transcript search to AI agents

    Limitations

    • A developer API, not a ready-made studio or editor
    • No voice cloning or voice creation
    • Speech Understanding add-ons are billed on top of the base rate, so cost stacks per request
    • Streaming is billed for as long as the connection stays open, not for the audio sent, and multichannel files are billed per channel
    • Universal-3.5 Pro covers 18 languages against 99 on the older Universal-2
  • Vozo

    Price Free tier; paid from $29 / month

    Not rated yet

    AI video localization with dubbing, lip sync and on-screen text translation across 160+ languages.

    Strengths

    • Very broad language coverage, 111 source and 165 target languages
    • Visual Translate rewrites on-screen text and keeps layout and animation
    • Free tier covers a first project without a card
    • Two dubbing modes, one keeping the original voice and one using a native accent

    Limitations

    • API access only on Enterprise plans
    • Metered in AI points, so cost depends on the mix of dubbing, lip sync and visual translation
    • Each tier caps video length, seats and concurrent jobs
    • Lip sync and watermark-free export need at least the Creator plan
  • Price Free tier, then from $4 per 1 million characters

    Not rated yet

    AWS text-to-speech API with Standard, Neural, Long-Form and Generative voices, billed per character.

    Strengths

    • Four voice engines let quality be traded against cost per character, from $4 to $100 per million
    • Generous free tier, with 5 million Standard characters a month that does not expire
    • Generated speech may be cached and replayed at no additional cost
    • Speech Marks and custom lexicons cover lip-sync and awkward pronunciations properly
    • Brand Voice cannot be used without the voice owner being cast and recorded

    Limitations

    • Text to speech only: no transcription, translation or dubbing in this service
    • No self-service voice cloning, and Brand Voice starts with an AWS account manager
    • The Generative and Long-Form engines cover far fewer voices than Standard and Neural
    • Long-Form voices cost $100 per million characters, twenty-five times the Standard rate
    • Useful only from code, so it suits developers rather than anyone wanting an editor
  • Price From $19 / month

    Not rated yet

    Enterprise AI voice platform with studio-quality avatars and ethically sourced voices.

    Strengths

    • Voices licensed from professional voice actors, with consent
    • AI Director controls for pitch, pace and tone, plus custom pronunciation rules
    • Shared workspaces, so a script gets reviewed before it ships
    • Real-time streaming API at sub-600ms latency, up to 96 kHz

    Limitations

    • Free tier has no commercial rights and only three download minutes a month
    • Paid tiers meter output in minutes rather than leaving it open
    • All languages and translation are Enterprise-only, the self-serve tiers are English
    • Adobe Premiere Pro integration starts at the $160 Business tier
  • Price From $0.129 / hour of audio

    Not rated yet

    Speech-to-text APIs built for accents, dialects and code-switching, with cloud, on-premises and on-device deployment.

    Strengths

    • 55+ languages with explicit accent, dialect and code-switching coverage
    • Runs in the cloud, in containers, on-premises or on-device
    • $100 in credit to start, with no card required
    • Data is not logged as standard

    Limitations

    • A developer API rather than a ready-made voice studio
    • Text-to-speech on the free tier covers English only
    • The Voice Agent API is still in early access
    • Volume discounts start at 500 hours a month, so small workloads pay the full rate
  • Vapi

    MCP-ready

    Price From $0.05 / minute of calls

    Not rated yet

    A developer platform for building voice AI agents that make and take phone calls, billed per minute.

    Strengths

    • Per-minute pricing with model provider costs passed through at cost
    • Bring your own provider keys and the pass-through drops to nothing
    • Free choice of speech-to-text, model and voice provider across dozens of options
    • Official MCP server lets external AI agents drive Vapi's APIs

    Limitations

    • HIPAA compliance and zero data retention are monthly add-ons rather than included
    • Total cost depends on which model, transcription and voice providers you choose
    • Concurrent lines beyond the first ten cost $10 per month each
    • A platform to build on, not a finished call-centre product
  • Fathom

    MCP-ready

    Price Free plan; Premium $20 / user / month

    Not rated yet

    AI notetaker that records, transcribes and summarises calls, with an unlimited free plan.

    Strengths

    • Free plan carries unlimited recordings, transcription and AI summaries
    • Records without a bot in the call from the desktop app, or with one
    • Transcripts in 38 languages
    • Documented REST API with webhooks, SDKs and an MCP server

    Limitations

    • Transcription and summarisation only, with no voice synthesis, cloning or dubbing
    • Summary translation covers six languages against the 38 it transcribes
    • Team and Business plans carry a two-user minimum
    • CRM field sync, coaching metrics and custom summaries start on the Business plan
  • Granola

    MCP-ready

    Price Free tier, then $14 / user / month

    Not rated yet

    AI meeting notepad that records from your computer audio instead of sending a bot into the call.

    Strengths

    • No bot joins the call, so nothing appears in the attendee list
    • Expands shorthand typed during the meeting instead of replacing note-taking
    • AI chat runs across the whole meeting history, not a single transcript
    • Opt-out of model training on every plan, including the free one
    • MCP server and API on Business, so meeting context reaches other AI tools

    Limitations

    • Capture depends on a machine being in the meeting, since it records local computer audio
    • The free plan limits how far meeting history reaches back
    • Integrations, API and MCP start on the Business plan
  • Price Commercial from $16.50 / user / month

    Not rated yet

    AI text-to-speech reader and commercial voice generator, with voices from Gemini, OpenAI, Azure and ElevenLabs.

    Strengths

    • Bundles voices from Gemini, OpenAI, Azure and ElevenLabs in one place
    • Doubles as a reader for your own documents, PDFs and ebooks
    • Commercial license and up to four voice clones on paid plans

    Limitations

    • Personal reader and commercial generator are separate apps with separate subscriptions
    • The Team plan requires a minimum of two users
    • No API, so the voices cannot be built into another product or used live
  • Price Free tier; paid from $5 / month

    Not rated yet

    Localization AI for dubbing and text-to-speech, with real-time dubbing (DubStream) and the MARS speech models.

    Strengths

    • Real-time dubbing (DubStream) and on-demand dubbing (DubStudio)
    • MARS speech models tuned for conversation, dubbing and narration
    • Broad language coverage for media and sports localization

    Limitations

    • Priced in credits, so cost depends on usage
    • Aimed at localization and dubbing more than general voiceover
  • Price Free tier; Pro from $5 / month

    Not rated yet

    Real-time AI voice (Sonic) built for low-latency voice agents, with instant cloning.

    Strengths

    • Lowest-latency real-time voice for agents
    • Instant three-second cloning
    • Cost-effective developer API

    Limitations

    • Less expressive on long-form and audiobook content
    • Newer and developer-oriented
  • Price Free tier; paid tiers above

    Not rated yet

    AI voiceover platform for business presentations and e-learning, with real-time voice agents.

    Strengths

    • Polished business voiceover
    • Flexible for training and e-learning
    • Also does real-time agent voices

    Limitations

    • Free tier has no downloads or commercial rights
    • Less storytelling realism than ElevenLabs
  • Deepgram

    MCP-ready

    Price Usage-based, pay-as-you-go

    Not rated yet

    A developer voice AI platform with speech-to-text, text-to-speech and voice agents through unified, real-time APIs.

    Strengths

    • Unified speech-to-text, text-to-speech and voice-agent APIs
    • Real-time streaming with low latency for live applications
    • Usage-based with $200 free credit and no card required
    • Runs self-hosted for data control and compliance

    Limitations

    • Developer platform, not a consumer voice studio
    • Several headline per-minute rates are promotional and sit below the list price, so the long-run cost depends on which rate applies

Direct Comparisons