
Hume AI built EVI, an empathic voice interface that reads emotion from a caller's voice and responds with matching tone and pacing. Founded by ex-Google emotion researcher Alan Cowen, it's the only commercial voice AI with bidirectional emotional intelligence built into its core architecture.
Hume AI is an AI research company founded in March 2021 by Alan Cowen, a cognitive scientist who spent years leading affective computing research at Google AI before completing his PhD in Psychology at UC Berkeley. The company is named after philosopher David Hume, whose argument that emotions drive human choice runs through every product decision. In March 2024, Hume raised a $50M Series B led by EQT Ventures, and simultaneously released its flagship product, the Empathic Voice Interface (EVI), publicly via API. The core problem Hume is solving is one that every other voice AI sidesteps: a caller who sounds frustrated, scared, or distressed should not receive the same tone as a caller who sounds cheerful, and the AI itself should adjust, not just generate expressive speech, but read the emotional subtext in what it hears.
EVI is a speech-to-speech model that processes voice directly, not a separate STT pipeline feeding an LLM feeding a TTS engine, which is how it achieves sub-300ms response latency while simultaneously analyzing prosodic cues in the caller's voice. The current version, EVI 3 (launched May 29, 2025), introduces custom voice cloning from under 30 seconds of audio, support for 200,000+ pre-designed voices, integration with Claude 4, Gemini 2.5, and Kimi K2, and expressive output across 30 distinct emotional and vocal styles ranging from calm reassurance to high energy enthusiasm. Alongside EVI, Hume offers the Octave text-to-speech engine (released December 2024) and an Expression Measurement API for analyzing emotional content across audio, video, and images. Plans start at $3/month for the Starter tier, with a free tier available for evaluation.
What Hume EVI produces in May 2026
EVI 3 is best understood as a unified speech-language model rather than a stitched-together pipeline. Its pretraining on trillions of text tokens and millions of hours of speech means it understands how vocal characteristics interact with language, not just what words were said. When a caller's voice rises in pitch or their sentences shorten and clip, EVI detects those signals and adjusts its own response: slowing down, softening tone, adding pauses that signal it is genuinely processing what was said before answering.
In terms of raw capabilities: EVI 3 delivers conversational latency around 300ms, with instant mode enabled by default that starts sending audio chunks within approximately 200ms. It supports over 11 languages, though the product is clearly English-first in optimization. Voice cloning from under 30 seconds of audio captures timbre, accent, rhythm, tone, and inferred personality, making it possible to build a custom voice for an enterprise product without weeks of studio recording. The Expression Measurement API, offered separately at pay-as-you-go rates, can analyze emotional content from video with audio ($0.0828/min), audio only ($0.0639/min), and images ($0.00204/image), which opens up analytics use cases beyond pure voice generation.
The EVI 3 launch in May 2025 was a credible milestone. In blind testing Hume published alongside the release, EVI 3 outperformed GPT-4o, Gemini, and Sesame in identifying eight of nine tested emotions from voice tone alone. Alan Cowen stated at the launch: "At Hume, we promised ourselves that before the end of 2025, we'd achieve a voice AI experience that can be fully personalized." EVI 3 delivered on that specific promise, custom voice creation without fine-tuning had not been possible in prior EVI versions.
Real-world deployments corroborate the technical claims. Vonova, a customer support platform, used EVI-powered voice agents to achieve a 40% reduction in operational costs versus traditional call centers, and a 20% improvement in AI resolution rate without human escalation. Thumos Care, a preventive healthcare platform, integrated EVI for patient conversations where emotional sensitivity is clinically meaningful. Their CEO Shan said: "Hume's EVI was transformative. It accelerated our development of emotionally intelligent features and helped expand our product offering." Ream, a mental health coaching app, reported doubling its daily active users after switching to EVI-powered voice coaching from text-based interactions.
Where Hume EVI sits versus ElevenLabs and Cartesia AI
ElevenLabs is the most direct comparison in market positioning, but the products are mechanically different. ElevenLabs' core technology is a text-to-speech model that does not natively read emotional state from the user's voice. Its strength is output quality and breadth: 70+ languages (versus Hume's 11), a voice library of comparable size, and speech naturalness scores of 89.6% in independent benchmarks versus Hume's 78.5%. ElevenLabs' v3 model delivers clean, consistent narration that "can pass for human speech" across content creation, audiobook production, and multilingual deployment. Emotion is available in ElevenLabs through SSML tags and prompting, but it is an output-side parameter, the model does not listen to how you sounded and adapt. For a customer support agent that needs to detect a frustrated caller and respond with genuine de-escalation, ElevenLabs provides no native mechanism to do that. For a podcaster who needs multilingual narration in 40 languages, ElevenLabs is the clear choice and Hume is not the right tool.
Cartesia AI operates at a different layer entirely. Cartesia's Sonic model is a pure TTS streaming engine targeting sub-90ms time-to-first-audio, compared to Hume's 200-300ms. Cartesia is approximately 33% cheaper per character ($0.00004 vs. $0.00006) and targets developers who assemble their own voice pipelines: they bring STT, route to their own LLM, and use Cartesia for fast, clean synthesis on the output side. Cartesia has no emotion detection in either direction. The tradeoff is explicit: if latency and cost efficiency are the constraints and you are willing to manage the emotional context layer yourself, Cartesia wins. If the emotional responsiveness needs to be baked into the voice model itself, which matters most in mental health, healthcare, and high-stakes customer interactions. Cartesia cannot substitute for EVI.
For developers building systems where callers might be distressed, angry, or scared, the bidirectionality is the whole point. EVI reads what it hears; ElevenLabs and Cartesia do not. That distinction narrows Hume's addressable market but makes it genuinely differentiated within that narrower space, much like how Whisper dominates speech-to-text accuracy but is not a conversational model. Pairing tools like OpenVoice for custom voice cloning or PlayHT for voice generation shows how crowded the adjacent space is, but none of those tools attempt bidirectional emotional intelligence at the model level.
"The accuracy of emotion detection blew me away. Hume AI could tell the difference between annoyed and angry callers.", reviewer, fahimai.com, 2026
The licensing and copyright reality
EVI 3's voice cloning capability raises the same questions that follow any voice AI platform. Hume requires users to confirm they have rights to any voice they clone, and terms of service prohibit cloning voices without consent. In practice, this is a self-policed guardrail. Hume does not have a separate voice verification layer, so the responsibility falls to the developer or enterprise customer. The Expression Measurement API, which processes third-party audio and video for emotional content, sits in a more complex legal landscape: using it to analyze customer calls without disclosure could create liability depending on jurisdiction, and Hume's documentation does not make that risk explicit in the onboarding flow.
Commercial rights for EVI-generated audio are included in all paid plans; the Creator plan ($7-$14/mo) is the entry point for a commercial license. Free tier output is for evaluation only and cannot be used in production products. For enterprises in regulated industries, healthcare, financial services. Hume positions EVI as HIPAA-compliant, which is a meaningful differentiation over less compliance-conscious voice AI tools, though enterprises should verify the specifics of their business associate agreement setup rather than assuming coverage.
Where Hume EVI reliably falls short
The language coverage gap is the clearest hard limit. EVI supports approximately 11 languages as of early 2026. ElevenLabs supports 70+. If your product serves non-English speaking users in any meaningful volume, EVI is not a viable primary voice layer today. This is not a minor gap to work around, it's a deployment blocker for any global or multilingual customer support or healthcare use case.
Access to EVI requires engineering work. There is no GUI where a non-developer can build, test, and deploy a voice agent. Every meaningful configuration, system prompts, voice selection, LLM routing, context injection, happens through the API or SDKs. Smaller teams without engineering resources cannot self-serve. The free tier's 5 EVI minutes gives almost nothing to evaluate with; meaningful testing requires the Starter tier at minimum.
Billing can surprise teams at scale. The overage model charges $0.04-$0.07 per EVI minute past the plan allocation. A customer support team with high call volume in a busy month can find that a $70 Pro plan (1,200 EVI minutes included) accumulates hundreds of dollars in overages. Unlike Cartesia's flat per-character pricing or ElevenLabs' clearer credit model, Hume's cost structure requires careful monitoring for production workloads.
Concurrency limits are plan-gated in ways that become painful. Scaling concurrent EVI sessions requires moving to Business ($500/mo) or Enterprise tiers. Teams that prototype on Starter and grow quickly hit these ceilings before they have the revenue to justify the plan jump.
"The first week was tough. Setting up API calls took longer than I expected.", developer user, referenced in fahimai.com 2026 review
Who Hume EVI is for
EVI is the right choice for engineering teams building conversational voice applications where the emotional register of both the caller and the AI response are load-bearing requirements. Customer support for high-stakes or sensitive interactions, mental health coaching, healthcare patient intake, and tutoring applications where a distressed or confused user needs more than correct information, they need a voice that responds to how they actually sound. In those contexts, Hume's bidirectional emotional intelligence is not a nice-to-have; it is the core feature that distinguishes it from bolting sentiment keywords onto a standard TTS prompt.
Skip Hume EVI if your primary requirement is multilingual coverage, content narration, audiobook production, or pure latency optimization for a voice pipeline. Those use cases are better served by ElevenLabs or Cartesia AI respectively, at lower cost and with less integration complexity. Also skip it if your team cannot dedicate engineering time to the integration, this is not a product for non-technical users who want to build a voice agent through a visual interface.
The $50M Series B from EQT Ventures in March 2024, combined with Northwell Holdings (a healthcare system) and LG Technology Ventures among the investors, signals where the company sees its future: regulated industries and enterprise deployments where the emotional intelligence layer is defensible IP, not just a prompt engineering trick. For developers building in those spaces, Hume's research pedigree, RLHE training methodology, and the measurable customer outcomes from Vonova, Thumos Care, and Ream make it the most credible specialist voice AI currently available.
User Reviews
No reviews yet. Be the first to share your experience!
Sign in to write a review.
Related articles
Guides and articles related to Hume AI.

OpenAI GPT-Realtime-2 (May 2026): Pricing, Latency & 30-Min Voice Agent
Build a 70-Language Live Voice Agent With GPT-Realtime-Translate (2026)

Grok 4.3 API for Agents (May 2026): Pricing, Benchmarks, Migration

Voice Agent for Missed Calls: Every Service Business Is Bleeding Leads After Hours (2026)

AI Receptionist for Solo Professionals: Replace the Answering Service (2026)
