AI Voice Agents for Business: How They Work, Latency, and What They Cost
Short answer
An AI voice agent is software that holds spoken conversations over the phone or web, typically by chaining speech-to-text, a language model and text-to-speech, or by using a speech-to-speech model. Good voice agents respond in under about a second, handle interruptions, connect to business systems to take actions, and hand off to humans when needed. All-in running costs commonly fall somewhere between a few cents and a few tens of cents per minute, depending on the stack.
Key takeaways
- Latency is the make-or-break metric — aim for well under a second per response.
- Voice agents are most valuable when connected to systems that let them complete tasks.
- Interruptions, accents, noisy lines and human hand-off must be designed for, not patched.
- Disclosure, consent and call recording rules apply — design for compliance from the start.
Voice AI crossed a threshold recently. Phone agents that once sounded robotic and broke the moment a caller went off-script can now hold natural conversations: handling interruptions, understanding accents, checking a calendar and booking an appointment in one call.
For businesses that handle lots of calls — clinics, home services, logistics, insurance, hospitality, sales teams — that changes the economics of answering the phone.
This guide explains how AI voice agents work, why latency matters so much, what they cost, and what it takes to build one that callers don't hang up on.
What is an AI voice agent?
An AI voice agent is software that holds spoken conversations with people — usually over the phone, sometimes in a web or mobile app — and can take actions during the call: book appointments, look up orders, update records, qualify leads or transfer to a human.
It's different from an old-style IVR ("press 1 for billing"), which follows a fixed menu. A voice agent understands natural speech and decides what to do, which makes it a type of AI agent with a voice interface.
How do AI voice agents work?
There are two main architectures.
The cascaded pipeline
Most production voice agents chain three components:
- Speech-to-text (STT): transcribes the caller's speech in real time (e.g. Deepgram, AssemblyAI, OpenAI or Google speech models)
- Language model (LLM): understands the request, decides what to do, calls tools and writes the reply
- Text-to-speech (TTS): turns the reply into natural-sounding speech (e.g. ElevenLabs, Cartesia, OpenAI or Google voices)
Around them sit:
- Telephony: phone numbers and call handling (e.g. Twilio, Telnyx, SIP trunks)
- Orchestration: manages turn-taking, interruptions and streaming between components — using platforms like Vapi or Retell, or frameworks like LiveKit Agents or Pipecat
- Integrations: calendars, CRM, practice management, order systems
The cascaded approach lets you pick the best component for each job, inspect transcripts at every step, and swap providers independently.
Speech-to-speech models
Newer realtime speech-to-speech models (such as OpenAI's Realtime API and Google's live audio models) process audio in and audio out with a single model. They can respond faster and capture tone and emotion better, but give you less control over each step and can cost more per minute.
| Cascaded (STT → LLM → TTS) | Speech-to-speech | |
|---|---|---|
| Latency | Good with streaming and tuning | Often lower |
| Control & debugging | High — text at every step | Lower |
| Voice choice | Any TTS provider | Limited to the model's voices |
| Component flexibility | Swap any piece | Single provider |
| Cost | Usually lower, more predictable | Often higher |
Why does latency matter so much for voice AI?
In human conversation, the gap between one person finishing and the other responding is very short — often a few hundred milliseconds. When a voice agent takes two or three seconds, callers assume it didn't hear them, start talking again, and the conversation falls apart.
A practical target is to start speaking within roughly 500 milliseconds to one second after the caller stops. Getting there requires:
- Streaming everything: transcribe while the caller speaks; start TTS on the first words of the LLM's reply
- Fast turn detection: knowing quickly and accurately when the caller has finished
- Fast models: smaller, low-latency LLMs for most turns
- Short replies: spoken responses should be brief
- Co-located infrastructure: components running close to each other and to the telephony provider
- Filler for slow actions: "Let me check that for you" while a tool call runs
What makes a voice agent feel natural?
- Barge-in: the caller can interrupt, and the agent stops talking and listens
- Robust understanding: accents, background noise, poor phone lines, numbers and spelled names
- Confirmation of key details: reading back dates, phone numbers and amounts
- Natural voice and pacing: appropriate pauses, no robotic reading of lists
- Graceful hand-off: transferring to a human with a summary, not making the caller repeat everything
- Honest identity: clearly saying it's an AI assistant
What can AI voice agents do for a business?
| Use case | What the agent does |
|---|---|
| Appointment scheduling | Books, reschedules and cancels; sends confirmations and reminders |
| After-hours reception | Answers every call, captures details, handles urgent routing |
| Order and delivery status | Looks up orders and explains status and next steps |
| Lead qualification | Answers inbound enquiries, qualifies, books sales calls |
| Patient intake & follow-up | Collects information, confirms appointments, runs post-visit check-ins |
| Payment reminders | Reminds customers and routes them to secure payment options |
| Internal helplines | Answers staff questions about policies or IT |
We cover more workflows in 12 business workflows to automate with agentic AI, and our healthcare work shows voice alongside WhatsApp, web and email.
How much do AI voice agents cost?
Running costs
Per-minute costs combine:
- Speech-to-text
- LLM tokens
- Text-to-speech (often the largest component for premium voices)
- Telephony
- Orchestration platform fees, if you use one
Depending on choices, all-in costs commonly fall between a few cents and a few tens of cents per minute. Premium voices, frontier models and speech-to-speech models push it up; self-managed orchestration and efficient models bring it down. Many of the techniques in how to cut LLM API costs apply.
Build costs
A focused voice agent for one use case with a couple of integrations is similar in scope to a focused text agent, plus extra work on latency, telephony and conversation testing. Complex multi-intent agents with many integrations cost more. See our AI agent cost breakdown for typical ranges.
What are the compliance requirements for AI voice agents?
Rules vary by country, state and use case. Common considerations:
- AI disclosure: many jurisdictions expect or require callers to be told they are speaking with an AI. The EU AI Act includes transparency obligations for AI systems that interact with people (EU AI Act checklist).
- Call recording consent: one-party vs all-party consent rules differ between jurisdictions.
- Outbound calling: in the US, the FCC ruled in February 2024 that AI-generated voices are "artificial" under the Telephone Consumer Protection Act, so prior consent rules for robocalls apply.
- Sector rules: healthcare (e.g. HIPAA in the US), financial services and debt collection have additional requirements.
- Data protection: transcripts and recordings are personal data under GDPR and similar laws.
This isn't legal advice — involve counsel for your markets — but design for disclosure, consent and data protection from the start.
How do you build and launch a voice agent?
- Pick one call type with high volume and clear outcomes (e.g. appointment booking).
- Collect real call recordings or transcripts to understand how callers actually speak.
- Design the conversation: goals, required information, confirmations, hand-off rules.
- Choose the stack for latency, voice quality, cost and compliance.
- Build integrations so the agent can complete the task, not just talk about it.
- Test at scale with simulated callers, accents, noise and interruptions — plus real staff.
- Launch gradually: overflow and after-hours calls first, with human fallback.
- Monitor every call: transcripts, outcomes, latency, hand-off rate and caller sentiment.
How we build voice agents at Keyved
We build voice agents as part of our AI copilots and assistants work, usually alongside WhatsApp, web chat and email so customers get the same assistant on every channel. Our stack typically combines Twilio telephony with best-fit STT, LLM and TTS providers or orchestration platforms like Vapi, tuned for latency, and connected to your systems through our AI agent layer.
See our healthcare workflow platform and other projects, or tell us which calls you want to automate and we'll estimate the per-minute cost for your volume.
Frequently asked questions
What is an AI voice agent?
An AI voice agent is an automated system that talks with people by voice, usually over the phone. It converts speech to text, uses a language model to understand and decide what to do, takes actions in connected systems, and replies with synthesised speech.
How much does an AI voice agent cost per minute?
Running costs depend on the speech recognition, language model, voice synthesis, telephony and platform you use. All-in costs commonly range from a few cents to a few tens of cents per minute. Build costs depend on integrations and conversation complexity.
What latency do AI voice agents need?
To feel natural, a voice agent should start responding within roughly 500 milliseconds to one second after the caller stops speaking. Delays beyond that make conversations feel awkward and lead callers to talk over the agent.
Can AI voice agents replace call center staff?
They can handle a large share of routine calls such as booking, order status, FAQs and lead qualification, and free staff for complex or sensitive calls. A well-designed voice agent hands off to a human with full context when needed.
Is it legal to use AI voice agents for calls?
Generally yes, with conditions. Rules vary by country and use: disclosure that the caller is speaking with AI, consent for call recording, and restrictions on automated outbound calls. In the US, the FCC ruled in 2024 that AI-generated voices count as 'artificial' under the TCPA, so outbound robocall rules apply. Get legal advice for your jurisdictions.