Voice AI Agents in Production, Architecture and Lessons

Vasim Gujrati
Solutions Architect, AI & Platforms, Unico Connect
In this article
Quick Answer
A production voice AI agent runs three integrated layers, namely automatic speech recognition (ASR) to transcribe user speech, a large language model (LLM) for reasoning and response generation, and a text-to-speech (TTS) engine to deliver the answer. Each layer adds latency. Run the layers one after another and the user waits 1.5 to 3 seconds, while a fully streamed pipeline can deliver first audio within about 800ms. The architecture choices below bring that wait down, and the failure modes they cover are what break a voice agent in production.
Key Takeaways
- The voice stack has three layers to fill (for example Whisper or Deepgram for ASR, Claude or GPT-4o for the LLM, ElevenLabs or Azure Neural TTS for speech synthesis), and the tool you pick at each one shapes your cost, latency and reliability
- Stream every layer if the agent has to feel conversational, since batch processing creates pauses users will not accept
- Plan engineering time for voice activity detection (VAD), the component most developers underestimate until background noise in production breaks their silence timer
- In a multilingual agent, detect the language before ASR, because an English optimised model turns Hindi speech into a useless transcript
- GDPR Article 13 requires you to tell users who collects their voice data and for what purpose when you collect it, and from May 2027 the DPDP Act in India allows processing only with consent or for a listed legitimate use, so put a consent gate on the first interaction
A voice agent demo is quick to build, and most of the engineering sits in the production system. The rest of this post covers the gap between the two, with the WhatsApp AI agents we built for a B2B logistics client as one example.
The Three-Layer Voice Stack
Every voice AI agent runs on the same basic architecture, and what you pick at each layer decides your cost, latency and reliability.
| Layer | Function | Tools We Use | Typical Latency |
|---|---|---|---|
| ASR (Speech Recognition) | Converts user audio to text | OpenAI Whisper, Deepgram, Google STT | 200-500ms |
| LLM (Reasoning + Response) | Generates the appropriate reply | Claude, GPT-4o, Gemini | 400-1,500ms |
| TTS (Speech Synthesis) | Converts text response to audio | ElevenLabs, Azure Neural TTS, Google TTS | 150-400ms |
| Sum of the layer ranges | The three layer latencies added together | 750ms to 2,400ms |
What We Built, B2B WhatsApp Voice and Ordering Agents
The clearest example we can share is a pair of WhatsApp AI agents we built for a B2B logistics client, a voice agent for customer enquiries and a separate ordering agent for repeat business customers.
Customers send WhatsApp voice messages, and the voice agent transcribes each enquiry, looks up real time data on the client platform and answers status, pickup window and billing questions, handing the conversation to a human with its context when needed. The ordering agent recognises returning customers, surfaces their typical order patterns and confirms each order in conversation before pushing it through the platform. Both agents are live, and they lower the cost to serve routine enquiries while capturing business orders in the channel customers already use.
The 5 Architecture Decisions That Matter in Production
1. Stream Everything or Introduce Unacceptable Latency
In a non-streaming pipeline, the user speaks, ASR transcribes the whole message, the LLM writes the full response and TTS renders all of the audio before the user hears a single word. That wait is 1.5-3 seconds at minimum. A streaming pipeline overlaps the layers. The LLM starts generating tokens as soon as ASR has a partial transcript, and TTS starts rendering audio once the first sentence is complete, so the user hears a response within 800ms of stopping speaking.
2. Voice Activity Detection Determines Turn Quality
Voice activity detection (VAD) decides whether the user has finished speaking. Most demos get by with a simple silence timer, which fails constantly in production once background noise gets into the audio. Silero VAD running client-side classifies each audio frame as speech or non-speech in real time, at under 10ms per frame.
3. Language Detection Must Happen Before ASR, Not After
A common mistake is to send all audio to English ASR and then try to detect the language from the transcript it produces. If the ASR model is English-optimised and the user spoke Hindi, that transcript is useless. Run a lightweight language identification model on the first 2 seconds of audio instead, and use its result to route the audio to the right ASR model.
4. Fuzzy Entity Matching for Domain-Specific Vocabularies
General-purpose ASR models struggle with product codes and domain-specific terminology. Add a post ASR correction layer, a fuzzy matcher that compares ASR output against the product catalogue so each misheard code maps to the closest catalogue item.
5. Compliance and Consent Are Architecture, Not Afterthought
Voice data is personal data under GDPR (EU), the DPDP Act (India) and PDPA (Singapore). Put a consent gate on the first interaction and audit logging on every voice transaction, and store audio in AWS Mumbai (ap-south-1) when you want voice data to stay in India.
Tool Selection
For ASR, Deepgram Flux has the best latency for real-time streaming. OpenAI Whisper, when self-hosted, is the most cost-effective option at scale and has the best multilingual support, while Google STT v2 performs well on Indian English.
On the TTS side, ElevenLabs produces the highest quality voices and Azure Neural TTS offers strong price-performance at enterprise scale. Google WaveNet is cost-effective for high-volume GCP deployments.
Our guide to building AI agents with Model Context Protocol goes deeper on production AI agent architecture. When you evaluate a conversational AI partner, ask them to demo the specific failure modes above, such as noisy audio, mixed languages and unfamiliar product codes, as well as the happy path. For the wider selection process, see how to choose an AI development company.
Frequently Asked Questions
What is the best text-to-speech engine for production voice AI agents?
ElevenLabs produces the most natural-sounding voices, which makes it the best choice for customer-facing use cases where voice quality drives trust. Azure Neural TTS offers the strongest price-performance at enterprise scale, and Google WaveNet is acceptable for cost-sensitive high-volume use cases.
How do I keep voice AI agent latency under 2 seconds?
Stream every stage of the pipeline, which means streaming ASR, streaming LLM token generation and streaming TTS synthesis. Add sentence boundary detection so TTS only renders complete sentences, and use lower-latency LLM variants such as Claude Haiku or GPT-4o mini for conversational turns.
What compliance requirements apply to voice AI agents?
Voice recordings are personal data under GDPR (EU), the DPDP Act (India) and PDPA (Singapore). The main requirements are a lawful basis such as consent before recording, transparent disclosure of AI processing, data retention limits and audit logging. Build the consent architecture first, then the agent.
Can voice AI agents handle multiple languages in one conversation?
Yes, with the right architecture. Run a language detection model on the audio stream before it is routed to ASR. When a speaker switches between languages within the conversation (code-switching), multilingual ASR models like Whisper large-v3 handle it better than routing to separate single-language models.
What is the difference between a voice chatbot and a voice AI agent?
A voice chatbot follows a scripted decision tree, while a voice AI agent reasons about each request. It understands intent, retrieves relevant information, makes decisions within defined parameters and handles unexpected inputs with a response that fits the context.
How do I measure whether a voice AI agent is actually working?
Track task completion rate, fallback trigger rate (escalations to a human), ASR confidence distribution and user correction rate. Latency metrics tell you whether the system is fast, and completion rate tells you whether it is useful.




