Glossary

What is Voice AI?

Voice AI is a family of models and APIs that handle spoken language end-to-end: speech recognition (STT), text-to-speech (TTS), voice cloning, and real-time conversational models that listen, reason, and respond as audio. Production voice AI powers call-center agents, accessibility tools, and voice-first apps.

By Antoni Elkenbracht·Updated January 15, 2026

The voice AI stack

  1. Voice-activity detection (VAD). Detects when a human is speaking vs silent. Run client-side to save bandwidth.
  2. Speech-to-text (STT). Streaming transcription. Streaming is mandatory for interactive agents.
  3. Language model. Generates the response, often the same LLM that backs your chat product.
  4. Text-to-speech (TTS). Synthesizes the audio response. Streaming first-token latency matters more than overall throughput.

Speech-to-speech models (e.g. OpenAI gpt-4o-audio, Google Gemini Live) collapse steps 2 to 4 into a single multi-modal model. Lower latency, slightly less controllability.

Use cases where voice AI is winning

  • Outbound and inbound call agents. Sales qualification, appointment booking, support triage. Average handle time reductions of 20-40% are widely reported.
  • Drive-through and IVR replacement. Conversational voice replaces dial-tree menus.
  • Accessibility. Read-aloud, dictation, and conversational interfaces for low-vision and motor-impaired users.
  • Healthcare ambient scribing. Capturing patient encounters as structured notes, via tools like Nuance DAX, Abridge, Suki.
  • Localization. Voice cloning plus translation for multi-language content scaling.

Production reliability checklist

Production voice systems need: WebRTC or low-latency websockets between client and server, jitter-buffer tuning to handle network variance, a fallback path when the real-time model degrades, transcript logging for QA, and explicit consent capture for voice cloning and recording. GDPR and HIPAA have strict requirements around audio data handling; default to short retention and named-purpose processing.

Market context

Voice AI is the fastest-growing enterprise AI category in 2025 and 2026 according to Bessemer and a16z teardowns, with conversational agents replacing or augmenting human call centers at meaningful scale. Cost-per-minute is on a steep downward curve: the realtime audio APIs that launched in 2024 at $0.18/min are now available at $0.06/min, and that trend is accelerating with on-device inference.

How Enqualia handles Voice AI

Enqualia exposes this capability as a versioned, metered API on isolated Cloud Run. Read the deployment guide to see schemas, latency targets, and pricing.

Read: Multi-Modal APIs

Frequently asked questions

What are the main voice AI capabilities?
Speech-to-text (transcribing audio, via OpenAI Whisper, Deepgram, AssemblyAI), text-to-speech (synthesizing audio, via ElevenLabs, OpenAI tts-1, Google Chirp), real-time speech-to-speech (full-duplex conversational models like OpenAI gpt-4o-audio and Google Gemini Live), and voice cloning (creating a synthetic voice from a sample, via ElevenLabs Instant Voice Cloning or OpenAI Voice Engine).
What latency targets matter for voice agents?
Conversational agents need end-to-end latency below ~800ms for a natural turn-taking feel. That is split across VAD (voice-activity detection, ~50ms), STT (~200-300ms), LLM completion (~200-400ms), TTS first-token (~150-300ms). Speech-to-speech models like gpt-4o-audio collapse the pipeline and reach 200-400ms.
How is voice AI priced?
STT: $0.003-$0.01 per minute of audio. TTS: $0.015-$0.30 per 1,000 characters of input. Real-time speech-to-speech: priced per minute (input and output separately), typically $0.06-$0.10 per minute combined. Voice cloning has a one-time setup fee plus per-character synthesis cost.
What are the production hazards in voice AI?
Background-noise robustness varies wildly between STT models; speaker overlap breaks most systems; voice cloning is a fraud vector that requires consent verification; TTS hallucination can cause an agent to say something the LLM did not intend (mismatched audio vs transcript). Most production deployments add an STT-verification loop on TTS output for compliance-sensitive use cases.

Related concepts