Skip to main content

Learn

Plain-English guides to Vapi voice agents, per-minute pricing, latency tuning, and the broader voice AI ecosystem. No fluff, no affiliate spam.

Vapi basics

Compare platforms

Optimization

How voice AI works

FAQ

What should I read first to learn Vapi?
Begin with how Vapi voice agents work to understand the real-time STT, LLM, and TTS pipeline, then read the per-minute pricing breakdown so you can budget calls, and finish with the latency checklist before you launch to real callers.
Do these guides cover other voice AI platforms?
Yes. The comparison guide covers Vapi, Retell AI, and Bland AI side by side on latency, pricing, provider flexibility, and outbound throughput, and the glossary applies to any real-time voice agent stack.
Are these guides independent of Vapi?
Yes. vapi.health is an independent status mirror and reference site with no affiliation to Vapi, Retell AI, Bland AI, or ElevenLabs. Guides are written from public documentation and hands-on testing.
What is Vapi used for?
Vapi is an orchestration platform used to build, test, and deploy conversational voice AI agents for inbound customer support, outbound calling, and interactive voice assistants over phone networks or web applications.
How does Vapi compare to building directly on OpenAI or Twilio?
Vapi manages low-latency streaming orchestration between speech-to-text, LLMs, and text-to-speech providers, eliminating the complex engineering needed to sync audio streams, manage interruptions, and minimize voice latency manually.
What are the core components of a Vapi voice agent?
A Vapi voice agent consists of a Speech-to-Text transcriber (like Deepgram), a core reasoning LLM (such as OpenAI or Anthropic), a Text-to-Speech voice engine (like ElevenLabs or Cartesia), and telephony transport via SIP/Twilio.