All articles
11 min read

The three levels of AI voice agents, and when to stop at each one

  • levels of AI voice agents
  • Vapi alternative architecture
  • build your own AI voice agent stack
  • AI voice agent SaaS architecture
Original source video by Shreyas Raj. This guide restructures the useful parts for readers and adds current implementation context.

Start with the job the call must complete

Higher is not automatically better. Most businesses should stop at Level 1 unless the workflow is proven. Most service providers can stop at Level 2 if they have a small number of bespoke clients. Level 3 is justified only when the same workflow must operate safely across many clients and someone can own reliability, security, billing and support.

The model in one table

  • 1. Hosted orchestrator | Prompt, knowledge, tools and call flow | Fast prototype or controlled pilot | Platform markup and platform limits | The workflow produces measurable qualified outcomes
  • 2. Owned modular pipeline | Speech, model, orchestration, infrastructure and integrations | Proven bespoke deployments needing control | Your team owns latency, routing, monitoring and failures | Multiple deployments repeat without manual rebuilding
  • 3. Productized platform | Shared product, tenancy, admin, billing, support and operations | A repeatable offer serving many clients | Security, isolation, billing and support complexity | Demand and operational maturity justify product ownership

Level 1: hosted orchestrator

The transcript starts with Vapi and Retell as examples of visual hosted platforms. Their advantage is speed: a builder can configure a prompt, voice and tools without first operating the entire realtime stack.

Stay here when the business question is still unproven. Optimize the call flow, measure outcomes and learn what fails. Moving to an owned pipeline before the workflow earns its keep creates engineering work without a business case.

Level 2: owned modular pipeline

At 06:08, the source separates the system into listening, reasoning and speaking. In production, add telephony, turn-taking, tool execution, state, observability and human handoff.

The transcript correctly identifies the tradeoff at 08:53: more control also means responsibility for provider routing, latency tuning and failure handling. This level is not merely a cheaper API bill.

Level 3: productized multi-tenant system

At 13:18, the video introduces a shared client dashboard and admin layer. A real Level 3 system also needs tenant isolation, role-based access, billing reconciliation, usage limits, safe support access, audit logs, secrets management, deletion and retention controls, incident response and migration paths.

An admin switch does not create compliance, and one dashboard does not create infinite scalability. Those claims should not appear in the article.

Cost claims need a date and full denominator

The recording compares per-minute figures at the time it was filmed, but vendor and model prices change. More importantly, an owned stack adds infrastructure, engineering, monitoring and support.

Publish a dated cost table only after current invoices and provider pages are checked. Otherwise use the full-cost worksheet and explain the components without a headline price.

Which level should you choose?

If the workflow is unproven, choose Level 1.

If the workflow is proven and platform limits materially hurt cost, control or required behavior, evaluate Level 2.

If the same stable workflow repeats across many clients and a team can own tenancy, billing, reliability, security and support, evaluate Level 3.

If none of those conditions is true, stop where you are.

What the video demonstrates, and what it does not prove

Demonstrated: a hosted agent, a modular realtime-agent call and a client plus admin dashboard walkthrough.

Not independently proven by the recording: universal language quality, current vendor price advantage, data-security superiority, HIPAA compliance, unlimited scalability or guaranteed demand.

Use the video as first-party build evidence and the decision matrix as the publication's added value.

Continue with the right implementation path

Use the educational guide to make the architecture and test decisions. Use the matching service or location page only when you want RapidXAI to scope and deploy the system.

Use it yourself

Architecture level decision card

Copy this into your project notes, then replace every blank or assumption with evidence from your own workflow.

WORKFLOW PROVEN WITH REAL CALLS: YES / NO
MEASURABLE BUSINESS OUTCOME: YES / NO
CURRENT MONTHLY CONNECTED MINUTES:
CURRENT FULL MONTHLY COST:
DOCUMENTED PLATFORM LIMIT:
WHY THAT LIMIT MATTERS:
TEAM OWNS REALTIME INFRASTRUCTURE: YES / NO
TEAM OWNS MONITORING AND ON-CALL: YES / NO
SAME WORKFLOW REPEATS ACROSS CLIENTS: YES / NO
TENANT ISOLATION DESIGNED: YES / NO
BILLING AND USAGE RECONCILIATION DESIGNED: YES / NO
ROLLBACK PATH:
DECISION
Unproven workflow: Level 1
Proven workflow with material platform limit and operating capacity: evaluate Level 2
Repeatable multi-client system with product operations: evaluate Level 3

Sources and further reading

Frequently asked questions

What are the three levels of AI voice agents?
Level 1 uses a hosted orchestration platform, Level 2 owns a modular voice pipeline and Level 3 productizes a proven pipeline for multiple tenants with billing and operations.
Is building your own AI voice stack always cheaper?
No. It can reduce platform markup, but it adds infrastructure, engineering, monitoring, support and failure ownership. Compare full cost at the same volume and reliability target.
When should I replace Vapi or Retell?
Only after the workflow is proven and a documented platform limit materially affects required behavior, control or full cost. A migration needs acceptance tests and a rollback plan.
What makes a voice agent platform multi-tenant?
Tenant isolation, roles, per-client configuration, usage metering, billing, audit logs, safe support access, secrets management, retention controls and reliable operations, not merely separate dashboards.
Can I stop at Level 2?
Yes. A service provider with a manageable number of bespoke clients may prefer Level 2 because it offers control without the product and support burden of a multi-tenant platform.
Does an admin compliance toggle make a platform compliant?
No. Compliance depends on the complete technical, contractual and operational system. A user-interface toggle alone proves nothing.

Want this working in your business?

Fifteen minutes. Your numbers, our honest read on what AI returns for you. No deck, no pressure.