7 min read

What Is a Voice AI Concurrency Limit, and How Does It Affect Agency Pricing?

A concurrency limit caps how many AI voice calls can be live at the same moment across your account, and it quietly decides how many clients you can safely resell to and what you can promise them.

A voice AI concurrency limit is the maximum number of calls your account can have live at the same time, inbound and outbound combined, and for an agency it matters because that one pool is shared by every client you resell to. It is not a monthly minutes cap or a per-call fee. It is a ceiling on simultaneous conversations, and when call number N+1 arrives at the wrong moment, something has to give.

How is concurrency different from minutes or call volume?

Minutes measure how much talking happens over a billing period. Concurrency measures how many conversations overlap at one instant. An account can use very few minutes a month and still hit its limit if calls cluster, or burn through thousands of minutes without ever coming close.

Each live call holds a slot for its full duration: ringing, the AI speaking, the caller pausing, and the wrap-up. Vapi, for example, documents concurrency as an account-level limit that applies to calls in progress, and other platforms follow similar logic. Check each vendor's own page for the exact behavior, because the details differ.

Why does concurrency affect what an agency can charge?

Concurrency is real infrastructure cost for the platform. Every live call holds open a speech-to-text stream, an LLM session, a text-to-speech stream, and a carrier channel. Vendors cannot offer unlimited simultaneous calls for free, so they either bundle a fixed number into each plan tier or sell extra capacity as an add-on.

For a reseller, that creates two pricing problems:

  1. Your clients share one pool. If you resell to eight local businesses, all eight draw from the same ceiling. Your cost does not scale cleanly per client the way minutes do.
  2. Your promises can outrun your capacity. If you sell "never miss a call" or "unlimited calls" to every client, you have made a commitment that only holds while your combined peak stays under the limit.

You cannot price a reseller offer well from average usage. You have to price it from peak overlap. For the broader economics, see the actual margin on reselling AI voice as an agency.

How many concurrent calls does an agency actually need?

There is no universal number, and anyone who gives you one without knowing your client mix is guessing. The underlying math is old telephony engineering: the Erlang B model estimates how likely a caller is to be blocked given offered traffic and number of lines. You do not need to run the formula, but the lesson holds. Blocking probability depends on peaks and call length, not daily totals.

Here is a hypothetical scenario, not a benchmark. You resell to a plumber, an HVAC company, and a dental office. Individually, each rarely has more than a couple of calls at once. But an HVAC client during a heat wave, a plumber after a burst-pipe storm, and a dental office at 8:30 on Monday can all spike while you are also running an outbound campaign for a fourth client. Their peaks may not line up, or they may line up exactly when you least want them to.

Practical approach:

  • Estimate each client's busiest hour, not their daily average.
  • Note which verticals have correlated spikes (weather, holidays, local events).
  • Count outbound campaigns separately, because they are the easiest way to consume your whole pool on purpose.

What happens when a call hits the limit?

This is the question to ask any vendor before you sign, because behavior varies. Depending on the platform, the over-limit call might get a busy signal, get rejected at the carrier, queue briefly, or fail silently. The caller experience in each case is different, and so is your exposure as the brand on the front of it.

Mitigations worth designing in from day one:

  • A human or voicemail fallback configured per client, so over-limit inbound calls land somewhere useful.
  • Missed-call text-back, which can recover some callers who never reached the agent.
  • Scheduling outbound off-peak, so a dialer campaign does not compete with inbound traffic for the same slots.

Ask whether inbound and outbound share one pool or have separate allocations. If outbound can crowd out inbound, your most urgent caller (someone with a flooded basement) is the one who loses. Outbound guardrails matter here too. NovaReps ships an outbound safety stack, and the same logic of pacing and controls applies to capacity, so campaigns do not run unattended into your inbound headroom. Read the explanation of the outbound safety governor for how those controls work.

Is the platform limit the only limit?

No, and this catches agencies out. Several ceilings stack, and the lowest one wins:

  • Platform concurrency on your plan.
  • LLM provider rate limits. With BYOK (bring your own key), your own API keys carry your own provider rate limits, which can be an advantage or a constraint depending on your tier with that provider.
  • TTS and STT provider limits, which may be separate from the LLM.
  • Carrier channel limits on the phone numbers or trunks involved.

Per-tenant carrier isolation helps with the last one. When each client's telephony setup is separated, one client's carrier-side problem or traffic burst does not automatically take down another's lines. It does not remove the platform-level pool, but it narrows the blast radius. If you are weighing how that differs from other tools, the NovaReps vs Vapi comparison is a reasonable place to see how two platforms frame account-level capacity, and Vapi's own docs are the source of truth for their side.

How should you price around a concurrency ceiling?

A few rules that hold up regardless of platform:

  1. Do not sell "unlimited" unless you control the pool. Sell included minutes, a defined business-hours scope, or a tier tied to the client's expected peak.
  2. Tier by risk, not just by volume. A client whose business depends on every urgent call (emergency trades) warrants a higher tier and a clearer fallback plan than a low-urgency one.
  3. Put outbound on a schedule. Campaign windows should be an explicit part of your service terms.
  4. Review capacity when you add clients, not when something breaks. Adding your ninth client is a capacity decision, not just a sales win.
  5. Know your upgrade path and its cost before you need it. Check NovaReps pricing for how plans and capacity are structured, since a limit you can raise in minutes is a different risk from one that needs a sales cycle.

What should you ask a vendor before reselling on their platform?

Bring these questions to any vendor, including us:

  • What is the default concurrency on my plan, and what is the process and cost to raise it?
  • Is the limit per account, per sub-account, or per number?
  • Do inbound and outbound share the same pool?
  • What does the caller experience when the limit is hit?
  • Can I see live utilization and historical peaks so I can plan, rather than discover limits from client complaints?

If the answers are vague, treat that as information. Concurrency is not the most glamorous line item on a plan page, but it is the one most likely to turn a happy pilot into an embarrassing outage once you have several clients live. Price from your peaks, keep a fallback for every client, and treat capacity as part of the product you are reselling.

Built for Agencies & Resellers

Launch Your Own Voice AI Brand.

Your clients never see us. Your brand, your pricing, your recurring revenue. Zero per minute markups, ever.

No per-minute taxes Bring your own Twilio & OpenAI 100% White-Label Client Portals
Continue Reading

Related Articles & Insights

View all articles