Skip to content

What an AI Phone Call Costs Per Minute in 2026

September 16, 2026. The unit everyone quotes for an AI phone agent is the minute. Almost nothing inside that minute is actually billed in minutes. The phone line is, and the speech recognition is, but the voice is billed per thousand characters, the model is billed per token, and the orchestration layer that glues the four together charges its own separate rent. Prices read today from the vendors' own pricing pages put a real number on that rent, and it is the largest single line in a cheap AI call.

The four units hiding inside one minute

A self assembled AI phone call has four legs, and every one of them meters a different thing.

  1. Telephony bills per minute of connected call. Twilio's Pricing Voice API documentation returns 0.0085 dollars per minute for an inbound US local number, 0.022 for inbound toll free, and 0.013 for an outbound US call. Vapi's transport table, read the same morning, lists Telnyx at 0.0055 per minute, less than half Twilio's outbound rate for an identical commodity.
  2. Speech to text bills per minute of audio, and it is the cheapest leg in the stack. Deepgram's pricing page puts streaming Nova-3 at 0.0048 dollars per minute today, against a list price of 0.0077. ElevenLabs Scribe v2 Realtime is 0.39 dollars per hour, which is 0.0065 per minute.
  3. Text to speech bills per thousand characters, not per minute. Deepgram Aura-2 is 0.030 dollars per thousand characters, Aura-1 is 0.0150, Flux TTS is 0.0450. On ElevenLabs' API pricing page, Flash and v3 Conversational are 0.05, and v3 and v2 Multilingual are 0.10.
  4. The language model bills per token. Vapi's calculator converts that to a per minute range for OpenAI of 0.0077 to 0.0452 dollars, a spread of almost six times on one leg of four.

Two of those four meter the clock your customer is watching. The other two meter things your customer never sees, which is why almost every all in figure you will read online is a guess wearing a decimal point.

The vendor published the conversion for you

The character world and the minute world can be reconciled, and the bridge comes from a vendor rather than from us. ElevenLabs' model pricing table prints the conversion beside the price: 0.10 dollars per thousand characters is labelled as about 0.10 per minute, and 10,000 included characters is labelled as about 10 minutes. A thousand characters of synthesised speech is roughly a minute of talking. Apply that to Deepgram Aura-2 and a minute of continuous agent speech costs 0.030 dollars. Apply it to ElevenLabs Flash and the same minute costs 0.05.

The talk ratio nobody asks you for

Here is the assumption buried in every voice AI price you have been quoted. Text to speech is the only leg that bills for the agent's half of the conversation. Telephony and speech recognition bill the whole clock whether the agent is talking or listening. So the share of the call your agent spends speaking moves your bill, and not one pricing calculator we opened this morning asks you what that share is.

Work it through on ElevenLabs Flash at 0.05 per thousand characters. An agent that talks for 40 percent of the call spends about 400 characters a minute, so 0.020 a minute on voice. Push it to 60 percent, which is what happens when an agent reads back an address, confirms an appointment window and recites a disclosure, and the same leg costs 0.030 a minute. That is 50 percent more on the most expensive variable component, caused entirely by script length, and it will not show up in any quote you were given.

What the bundles charge to do the gluing

Now price the other road. Deepgram's Voice Agent API sells the whole pipeline at 0.075 dollars per minute on the Standard tier, 0.163 on Advanced, and 0.050 on the Custom tier where you bring your own model and your own voice. ElevenLabs' Speech Engine is 0.08 per minute, with burst pricing at 0.160, exactly double. Vapi charges a 0.05 per minute hosting fee and passes the four components through at cost.

The Custom tier is the revealing one, because on that tier Deepgram is selling you orchestration plus its own speech recognition and nothing else. Deepgram's own streaming rate for that recognition is 0.0048 a minute. Subtract it from the 0.050 bundle and the gluing alone is about 0.045 a minute, roughly nine times the speech recognition it wraps.

Then look at Vapi, a company with a completely different pricing model, and read its hosting fee: 0.05 per minute. Two vendors who agree on nothing else arrive at about five cents a minute for the part that neither listens nor thinks nor speaks. Orchestration is the biggest line item in a cheap AI call.

Why one vendor's own calculator returns a range

Vapi's calculator is unusually honest, and it is worth reading as a document rather than a quote. Set it to 1,000 minutes a month and it returns 82 to 129 dollars. The top of its own estimate is 57 percent above the bottom, for the same thousand minutes on the same platform.

The spread has a single source. Hosting is fixed at 50 dollars and transcription is near fixed at about 9.50, so roughly 73 percent of the low estimate cannot move. Everything that moves is the model, quoted at 8 to 45 dollars, and the voice, quoted at 15 to 24. In other words the two legs that are not billed per minute are the two legs that decide whether your per minute cost is eight cents or thirteen.

What it means for operators

Assemble the cheap parts at list and a minute of outbound US calling comes to about 0.033 dollars: 0.0055 of Telnyx, 0.0048 of Deepgram streaming, 0.0077 of model at Vapi's low figure, and 0.015 of Aura-2 for half a minute of speech. Swap in ElevenLabs Flash and it is about 0.043. Buy the same minute as a bundle and it is 0.075 to 0.129. That gap is the price of not maintaining it, and for most small businesses it is worth paying. It stops being worth paying at volume, which is the only real decision here.

Three things matter more than the per minute number, and they are the ones that surprise people.

  1. Concurrency is the ceiling, not minutes. Vapi's usage tier includes 4 concurrent calls and sells extra lines at 10 dollars per line per month. Deepgram's Voice Agent API tops out at 45 concurrent connections on pay as you go and 60 on Growth. Minutes are elastic. Lines are not. If you run AI calling for home services, a storm does not spread its calls politely across the month, it delivers them in the same twenty minutes, and four lines will take four of them.
  2. Compliance is a step function, not a rate. Vapi lists HIPAA eligible data handling as a 2,000 dollar per month add on. At a thousand minutes that single line is 2.00 a minute, more than fifteen times the entire rest of the estimate. Regulated work does not buy minutes, it buys a floor, and the same logic applies to consent capture and disclosure under TCPA compliant AI calling.
  3. Some of today's prices are promotions. Deepgram's streaming table is explicitly on limited time promotional rates: Nova-3 at 0.0048 is 38 percent below its own 0.0077 list. Its Flux TTS matching credit offer ends December 31, 2026. Build the business case on the list price, and treat the promotional rate as margin you get to keep for a while.

How to price your own calls this week

Take your last hundred recorded calls and measure three numbers before you sign anything: mean call duration, the percentage of that duration your agent spends speaking, and your peak concurrent calls in a single fifteen minute window. Those three numbers turn every table above into your actual bill. Duration multiplies telephony and recognition, talk share multiplies voice, and peak concurrency decides which plan you are allowed to buy at all.

Then apply the rule the arithmetic keeps producing. Below roughly a few thousand minutes a month, the bundle wins, because five cents of orchestration is cheaper than an engineer. Above that, the orchestration rent starts to look like a salary, and assembling the parts yourself becomes a real option. We build both, and the honest answer usually depends on your concurrency profile rather than your minute count. If you want a second opinion on your own numbers, our team builds AI voice agents against exactly this arithmetic.

One last caution about every figure on this page, including ours. The 0.033 dollar build cost assumes the agent speaks half the time, that a thousand characters is a minute, and that nothing retries. Change any of those and the number moves. The reason vendors quote a range is not evasion, it is that a minute of conversation is not a unit anyone actually sells.

Want an AI voice agent priced against your real call data?

We design, build, and run it for you, integrated with the tools you already use. Free audit in 24 hours.

Get Your Free Audit

Frequently Asked Questions

Assembled from parts at list prices read on 16 September 2026, an outbound US minute costs roughly 0.033 to 0.043 dollars: about 0.0055 for telephony on Telnyx, 0.0048 for Deepgram streaming speech to text, 0.0077 for the model at Vapi's low OpenAI figure, and 0.015 to 0.025 for text to speech assuming the agent speaks half the call. Bought as a managed bundle it is 0.075 to 0.129 dollars a minute. The gap is the orchestration fee.

Because two of the four legs are not billed in minutes. Text to speech is billed per thousand characters and the language model is billed per token, so the per minute cost depends on how long your script is and how big your model is. Vapi's own calculator returns 82 to 129 dollars for the same 1,000 minutes, a 57 percent spread, and all of that spread sits in the model and the voice.

Speech to text. Deepgram's streaming Nova-3 is 0.0048 dollars a minute at today's promotional rate and 0.0077 at list, and ElevenLabs Scribe v2 Realtime is 0.39 an hour, which works out at 0.0065 a minute. Recognition is typically under six percent of the bill. The orchestration layer, which does none of the listening or speaking, is usually the largest single line.

Below a few thousand minutes a month, buying wins. Orchestration costs about five cents a minute from both Vapi and Deepgram, which is far cheaper than the engineering time to maintain a four vendor pipeline. Above that volume the same five cents starts to resemble a salary, and assembling the parts becomes defensible. Concurrency requirements usually decide it before minute volume does.

About one thousand. ElevenLabs publishes the conversion on its own API pricing page, labelling 0.10 dollars per thousand characters as roughly 0.10 per minute and 10,000 included characters as roughly 10 minutes. Because text to speech only bills for the agent's share of the conversation, a minute of call time usually costs less than a minute of speech.

Three. Concurrency limits, because Vapi's usage tier includes only 4 concurrent calls and extra lines are 10 dollars each per month. Compliance add ons, because HIPAA eligible handling is listed at 2,000 dollars a month, which dwarfs per minute cost at low volume. And promotional pricing, because Deepgram's current streaming rate is 38 percent below its own list price and is labelled limited time.

Free Strategy Audit

Ready to put this to work?

Join 200+ businesses already scaling with AI and automation. Get your free audit and a custom roadmap within 48 hours.

Website & marketing performance analysis
AI & automation opportunity mapping
Custom growth roadmap with ROI estimates
Delivered within 48 hours, 100% free
200+
Clients served
48hr
Turnaround
100%
Free, no strings

Get Your Free Audit

Takes 30 seconds. No credit card required.

Prefer to chat?

WhatsApp us