Insights · AI voice
Most AI voice pricing is quoted as a single per-minute number, which is convenient and misleading. There are four separate cost lines under every AI agent on Amazon Connect, three of them are usage-based, and only one is a licence. Here's the whole model, a worked example you can put your own numbers into, and the places the curve bends.
Whatever you buy and whoever you buy it from, an AI-handled call on Amazon Connect generates charges in four places.
1. Amazon Connect service and telephony. AWS bills Connect voice per minute for the service, plus telephony per minute for the number and the carriage — and telephony rates vary a lot by country, by inbound versus outbound, and between a direct-dial number and a toll-free one. This line is unchanged by putting an AI on the call. An AI-handled call costs exactly what a human-handled call of the same length costs on this line. The saving is agent time, not telephony. Price yours from the AWS pricing page for your region; don't accept a vendor's estimate for a rate AWS publishes.
2. Speech. Speech-to-text and text-to-speech, billed per minute or per character by whichever provider you use. This is where residency choices show up as money as well as compliance, because the onshore-capable providers and the premium-voice providers are usually not the same vendor.
3. The language model. Billed per token — input and output priced separately — or rolled into a vendor's per-minute figure. Voice conversations are short in token terms compared with document work, but the system prompt, the tool definitions and the running conversation history are re-sent on every turn, which is what makes long calls disproportionately expensive.
4. The agent licence. The product itself: a subscription, a bundle, or your own engineering team's time if you built it. Approach matters here — we compare the four build-versus-buy routes in how to add an AI voice agent to Amazon Connect.
Line 1 belongs to AWS. Lines 2 and 3 belong to AWS or a speech and model vendor. Line 4 is the only one a Marketplace listing price refers to. A quote that shows you one number without saying which lines it covers is not a quote, it's a hope.
Use your own rates; the structure is the point. Assume 1,000 AI-handled calls a month averaging four minutes each — 4,000 voice minutes.
Two habits make this exercise honest. First, work in minutes, not calls — average handle time drives three of the four lines, and a 40-second containment and a nine-minute conversation are not the same product. Second, model the abandoned and mis-routed calls too. A caller who says three words and asks for a person still consumes telephony and speech.
Four inflection points do more to your bill than any rate negotiation.
Average handle time. Three of the four lines are per-minute or per-turn. A well-scoped agent that resolves a call in ninety seconds is not marginally cheaper than one that meanders for five minutes — it is roughly three times cheaper, on every line except the licence. The cheapest optimisation available is a tighter prompt and a shorter greeting.
Concurrency, not volume. Speech and model providers sell concurrency tiers. Ten thousand calls spread evenly across a month is a small problem; the same ten thousand arriving between 9am and 11am on the first Monday after a billing run is a different tier. Size against your peak concurrent calls, not your monthly total, or your first busy day will be an outage rather than an invoice.
Containment rate. A call the AI resolves costs you four lines once. A call the AI works for two minutes and then transfers costs you four lines plus the full human handling cost. Below a certain containment rate an AI agent adds cost rather than removing it — which is why the warm handoff matters commercially and not just for customer experience. When the summary, intent and transcript land in the agent's CRM screen-pop, the human leg is shorter, so a transferred call is not a total loss.
Residency. Requiring an onshore Australian speech path narrows your provider list, and a narrower list is a weaker negotiating position. That is a real cost of compliance and it should appear in the business case rather than as a surprise in month two.
The same product bought two ways lands in different budgets, and for larger organisations that matters more than the sticker price.
Through AWS Marketplace, the subscription appears on the AWS bill you already receive and reconcile. There is no new vendor to onboard, no new payment instrument, and for many organisations eligible Marketplace spend can count toward an existing AWS commitment — worth confirming with your account team. Procurement time is often the real saving.
Direct gives you a normal commercial conversation: negotiated terms, invoicing in your own currency arrangements, and bundling across products. If you are also buying CRM connectors or wallboards, a combined quote is usually the better route — the pricing page lists the direct plans.
One planning note: ProUCX's AI products are priced in USD, so Australian buyers should build an FX assumption into a twelve-month business case rather than converting at today's rate and forgetting about it.
The comparison people reach for is "AI agent versus a new hire", and it's the right instinct done wrong. Don't compare against a salary; compare against your own fully-loaded cost per handled contact, which you can calculate from figures you already have: total cost of the team (salary, on-costs, supervision, recruitment, training, attrition, seat and licence) divided by contacts handled. Most contact centres have never worked this out and are startled by it.
Then set the AI's cost per handled contact — all four lines, divided by contacts it actually contained — against that number, and apply three adjustments most business cases miss:
A fair business case usually lands on "this changes what we can offer" rather than "this removes N people", and it survives scrutiny far better as a result.
Any vendor who can't answer all six in a single email is quoting you a hope too.
Four lines: Connect and telephony, speech, model, licence. Three are per-minute, which makes average handle time the dominant variable and containment rate the one that decides whether the whole thing is worth doing. Size on peak concurrency, not monthly volume. And if you want to see how a call actually behaves before modelling it, the demo console at dev.ai.liveucx.com runs one end to end.
See what Live AI Agent does, or subscribe on the AWS account you already have.