Insights · Translation

Real-time call translation in Amazon Connect: what's actually possible

"Real-time translation" gets used for at least three different things, only one of which lets a monolingual agent and a caller who shares no language hold a conversation. Here's what separates them, what live two-way translation on an Amazon Connect call genuinely requires, and — for Australian buyers — why the data-residency question decides the shortlist before anything else does.

Post-call transcription is not live translation

Amazon Connect already does a lot with language, and it's worth being precise about what each capability is for.

Contact Lens transcribes and analyses calls, with post-call analysis for quality and coaching and a real-time capability that can surface rules and alerts to a supervisor during the call. It is superb for finding out what happened and for catching a call going wrong. It does not put words in the agent's mouth or the caller's ear.

Amazon Translate converts text between languages, and it is genuinely useful stitched into a chat flow, where the interaction is already text and a second or two of processing is invisible.

Neither of those is what a service manager means when they ask about translation. They mean: a Vietnamese-speaking caller rings, my English-speaking agent answers, and the two of them sort the problem out. That is a different engineering problem, because it has to happen inside the natural rhythm of a spoken conversation instead of alongside it.

Two lesser things also get sold as real-time translation. One-way translation — the agent sees an English transcript of what the caller said, and replies in English regardless — helps the agent understand and leaves the caller no better off. And translated callbacks, where the enquiry is captured and answered later through a translated channel, are a queue with extra steps. Both are legitimate products. Neither is a conversation.

What two-way live translation on a call actually requires

For a caller and an agent to converse across a language gap, five things have to happen continuously and in order, in both directions.

  1. Identify the language. Ideally from the caller's first utterance, without asking them to navigate an English menu to declare it — which is precisely the thing they can't do.
  2. Recognise the speech. Streaming recognition on telephony-grade audio, which is narrowband, compressed and frequently noisy. This is meaningfully harder than recognising a podcast, and it is where accents and code-switching — a caller dropping English words into another language, which is completely normal in Australian migrant communities — cause most errors.
  3. Translate the meaning. Not the words. Idiom, politeness registers and the way different languages order information all matter, and a literal rendering of a polite refusal can read as rude.
  4. Speak it. Synthesised speech played onto the live call, and only to the party who needs it — the caller hears their own language, the agent hears English, and neither hears the other's audio bleeding through. Getting that crossing wrong produces an unusable call, and it is the least glamorous, most decisive part of the build.
  5. Show both sides. A panel with the conversation in both languages, side by side. This is a trust feature, not a nice-to-have: it is how an agent, a bilingual supervisor or the caller themselves catches a mistranslation while it can still be corrected.

There is a sixth requirement that isn't about language at all: the call has to stay one contact. If translation is delivered by conferencing the caller out to another platform, you have split the interaction — routing context lost, recording in two places, reporting confused. Live Interpreter joins the same Amazon Connect contact as the caller and the agent, so the call remains a single routed Connect contact and your recording, analytics and reporting continue to apply exactly as they do today.

Latency: what a conversation will tolerate

Every stage above adds delay, and delay is the constraint that decides whether the result feels like a conversation or like a badly-connected international call.

The behaviour to design for is not raw milliseconds but turn-taking. Human conversation runs on very short gaps between speakers, and when those gaps stretch, people do predictable things: they repeat themselves, they talk over the translation, they assume the line has dropped, or they fall into a stilted walkie-talkie rhythm. All four degrade the call.

Three practical implications follow.

  • Endpointing matters as much as speed. Deciding when a speaker has finished is what determines when translation can start. Cut too early and you translate half a sentence; wait too long and you add dead air to every turn.
  • Geography is latency. Every hop between your Connect region and a speech endpoint on another continent is added round-trip time on every turn. An Australian call processed through a European or US endpoint pays that tax repeatedly. This is the practical argument for onshore processing even when compliance isn't driving it.
  • Prepare the agent. Agents who are told to leave a beat after speaking, and not to fill silence, get dramatically better calls than agents who are handed the tool with no coaching. Half an hour of training is worth more than any tuning.

Nobody should quote you a universal latency figure, because the tolerable gap varies with the caller, the language pair and how upset they are. Test it on your own calls, in your own languages, and listen to the recordings.

Data residency, and why it decides the shortlist in Australia

For Australian government, health and financial-services buyers, this section is usually the whole evaluation.

A translated call sends the caller's voice — and a transcript of it — to a speech recognition service, then to a translation model, then to a speech synthesis service. Each of those is a place the audio and its text exist. Most live-translation offerings route that through whichever cloud region the vendor happens to operate in, which for many is the United States or Europe. If your obligations require personal information to stay in Australia, that is not a detail to resolve during implementation. It disqualifies the product.

What makes an onshore path possible today is that the underlying pieces now exist in Australian regions — including a generally available Sydney speech-recognition endpoint and Australian-hosted speech synthesis — so recognition and synthesis can both run in-country rather than being shipped offshore and back. Live Interpreter is built to use them, which is why we can state the onshore claim rather than imply it. Your Amazon Connect instance, recordings and reporting never leave your own AWS account and region in either configuration.

Two things to check before you sign, whoever you buy from:

  • Which languages are actually covered onshore. Onshore coverage is good but not universal across every language pair, and the honest answer is per-language. Confirm the specific pairs you need.
  • Where transcripts and logs are retained, not just where processing happens. Those are separate questions and vendors frequently answer only the first.

Notice that this is the opposite trade-off from the one on AI voice agents, where the highest-quality voices are offshore — as we've written about ElevenLabs. Residency and voice quality pull in different directions, and which one leads is a business decision, not a technical one.

Where a human interpreter is still the right answer

We'd rather say this plainly than bury it in a disclaimer. Machine translation is built for speed and coverage on everyday service conversations — appointments, accounts, enquiries, status updates, the ordinary business of a contact centre. It is not a certified interpreter and does not carry a certified interpreter's accountability.

Use a certified human interpreter where the law, your regulator or your own policy requires one: legal proceedings, clinical consent, child protection, anything where a misunderstanding creates legal consequence. In Australia that boundary is reasonably well understood in health and justice settings and much less so elsewhere, so the useful test is simple — if a mistranslation on this call could produce a consequence you'd have to defend, use a person.

The realistic position for most contact centres is both. Machine translation handles the volume of routine multilingual contact that currently goes unanswered, gets an engaged signal, or waits on hold for an interpreter line. Certified interpreters handle the calls that warrant them. That combination serves more people than either alone, and it is a much easier proposition to defend internally than "we replaced interpreters with software".

The short version

Post-call transcription tells you what happened. Live two-way translation lets the conversation happen. Doing the second on Amazon Connect means recognition, translation and synthesis running continuously in both directions on a single routed contact, with both languages visible so errors get caught. For Australian buyers, decide the residency question first — it eliminates most of the market before you compare anything else. And keep certified interpreters for the calls that need them.

Serve every caller, with the team you already have

Live Interpreter adds real-time two-way translation to your Amazon Connect calls, with an onshore Australian speech path where residency matters. Listed on AWS Marketplace at US$300 per session.

See Live Interpreter   View on AWS Marketplace