Connect by JBRH Open Connect

Multilingual calling

A call opens in one language and can move to another on evidence. The realtime model speaks each one directly rather than switching a synthesiser, so the practical question is not which are supported but how much proof the move takes. That depends on what Connect already knows: a whole sentence where the language was remembered or unknown, two or three words where the contact's city makes one likely, and nothing at all on a pinned line.

Status
Available What this means
Audience
both
Channels
phone
In the app
#/calls
Last verified
Product version
6.3.2

What the call opens in#

Three settings meet at the start of a call, and their order surprises people: the greeting's script is what the call opens in, not the line's setting. The brief is rendered from the line, but the caller's first evidence comes from the sentence they actually hear — so if a workspace has written its greeting one way and set its line another, the greeting wins in practice and the mismatch is worth removing.

The line's setting
What the brief is written in, and the default for everything the model is told.
The greeting
Pre-synthesised audio, fixed before anybody has spoken. It is the first thing said and it sets the caller's expectation.
A pinned language code
An engine setting in the Voice Lab. When it is set the call stays there, and no steering note is ever added.

Two pieces of machinery, different reach#

This is worth separating. The realtime model is speech-to-speech: it understands and produces the speech itself, which is why there is no pitch dial and no per-locale voice to pick. Connect's own text-to-speech — used for the greeting, the closed-line message and the turn-based carrier path — is a separate service, where gemini-tts offers a single voice across the range with a natural-language style instruction and chirp3-hd offers one voice per locale with a speaking rate.

That second service is the reason the phone line stopped being English-shaped. Until 2026-09-04 every word a caller heard was the carrier's own text-to-speech: two voices, sixteen European locales, nothing Indian at all. Connect's own voice covers how that sentence is produced and served now.

Switching part-way through#

Two separate pieces of machinery decide this. voice_region.plan writes the language rule into the opening brief, and how much evidence a move takes is set there — differently in each of its four tiers. Mid-call, LanguageTracker appends one chat-context note (language_steer) when a reply comes back in a script the caller has not used. The rules around both are strict, and each one exists because the loose version misbehaved:

  • A remembered preference is spoken from the first reply; a whole sentence in another language is what overrides it — a greeting word is not a choice.
  • A mixed line is not a switch where the language is remembered or unknown. Hinglish is one way of speaking, not a request to change.
  • Where the contact's city or state makes a language likely, that same mixed line is the evidence. Two or three words in it are enough, and the reply stays in it from then on.
  • A pinned code is never steered. If somebody has said the call is held one way, no evidence overrides that.
  • A Telugu reply transcribed in Latin letters is still Telugu (romanised_script) — the script the transcriber reached for is not the thing being spoken.
  • One garbled line does not undo what is already established.

Detecting the caller's language gives the evidence rules in full, and light code-mixing and slang covers a Hinglish conversation that never resolves into one or the other.

What is verified today#

Everything above was reconciled on 2026-09-10 against backend/app/voice_region.plan as production runs it. The part most often mis-stated is that there is a single bar. There is not: a whole sentence is the bar for a remembered language and for a caller nothing is known about, two or three words — mixed with English or not — is the bar where the contact's city or state makes one likely, and a pinned line has no bar because nothing moves it. Each was arrived at by watching a looser or stricter rule fail on a real call.

Questions#

Can one line take calls in several languages?

Yes, unless it is pinned. The line has one setting for its brief and its greeting, and the call can move on evidence from there. Pinning is the way to say it must not — useful when a line exists precisely to be answered one way.

If the caller switches, does the voice change too?

On the realtime engine there is no separate synthesiser to change: the same model speaks the other one. The greeting is the exception, because it is pre-synthesised audio chosen before anybody has spoken.

Does a multilingual call cost more?

Not for that reason. A live call is billed on audio tokens with the whole context re-billed every turn, so what raises the cost is the number of turns and the size of the brief. What a call costs sets out the metering.