Regional speaking style
A regional style layer adds directions about how to speak — rhythm, formality, the politeness markers a place actually uses — on top of the voice already chosen. It is prompt text, so the model approximates it rather than executing it, and one carefully written cadence rule was measured, found to change nothing, and taken out. Which region applies is ranked on evidence — a contact's city or state, or an Indian landline's STD code — and never inferred from a mobile number.
What the layer is made of#
Everything in the layer is a steering direction rather than an engine parameter, and the difference decides what you should expect. Engine settings become session parameters and take effect exactly: a model id, a voice, a silence in milliseconds. Steering settings become sentences in the instructions, and the model does its best with them. Accent and pronunciation have always sat on the steering side, and so does everything regional.
The material comes from documented speech behaviour rather than impression: syllable-timed rhythm for Indian English, speaking rates that fall with age (122–168 words per minute across the range), filler-word frequency of one to three per hundred words, and the politeness markers a place genuinely uses — *andi*, *ji*, *-nga*, *po*. The same body of research sits behind human voice profiles, which is the whole-person version of the same idea.
What the brief actually asks for comes to a sentence: sound like a courteous local professional on a business call. Everyday speech; the English business words people there use in English left in English — *price*, *demo*, *meeting*, *software*, *installation*, *service*, *payment*; and the respectful forms the place genuinely has, *garu* after a name in Telugu, *ji* in Hindi, *-nga* endings in Tamil. Match how much the caller mixes languages, and get there gradually rather than in one jump.
- No imitated accent.
- No slang and no film phrases before the caller has used them.
- No caricature of the region.
- No assuming the regional language from where somebody is — a region makes it likely, and the caller's own words decide.
What it deliberately does not change#
- The language of the call
- A place and a language are different things. A region makes a language *likely*, never certain: the call still opens in the line's language, and moves only when the caller uses the regional one — two or three words are enough. That runs on the evidence rules in detecting the caller's language.
- The voice itself
- The realtime engine is speech-to-speech. There is no synthesiser underneath with a pitch dial, so nothing here reaches into the timbre of the voice.
- Anything measured in Hz, dB or milliseconds of pause
- Those are in the unsupported list and are shown as unsupported. A slider that cannot move anything is worse than an absent one.
The cadence rule that was measured and removed#
The most instructive thing on this page is a failure. A respect-word cadence rule — placing *andi* the way a speaker would — was written, rendered into the instructions and measured. It changed nothing: the marker appeared in 6 of 8 replies with the rule and without it, and on one occasion an English sentence came back answered in Telugu.
It was rolled back rather than left in place looking like a feature. That decision is the honest half of the whole tuning surface: a direction the model approximates but does not follow is a control in appearance only, and shipping it would have had somebody adjusting a setting that does nothing while a real problem went unlooked-at.
What a setting cannot change collects these, and findings a setting cannot fix is the longer argument for reporting a limit instead of a dial.
Using it without over-spending#
- Every regional direction is prompt characters, and prompt characters are first-token seconds — 7,500 characters measured 1.2–1.8 s, 9,600 measured 2.3–3.4 s.
- A whole-person profile is capped under about 2,500 characters for the same reason. A style layer that grows past a paragraph is buying less than it costs.
- Judge it on a test call from the Voice Lab, where the resolved settings are snapshotted on the call so the review knows what actually ran.
This page was reconciled on 2026-09-10 against backend/app/voice_region.plan at f5ea518. Two things there are easy to get wrong from the outside: the number itself yields only the country, from the dialling code, plus a city when an Indian landline's STD code gives one — a mobile never names a region, because number portability made the series meaningless — and the style layer applies to a language the caller has actually used, not to one their address suggests. The parts established by measurement stand: what a steering direction can and cannot do, and the cadence rule that did not survive it.
Questions#
Can I make the voice sound like it is from a particular city?
You can ask for the way people there speak — rhythm, formality, the markers they use — and the model will approximate it. You cannot set an accent as an engine parameter, because on a speech-to-speech engine there is no synthesiser to set it on. Expect an approximation, and judge it by listening rather than by reading the setting.
Does the style layer decide which language the call is held in?
No, and keeping the two apart matters. Which one is spoken follows the greeting, the line's setting, a pinned code and the caller's own evidence — and where the contact's city or state is on file, two or three words of the regional language are enough to move it. The style layer only affects how it is spoken.
Is this available to customer workspaces?
Yes. The capability registry records it as available, and phone behaviour here is one implementation reached through two routers rather than an operator feature with a customer subset.