Reduce what Connect costs you
Voice dominates, and inside voice the dominant term is not what the model says but what it carries: a live session is re-billed for its entire context on every turn, and measured production calls were 92% carried context by input token. Shorter replies, a bounded prompt, fewer calls that were never conversations, and research that only goes deep where it changes a decision are the four levers that matter.
Where the money is#
Two facts explain most of a voice bill. Audio tokens are priced several times higher than text tokens, on the way in and on the way out. And a live session is re-billed for its whole context every single turn, so a long call does not cost linearly in what was said — it costs in what is being carried, repeatedly. On production calls that carried context was 92% of input tokens.
Cost is metered from what the model actually reports — audio tokens in, text tokens in, reasoning tokens out, session duration — rather than from a field that sounds right. Looking for an audio-duration figure the provider does not emit once recorded zero seconds for every call ever metered, which is why the ledger now reports what was measured and says what was not, instead of estimating. A call that reports no tokens at all is still charged from its own duration rather than at zero.
The four changes that move the number#
Shorten what the voice says. Reply length is an engine setting rather than a request, so it is enforced. Median words per reply moved from the twenties and thirties to nineteen when this was set deliberately.
Result Output audio tokens fall directly, and calls get shorter, which compounds because every turn re-bills the context.
Bound the prompt. Knowledge carried into a call is budgeted by characters and by number of facts, memory and the contact block have their own character budgets, and a person's whole voice persona is kept under roughly 2,500 characters.
Result Less context to re-bill on every turn, and a faster first token: instructions around 7,500 characters produced a first token in 1.2 to 1.8 seconds, and around 9,600 characters pushed it to 2.3 to 3.4 seconds.
Stop paying for calls that were never conversations. Set the line's hours so out-of-hours calls get the closed-line message and end, and let the stale-session sweep hold sessions to their cap — a browser call once sat active for nearly twenty-five hours.
Result Refusals are cheap and recorded; abandoned sessions are closed and charged for what they were rather than left running.
Route research spend. A deeper prospect pass runs only where it can change a decision, and identity resolution stops an existing customer being researched as a stranger.
Result The expensive passes land on the candidates where the answer is still open, which is the only place they can pay for themselves.
Things that look like savings and are not#
- Turning the daily allowance down
- It does not reduce work, it defers it. Held mail is still mail somebody must deal with, and the read cursor does not advance, so tomorrow starts behind.
- Cutting Knowledge to shrink the prompt
- Knowledge is what stops answers being improvised. Removing it saves tokens and buys escalations, refusals and wrong answers — the budgets exist so you can keep the material and bound what is carried.
- Shortening calls by hurrying the model
- A steering direction to be quick is a request, not a control, and the measured floor on the realtime path is the model's own first token plus its end-of-turn detection — 3.3 seconds median on the best call measured. Chasing below that costs an afternoon and saves nothing.
- Turning the engine off overnight
- It stops new work but leaves a live line unhandled. Closing the line is the cheaper and better-behaved version of the same intention.
- Switching everything to
draft_only - Drafting is where much of the model cost is, so this saves the send and not the thinking — while producing replies nobody receives.
Knowing whether it worked#
Change one lever and compare like with like: the same kind of call, over the same sort of week. Per-call cost is the honest unit, because total spend moves with volume and volume is the thing you are usually trying to increase. A budget check runs before an outbound call is placed and before an inbound one is answered, so a workspace at its ceiling shows up as refusals — with the spoken closed-line message rather than silence — which is a signal to read rather than a fault to fix.
Resist judging by a single call. Call durations vary far more than settings do, and one long conversation with a good customer is not a regression. If a number moves and you cannot attribute it, the decision log and the call records carry what was in force at the time, including the settings snapshotted on each call.
Questions#
Why is a long call disproportionately expensive?
Because the session is re-billed for its entire context on every turn. The twentieth exchange carries everything from the first nineteen, so cost grows faster than length. Keeping replies short and the carried context bounded is the same lever applied twice.
Does compressing the context happen automatically?
Compression is sized rather than left at a default, and the default trigger is the model's whole context window — which means it is a deliberate setting, not something to assume is already working for you.
Can I see the cost of one specific call?
Per-call usage profiling exists and is deliberately operator-only, like the rest of the ledger. What it reports is what was measured; where something was not measured it says so rather than filling the gap with an estimate.
Is email cheaper than phone?
Written channels do not carry audio tokens and do not re-bill a live session per turn, so per interaction they are a different order of cost. The constraint on mail is the daily allowance rather than per-message spend.