Voice Lab
Voice Lab is the screen where a workspace tunes the realtime phone voice and reviews how it performed. Its settings split in two: some become parameters the model session is opened with, and the rest become written directions inside the instructions. The review half turns a finished call into measurements and findings, each paired with the setting that would address it.
Two kinds of control, and why the difference matters#
| Kind | Examples | How it takes effect |
|---|---|---|
| Engine | Voice name, model, language, temperature, end-of-speech sensitivity, interruption thresholds, reply length, silence timers | Session parameters the voice worker passes to the model |
| Steering | Speed, pitch, energy, warmth, pauses, fillers, formality, accent, pronunciation | Lines written into the system instructions |
The realtime engine is a speech-to-speech model: there is no separate synthesiser with a pitch dial to turn. So a steering control is a direction, followed well but not guaranteed, and the screen says so beside each one rather than presenting every control as if it were a switch. Knowing which kind you are adjusting is the difference between *this did not work* and *this was asked for and not always obeyed*.
Reviewing a call on evidence#
The review half reduces what the voice worker measured on the wire — turn timings, interruptions, silences, false interruptions — to the numbers the Lab shows: how long from the caller's last word to the first audio back, at the median and at the ninetieth percentile, how long the model took to its first token, how quickly an interruption stopped the voice, and how loaded the host was.
Nothing is estimated from the timestamps of transcript rows, and that restraint comes from a real failure: transcript rows arrive when a sentence is complete rather than when it started, and the first live call recorded a latency of eight milliseconds while the caller sat waiting for several seconds. A number computed from the wrong clock is worse than no number, because it is believed.
Findings are then produced deterministically — slow replies, long gaps, repetition, ignored interruptions, wrong language — each with the setting that would address it. The model is asked exactly one question, about delivery, and its verdict is labelled as the model's opinion rather than mixed in with the measurements.
- A reply at or under 2,500 ms is the target.
- Past 4,000 ms the reply is treated as slow and shows as a finding.
- About 60% of the delay on a well-tuned call is the model itself, which no setting removes — see findings a setting cannot fix.
Not a studio, not a profile, not the greeting voice#
- Not a recording studio
- Nothing is pre-recorded here. Every word on a call is generated live, and a change is heard on the next call rather than rendered into a file.
- Not the human voice profiles
- A profile is a named persona layered over the workspace's settings. Voice Lab is where the settings underneath live, and where a profile's effect is reviewed.
- Not Connect's own text-to-speech voice
- A separate capability used where a fixed line is spoken, such as a greeting prepared before the call connects. Voice Lab tunes the conversation, not that.
- Not a transcript reader
- The transcript is on the call record. Voice Lab is about how the voice performed, which is a question the words alone cannot answer.
Questions#
Will a change here affect calls already in progress?
No. Settings are read when a call's model session is opened, so a change applies to calls that start after it, not to one somebody is on.
Why does a finding sometimes have no setting attached?
Because some causes are not settings. A slow first token on a loaded host, or a model's own floor, is reported as what it is rather than pointed at a control that would not change it.
Can a customer workspace use Voice Lab?
Yes. It is one implementation serving both audiences, like every other capability — what differs between the Owner and a customer is the plan's allowance, not the feature.