Auditing what an answer was based on
Every reply keeps the record of what it was grounded in: the score, the titles matched, which of them carried authority, any commitment topics detected and the citations for the passages used. Auditing means reading those, resolving each citation back to a source and section, and comparing the cited text with the sentence in question. Nothing is reconstructed afterwards.
The stages#
- Trigger — somebody questions a sentence: a customer quotes it back, or a review notices it.
- User event — the thread, or the call, is opened in the workspace that sent it.
- Authentication and workspace resolution — the audit reads only that workspace's records; there is no cross-workspace view of an answer.
- Ingest — the reply and the message it answered are located together, so the sentence is read against the question rather than alone.
- Canonical record — the thread or the call is the anchor, and the entry written at the time of the Knowledge check is the audit record.
- Reasoning — that entry is read: the score of the best match, the titles retrieved, which of them carried fact authority, which commitment topics fired, the authority verdict, and the citations.
- Knowledge, memory and rules — each citation resolves to a source and a passage position, and the memories that were marked used at that moment are visible in the memory viewer for the person involved.
- Autonomy and approval — the decision record says whether the reply went automatically under the channel's autonomy, or was approved, and by whom.
- Action and provider — the send evidence sits with the reply: the provider's own acknowledgement for an email, the transcript for a call.
- Result — the sentence is compared with the cited passage. This is the audit; everything before it was gathering.
- Relationship and timeline — the exchange sits on the person's timeline with everything else that has happened with them, on any channel.
- Audit and usage — none of this is generated on request. It was written as the reply was produced, which is what makes it evidence rather than a reconstruction.
Reading the record#
| Field | Question it answers |
|---|---|
| Score | How well anything matched at all — a near miss looks very different from a strong match |
| Matched titles | What was retrieved and considered |
| Authoritative titles | Which of those could prove something. Empty beside a full matched list is a complete diagnosis on its own |
| Critical topics | Whether this was treated as a commitment needing explicit authority |
| Authority verdict | The short reason the gate passed or failed |
| Citations | The exact passages, by source and position |
An audit that stops at *a citation exists* has not audited anything. The question is always whether the cited text supports the sentence, and the three verdicts are: it does; it is about the topic but does not say that; or the sentence went further than the passage. The second and third are different faults with different repairs — the second is usually authority set too generously on a mixed document, the third is a case for tightening the source so there is less room to extrapolate.
What can be traced, and what cannot#
- Traceable
- Which passages were used, from which source, at which position, under which authority; whether the material was judged sufficient; who approved the send, if anyone; and the provider's acknowledgement.
- Traceable with care
- Memory that informed the reply. Memories can be edited or superseded afterwards, so what is on the record today is not necessarily what was in front of the model then — the viewer shows provenance and supersession for exactly this reason.
- Not traceable
- The model's reasoning. There is no stored chain of thought, and inventing one after the fact would be worse than admitting its absence.
- Fragile
- A passage position after a re-index. The source identifier in a citation is stable; the position within it can move when the text is rebuilt. A deleted source leaves a citation that resolves to nothing, which is itself informative.
Auditing a call rather than an email#
The same records exist, with the transcript standing in for the sent message. Two differences are worth knowing before starting. The call brief is capped at four Knowledge entries and 2,000 characters, so the material in front of the model was a subset of what retrieval would have returned for an email; and the assembled instruction size is recorded on the call, which is often the fastest explanation for a thin answer.
Where a caller was told something no source supports, the sequence to check is the same: what was retrieved, what carried authority, and whether the sentence goes beyond the passage. A call adds one more possibility — that the sentence was never grounded at all but was conversational filler, which the transcript shows and the citations do not.
Questions#
How long are these records kept?
They live with the thread, the call and the decision log rather than in a separate store with its own lifetime, so they last as long as the conversation does. Deleting a Knowledge source does not delete the citations that named it — it makes them stop resolving, which is the correct outcome for an audit.
Can I audit a reply that was never sent?
Yes, and it is often more useful. A held draft carries the same Knowledge check as a sent one, so reviewing what a draft was grounded in before approving it is the cheapest possible audit — done before the customer is involved rather than after.
The citation points at the right document but the wrong section.
That normally means the source was re-indexed after the reply went out and positions moved, or that extraction produced a structure different from the document a person is reading. Compare the passage text rather than the heading; the text is what the model saw.