AGENSPHERE/ JOURNAL
← JOURNAL
ENTERPRISE AI · PART 8 OF 9
J-032NOTE3 MIN READ

No audit trail: when nobody can explain what the AI did

A customer disputes a decision, a regulator asks how it was made, or an agent did something unexpected overnight. Without an AI audit trail showing which model, prompt, context and tools produced the outcome, you cannot defend it, fix it or learn from it.

IN SHORT
  • An AI audit trail is a record, for each AI-assisted decision or action, of what went in, what the system did, and why: model and version, prompt version, context and sources used, tool calls, outputs, and any human approval.
  • Without it, companies cannot answer disputes, regulator questions or incident reviews, and cannot reproduce a failure to fix it.
  • Chat logs are not enough. A useful trail is structured, links every step of a run, records versions and sources, and is retained under a clear policy.
  • Design the trail with privacy in mind: log references and metadata by default, and keep sensitive raw content only where needed, with restricted access.

This is part 8 of Enterprise AI. The previous parts kept coming back to one requirement: being able to show what happened. This part makes it concrete.

01When does the question get asked?

Usually at the worst moment:

  • A customer says the assistant promised a refund. Did it? Based on what?
  • A loan, claim or job application was declined with AI involved. The applicant, or a regulator, asks how the decision was made.
  • An agent changed records overnight. Which ones, why, and on whose authority?
  • Quality dropped last week. Was it a model update, a prompt change, or new data?

If the honest answer is “we don't know, the model must have decided”, the company has a governance problem, a legal exposure and a debugging problem at the same time.

02Why are chat logs not enough?

Many systems keep a transcript of user messages and model replies. It shows what was said, not why:

  • It does not record which model and version produced the answer.
  • It does not record which prompt version was live.
  • It does not record which documents, records and rules were in the context (see out of context).
  • It does not link tool calls and their results, or show which actions actually happened in other systems.
  • It does not record who approved what (see human approval is a durable wait).

Without those, you cannot reproduce the outcome, so you cannot explain it.

03What should an audit trail record?

One record per run, with linked steps, written as the run happens.

audit/run-7f3a.json
{
  "run_id": "run_7f3a", "task": "claims.triage", "actor": { "type": "user", "id": "u_1042" },
  "started_at": "2026-10-06T09:14:03Z",
  "steps": [
    { "n": 1, "kind": "model_call", "model": "provider_b/mid-tier@2026-08", "prompt": "claims_triage.v12",
      "context": ["policy:motor-v31#4.2", "claim:CLM-88231", "rule:fasttrack-limit-v5"],
      "output_ref": "blob://runs/7f3a/1", "tokens": { "in": 3120, "out": 214 } },
    { "n": 2, "kind": "tool_call", "tool": "set_claim_route", "args_ref": "blob://runs/7f3a/2a",
      "effect": "idempotent", "key": "route:CLM-88231", "result": "ok" },
    { "n": 3, "kind": "approval", "required_by": "rule:amount>200000", "decided_by": "u_0311",
      "decision": "approved", "at": "2026-10-06T11:02:47Z" }
  ],
  "outcome": "routed_fasttrack", "retention": "7y", "content_access": "audit-restricted"
}

The essentials:

  1. Identity: who or what initiated the run, and on whose behalf it acted.
  2. Versions: model and version, prompt version, tool definitions, policy and rule versions.
  3. Context with provenance: references to every document, record and rule used, by ID and version.
  4. Steps in order: each model call, tool call and result, linked to the run.
  5. Actions and effects: what changed in which system, with idempotency keys (see effect classes).
  6. Human decisions: approvals, edits and rejections, with who and when.
  7. Outcome: what the run concluded.

04How do you keep it from becoming a privacy problem?

An audit trail full of raw prompts is itself sensitive data (see data compliance). Design it deliberately:

  • Log references, not copies, by default: document IDs and versions instead of document text.
  • Store raw content separately, encrypted, with restricted access and its own retention period, only where the use case needs it.
  • Apply retention by task, set with your legal team: some records must be kept for years, most raw content should not be.
  • Make access itself audited, so reading the audit trail leaves a trace.

05Can you reproduce a run from the trail?

That is the real test. If the trail records the inputs, versions and tool results, you can replay the run offline to see exactly what happened, and re-run it on a fixed prompt or a new model to check whether a change would have helped. Durable execution runtimes like KEEL make this a property of the system: the event log that lets a run survive a crash is also its audit trail.

If you cannot replay it, you cannot explain it. If you cannot explain it, you should not automate it.

06Where does this fit?

The trail lives in L0 infrastructure, next to tracing and cost (see trace every model call), and records everything that happens at L4 actions. It is the evidence layer of the owned intelligence layer, described next.

Previous: Pilot purgatory. Next: The owned intelligence layer.

Building something like this?DESCRIBE A SYSTEM →