AI audit trail & evidence
Architecture directionIf AI acted, the business should be able to reconstruct what happened.
AI audit is not a demand to expose private model reasoning. It is a requirement to preserve the operational evidence around material work: who initiated it, which agent acted, what context was used, which tools were called, what policy decided, who approved and what action was actually released.
That is the evidence model behind the TEMRIK AI control plane: organisational authority should remain reconstructable even when the underlying model changes.
Identity · context · models · tools · policy · approval · action · outcome
Operational evidence record
Actor
User, agent, service identity and tenant.
Goal
The authorised task or business objective.
Context
Sources retrieved, records referenced and evidence identifiers.
Model
Provider, model identifier and relevant runtime metadata.
Tools
Tool requested, scoped arguments, result status and duration.
Policy
Allow, deny, restrict, escalate or request-more-information result.
Delegation
Which agent or service handed work to another actor.
Approval
Who approved a consequential release, and when.
Action
What was actually sent, changed, created, released or executed.
Outcome
Success, failure, exception, reversal or later business result.
Observability does not require private model reasoning.
It requires operational evidence.
The audit question
You do not need to know every thought.You need to know every material action.
Traditional application logs often tell engineers that something failed. Enterprise AI evidence needs a wider question: can the organisation prove the path from intent to action?
Reconstructable action
Interactive explorer
Inspect the operational evidence behind each event.
Operational evidence is inspectable. Private model reasoning is not required. Select an event to see what happened, who acted, the policy result and what was recorded.
Selected control
Goal received
The workflow receives a defined business objective.
What happened
A Commercial Manager asks the variation workflow to assess Rev D.
Who / what acted
Human / workflow initiator
Source / input
The workflow receives a defined business objective.
Policy result
Accepted into approved workflow
What was recorded
Goal + actor + time
Output
Workflow created
A useful audit trail follows the work, not just the model call.
The evidence boundary should continue across retrieval, reasoning interfaces, tools, policy, human authority and the final business system. The model response is one event inside a larger operational run.
Event 01
Goal
A bounded task enters the system with actor and tenant context.
Event 02
Context retrieved
Permission-aware sources are selected and referenced.
Event 03
Model called
The chosen provider/model processes the defined task.
Event 04
Tool requested
The agent requests a capability with structured arguments.
Event 05
Policy result
A control layer allows, restricts, blocks or escalates.
Event 06
Human approval
A named authority approves when the action threshold requires it.
Event 07
Action
The permitted business action is released to the target system.
Event 08
Outcome
Result, error, exception and relevant evidence are correlated.
Material-action timeline
01
GOAL
02
CONTEXT RETRIEVED
03
MODEL CALLED
04
TOOL REQUESTED
05
POLICY RESULT
06
HUMAN APPROVAL
07
ACTION
08
OUTCOME
On mobile this sequence remains a vertical narrative: each event stands alone and the evidence chain still reads in execution order without relying on animation or horizontal scrolling.
The vocabulary matters
Logging, tracing, audit, evidence, monitoring and evaluation are related.They are not the same thing.
Treating every telemetry stream as an audit record produces either too much data or too little proof. Each layer serves a different operating purpose.
Logging
Discrete records of events or state, usually optimized for search and diagnosis.
Tracing
A correlated sequence of spans showing how one request moved across models, tools, agents and services.
Monitoring
Ongoing visibility into health, latency, errors, volume and operational conditions.
Audit
A business-control record designed to demonstrate who did what, under which authority, and with what result.
Evidence
The records, references and approvals needed to support a material claim or action.
Evaluation
A structured assessment of quality, safety, policy compliance or task performance.
What to capture
Record enough to explain the business event.Not everything because storage is cheap.
A defensible evidence design begins with purpose. Decide which events matter, how long they matter, which identifiers link them, who can inspect them and which fields should be minimized or redacted.
Actor
User, agent, service identity and tenant.
Goal
The authorised task or business objective.
Context
Sources retrieved, records referenced and evidence identifiers.
Model
Provider, model identifier and relevant runtime metadata.
Tools
Tool requested, scoped arguments, result status and duration.
Policy
Allow, deny, restrict, escalate or request-more-information result.
Delegation
Which agent or service handed work to another actor.
Approval
Who approved a consequential release, and when.
Action
What was actually sent, changed, created, released or executed.
Outcome
Success, failure, exception, reversal or later business result.
Privacy and telemetry
What should never be logged carelessly?
Trace systems can become one of the richest datasets in the organisation. That makes collection policy, redaction, retention, access control and deletion part of the architecture—not an afterthought.
Passwords and secrets
Credentials, API keys, tokens and private keys should not become ordinary telemetry.
Unnecessary sensitive data
Collect only what is proportionate to the operational purpose and retention need.
Private chain-of-thought
Audit should capture material actions and evidence, not demand hidden model reasoning.
Unrestricted raw payloads
Full prompts, files, tool inputs and outputs can create a second sensitive-data store.
Security material without controls
Trace stores need their own identity, access, encryption, retention and deletion rules.
Private reasoning is not the audit target.
For enterprise control, the more reliable audit question is whether the organisation can reconstruct the observable operations: inputs selected, external evidence referenced, tools requested, policy decisions, approvals, actions and outcomes. This avoids making hidden chain-of-thought a governance dependency.
Observability architecture
Agent observability should cross provider boundaries.
OpenTelemetry provides a widely used foundation for correlated traces, metrics and logs, while major agent platforms are adding their own observability layers. The enterprise design problem is to preserve a useful business record even when individual providers expose different telemetry.
Execution telemetry
Trace IDs, spans, latency, status, model and tool operations.
Business evidence
Source references, policy outcomes, approvals, released actions and business results.
Control evidence
Identity, tenant, permission, delegation and exception decisions made outside the model.
TEMRIK architecture direction
01
Correlation
Carry durable run, actor, tenant and workflow identifiers.
02
Policy evidence
Record the allow / deny / escalate decision independently of the prompt.
03
Authority
Bind human approvals to the material action they release.
04
Outcome
Correlate the final system response or business result back to the run.
Human authority
Approval should be evidence, not just a button click.
When a workflow crosses an action ceiling, the audit record should be able to identify what the person saw, which action was proposed, who had authority, when they approved it and what was ultimately released.
See how this fits into TEMRIK Human Control and the wider agentic AI architecture.
Proposed action
What the agent wants the business to release.
Evidence bundle
The sources and context relevant to the decision.
Policy state
Why the request requires approval rather than automatic execution.
Decision owner
The named role or person with authority over the consequence.
Decision
Approve, reject, amend, request more information or escalate.
Released action
The exact action that reached the external system.
Security boundary
Audit data is security-sensitive data.
Observability can reveal user inputs, retrieved records, tool arguments, operational metadata and system structure. The evidence platform therefore needs its own access model and should follow the same least-privilege discipline as the workflow it observes.
The related AI security architecture explains why data boundaries, model choice, permissions and actions should be controlled outside the model itself.
Access
Limit evidence access by tenant, role and operational need.
Retention
Keep evidence for a defined purpose and period rather than indefinitely by default.
Integrity
Protect high-value audit events from silent alteration or deletion.
Redaction
Remove or mask sensitive fields that are unnecessary for the evidence purpose.
Separation
Do not make the same compromised agent the sole authority over its own audit record.
Export
Design for defensible review without uncontrolled bulk extraction of sensitive telemetry.
Research basis
Primary sources behind this architecture.
These sources describe current tracing, agent observability, telemetry conventions, AI risk-management and runtime-control patterns. They are references, not endorsements or partnerships.
OpenAI
Agents SDK — integrations and observability
Structured traces can record model calls, tool calls, handoffs, guardrails and custom spans.
OpenTelemetry
Semantic conventions
Common naming conventions for traces, metrics, logs, resources and other telemetry.
NIST
AI Risk Management Framework 1.0
Transparency, accountability, monitoring, evaluation and documentation are recurring AI risk-management themes.
OWASP
Agent Control Standard
Agent systems should be inspectable, traceable, instrumentable and subject to runtime controls.
Microsoft Foundry
Agent tracing overview
OpenTelemetry-based agent tracing across model, tool, memory and workflow operations.
AWS
AgentCore observability
Traces, spans, request context, tool invocations, errors and resource-use telemetry for agent applications.
Google Cloud
Gemini Enterprise Agent Platform release notes
Google documents agent observability capabilities for deployed agents and MCP servers.
Anthropic
Claude on Amazon Bedrock — activity logging
Official documentation describes invocation logging for prompts and completions in Bedrock-hosted Claude usage.
How it fits together
Auditability is the memory of controlled AI.
TEMRIK's operating approach is to place identity, company context, policy, playbooks, tools and human authority around model capability. A reconstructable evidence trail is how those controls become reviewable after the event.
Control loop
01
Identity establishes the actor.
02
Policy establishes the boundary.
03
Tracing establishes the execution path.
04
Evidence establishes what supported the action.
05
Human approval establishes consequential authority.
06
Audit establishes what happened afterwards.
Start with one workflow
Map the evidence trail for one AI workflow.
Identify the actor, data, model, tool, policy, approval point, released action and outcome that would need to be reconstructed if the workflow were challenged six months later.
Agents need rules as well as traces. Get the free 27 Rules of Peace field guide, or return to the TEMRIK platform.