An AI agent audit trail is the record of what an AI agent tried to do, which rule or person decided whether it could, and what happened next, kept in a form that someone outside the team can check later. It differs from an application log in purpose: a log helps engineers debug a service, while an audit trail has to convince an auditor, a regulator or a customer that a control was applied, and show whether the record changed after it was written.
This guide is for engineering, security and compliance leads who run agents that touch money, production systems, customer data or MCP tools. It covers what a useful record holds, why ordinary logs fail an audit, how tamper evidence and offline verification work, what the EU AI Act, SOC 2 and ISO/IEC 42001 ask for, a practical checklist and the questions to put to any vendor. The last section says what Praesidia records, and what it does not.
What is an AI agent audit trail?
An AI agent audit trail is an append-only sequence of decision records about the actions agents attempt through a control point. A useful record names the agent, the action, the target, the policy that applied, the decision, the result and the evidence that ties them together.
Three terms get used as if they meant the same thing. They do not, and the difference is what an auditor tests.
Application log, audit trail and tamper-evident evidence compared
| Application log | Audit trail | Tamper-evident evidence | |
|---|---|---|---|
| Written for | Engineers debugging a service | Reviewers asking who did what, and under which rule | Someone who does not have to trust the operator: an auditor, a regulator, a court |
| Typical content | Free-text messages, stack traces, request IDs | Structured records: actor, action, target, decision, time | The audit records plus the signatures and chain links that let them be checked |
| Who can change it | Anyone with write access to the log store, usually without a trace | Limited by access control, but an administrator can still edit the store | Anyone with access to the store, but an edit after signing makes the check fail |
| Shows what is missing | No | Rarely | Partly: a broken link in the middle of a chain shows; an action that was never written down does not |
| Checked with | Search in an observability tool | Queries in the operator's console | A verifier the reviewer runs on their own machine |
| Kept for | Often days to weeks, set by storage cost | Set by internal policy | Set by policy and by the law that applies |
The right-hand column is what this guide means by a trail that holds up. It does not replace the middle column; it is the middle column made checkable.
What an AI agent audit trail should record
A useful record answers seven questions about one action: who acted, what they tried to do, against what, under which policy, what was decided, what happened, and what lets someone check it later. If any of the seven is missing, the reviewer has to rebuild it from other systems, and that rebuild is where audits stall.
| Field: the question it answers | Illustrative scenario: a refund agent |
|---|---|
| Who: which agent acted, for which application or person? | Refund agent |
| What: what did it try to do, with which parameters? | Stripe refund, €8,250 |
| Target: which tool, system or data would it touch? | Payments tool (Stripe) |
| Policy: which rule applied, at which version? | Refunds over €500 need approval |
| Decision: allowed, denied or held, and by whom? | Held for approval |
| Result: what actually happened afterwards? | A person approves, then the refund executes |
| Proof: what lets someone check this later? | Decision receipt on the audit record |
Illustrative scenario: the agent, amount and rule are examples, not customer data. In Praesidia a hold applies once a policy is set to enforce; the default, observe mode, records the decision and lets the call through. Approving a held call is in every plan; multi-step approval workflows are Enterprise.
Details that decide whether a record holds up
- The policy version, not only its name. A rule edited next quarter must not change what last quarter's record means. Pin the version the decision used.
- The person in the loop. "Approved" is not evidence. The approver's identity and the time of the approval are.
- Denials and holds as well as allowed calls. A record of what a control stopped is the evidence that the control works. A trail of successes only shows that things happened.
- The delegation chain. When one agent calls another, the record should keep the path back to the application or person that started it. Without it, the last agent in the chain looks like the actor.
- Parameters, with secrets and personal data kept out. Keep enough of the request to show what was asked (an amount, a record count, an environment) and redact or reference the rest.
- Time and order. A timestamp in UTC at decision time, plus an order that does not depend on clocks alone, so a reviewer can tell which of two near-simultaneous actions came first.
- A link to the target's own record. A correlation ID that joins the decision to the payment, deployment or ticket on the other side lets a reviewer confirm the outcome from a second source.
Why ordinary application logs fail an audit
Application logs fail an audit for three reasons: whoever runs them can change them, they rarely hold the decision context an auditor asks about, and they give no signal of what is missing.
Whoever runs them can change them
A log line in a file, a database table or a search index can be edited by anyone with write access to that store, and the line itself carries no sign that it happened. Database audit features move the problem up one layer: the audit table has administrators too. Retention jobs and index rollovers remove entries on a schedule, which is ordinary housekeeping, but a reviewer looking at the log alone cannot tell housekeeping from a cover-up.
They lack the decision context
An application log says that a request to the refunds endpoint returned success. It rarely says which agent sent it, on whose behalf, which rule applied, at which version, and who approved it. Rebuilding that means joining the agent framework's logs, the gateway's logs, the tool server's logs and the approval tool's history, each with its own clock and identifiers. An auditor who asks for one decision and receives four log exports has to trust your join.
They give no completeness signal
A log with no entry for 14:02 could mean nothing happened, the logger was down, the pipeline sampled the event away, or the agent called the tool directly and never passed through anything that writes logs. Observability pipelines often sample or drop events under load by design. No log, signed or not, can show that an action took place outside its view. The honest answer is to state coverage: which agent paths pass through a control point that records decisions, and which do not.
Tamper evidence and offline verification
Tamper evidence does not stop anyone from changing a record; it makes a change visible to someone who checks. Offline verification means that check runs on the reviewer's own machine, against an exported bundle, without trusting or contacting the system that produced it.
The techniques are standard and well documented. What matters for an audit is what each one lets a checker show, and what it does not.
| Technique | What it lets a checker show | What it does not show | Primary source |
|---|---|---|---|
| Hash chain: each record carries the hash of the one before it | An edit, an insertion, or a removal from the middle of the chain | That the newest records were not cut off the end; that an unrecorded action took place | Haber and Stornetta, How to time-stamp a digital document (Journal of Cryptology, 1991) |
| Digital signature, for example Ed25519 | The record came from a holder of the signing key and was not changed after signing | That the signer wrote down everything, or that the key holder did not misuse it | RFC 8032 |
| Merkle tree over a batch of records | A record belongs to a batch with a given root, provable with a short proof; a later version of the log extends an earlier one without changing it | That the batch held each relevant event | RFC 9162 (Certificate Transparency 2.0) |
| Anchoring in a public transparency log, for example Sigstore Rekor | A root existed by the time it was logged, in a log the operator does not control, so a rewrite of the history up to that root no longer matches it | Records written after the last anchored root | Sigstore Rekor |
| Trusted timestamp | A hash existed at a time a timestamping authority vouches for | What the record means, or that anything else was recorded | RFC 3161 |
Two of those limits are easy to miss. First, a chain on its own cannot show that its newest records were cut off: the shortened chain is still a valid chain. Only a reference held outside the operator, such as an anchored root, covers that, and only up to the last anchor. Second, every technique in the table protects records that exist; none of them can show that an action happened which nobody recorded.
What an offline verifier can and cannot tell you
A good verifier gives a clear result and a reason, not a green tick. Expect it to tell you:
- whether the records and their batch roots carry valid signatures from keys that belong to the organization;
- whether the records form one unbroken chain;
- whether the roots match their external anchors, when anchoring was on;
- whether anchoring was on at all, which should be read from the result rather than assumed from a pass.
It cannot tell you that every action was captured, that records newer than the last anchor were not removed, or anything about systems outside the bundle it was given. A verification that passes is a statement about the records in the bundle, not about the actions outside it. A vendor that says otherwise is overselling.
What regulators and auditors ask for
Regulators and auditors ask for the same three things in different words: show the events, show that the record is reliable, and keep it for as long as the rules require.
| Framework | What it asks of the record | Read next |
|---|---|---|
| EU AI Act, high-risk systems | Automatic recording of events over the system's lifetime (Article 12); logs kept by providers (Article 19) and deployers (Article 26(6)) for an appropriate period of at least six months, unless other law provides otherwise | EU AI Act compliance for AI agents |
| SOC 2 | Evidence that access, monitoring and change controls operated across the audit period, from a population the auditor can sample (CC6, CC7, CC8) | SOC 2 controls for AI agents |
| ISO/IEC 42001 | A decision on when AI event logs are recorded, at a minimum while the system is in use (Annex A control A.6.2.8), with the records protected and retained as documented information | ISO/IEC 42001 for AI agents |
| GDPR | Personal data in logs minimised and kept no longer than necessary (Article 5(1)(c) and (e)) | GDPR and AI agents |
EU AI Act: Articles 12, 19 and 26
Article 12 of the EU AI Act requires high-risk AI systems to technically allow the automatic recording of events (logs) over the lifetime of the system. The logging has to enable the recording of events relevant to three purposes: identifying situations in which the system may present a risk or undergo a substantial modification, facilitating the provider's post-market monitoring under Article 72, and monitoring the system's operation by deployers under Article 26(5). For remote biometric identification systems, Article 12(3) sets minimum contents, including the period of each use, the reference database checked, the input data that led to a match and the people who verified the results.
Keeping the logs is a separate duty. Article 19 (providers) and Article 26(6) (deployers) require the logs a high-risk system generates automatically to be kept, to the extent they are under the provider's or deployer's control, for a period appropriate to the system's intended purpose, of at least six months, unless other Union or national law, in particular on the protection of personal data, provides otherwise. Article 26(5) adds that deployers monitor operation on the basis of the instructions for use and, where they have reason to think the system presents a risk, inform the provider or distributor and the relevant market surveillance authority, and suspend its use.
For agents, the practical reading is this. Most agents are not high-risk systems; whether yours is depends on what it is used for (Annex III lists the areas) or on the regulated product it is a safety component of (Annex I), not on the technology. Where an agent is, or is part of, a high-risk system, a decision trail like the one described here is the kind of event log Article 12 asks the system to support, and Article 26(6) makes keeping it the deployer's job. The obligations apply in phases; the dated timeline and the article-by-article mapping are on EU AI Act compliance for AI agents. This guide is not legal advice on your classification. Which events an agent's log should hold for Article 12, and who keeps it under Articles 19 and 26(6), is worked through in EU AI Act Article 12 logging for AI agents.
SOC 2
SOC 2 is an attestation, performed by an independent CPA firm, over controls you scope against the AICPA Trust Services Criteria. It does not mention AI agents. An agent audit trail is evidence for the common criteria on logical access (CC6), on system operations, which include monitoring for anomalies and evaluating security events (CC7), and on change management (CC8).
In practice the auditor asks for the population of relevant events for the audit period, picks samples from it, and tests that the report you produced is complete and accurate. A trail that can be exported for a chosen period and checked independently shortens that conversation. It does not replace the auditor's completeness testing, and a vendor's evidence supports your attestation without granting it. The criteria-by-criteria mapping is on SOC 2 controls for AI agents, and a worked build is in AI agent audit trail for SOC 2 evidence.
ISO/IEC 42001
ISO/IEC 42001:2023 is the certifiable management-system standard for AI. Its Annex A control A.6.2.8, AI system recording of event logs, asks the organization to decide at which phases of the AI system life cycle event logs are recorded, at a minimum while the system is in use. Clause 7.5 asks that documented information be controlled, which covers protecting it and deciding how long to keep it. A certification body samples the evidence that both are in place. The control mapping is on ISO/IEC 42001 for AI agents.
GDPR: the tension to plan for
Agent audit records often hold personal data: a customer ID in a tool call, an email address in a prompt. The data-minimisation and storage-limitation principles in GDPR Article 5 apply to them like any other processing. Plan for it at the design stage: record references and redacted summaries rather than raw payloads where you can, set retention deliberately, and know how an erasure request interacts with a signed chain before the first request arrives. More on GDPR and AI agents.
A practical checklist for an AI agent audit trail
Use this list to check an AI agent audit trail before an auditor does. Each item is something a reviewer can test, not a property to assert.
- Coverage is written down. List which agent paths pass through a control point that records decisions, and which do not: direct API calls, local tools, another team's agents. Put the list in your technical documentation.
- The seven fields are present for denials and holds as well as for allowed calls.
- Records carry the policy version that made the decision, not only the rule's name.
- Approvals name the person and the time.
- Integrity is checkable. Records are hash-chained and signed, you know whether signing can be switched off, and you know who is able to do it.
- Truncation is accounted for. External anchoring is on and you know how far the last anchor can trail the newest record, or it is off and you know that removal of the most recent records then cannot be detected.
- Export works for a chosen period and produces a bundle that a reviewer can check offline. The export itself should leave a record.
- A reviewer has run the verifier without your help at least once before the audit, and read the full result, including the anchoring line and any unsigned records.
- Retention meets the strictest rule that applies, from the EU AI Act minimum for high-risk systems to your SOC 2 audit period, and stays within what data-protection law allows.
- Personal data and secrets are kept out of the record or redacted, with references in their place.
- Monitoring reads the same records. Alerts and evidence come from one source, so what you monitor is what you can prove.
- Someone owns it. A named owner reviews coverage when a new agent, MCP server or tool goes live.
Questions to ask any vendor about its audit trail
These questions separate an audit trail that holds up from a log with a compliance label. Ask them of any vendor, including us; the answers that matter are specific and come with their limits.
- Which actions does your trail record, and which agent calls never pass through you?
- Is signing on by default? Who can switch it off, and how would I know it was off?
- What does a passing verification prove, and what does it not prove? Can I have that in writing?
- Can removal of the newest records be detected, and what has to be configured for that?
- Can my auditor verify an export offline, without an account on your platform and without contacting you? What do they need?
- Does a record carry the policy version and the approver's identity?
- When one agent calls another, does the record keep the chain back to the person or application that started it?
- Where do tool-call arguments and results go: into the signed record, into a separate log, or nowhere?
- How long are records kept, can I set it, and how do retention and erasure interact with the chain?
- Can I stream the same decisions into my SIEM?
- If your service is unavailable, are agent calls refused, allowed, or allowed and recorded later?
If the answer to question 3 names no limits, ask again: every technique in this guide has some, and a vendor that knows its trail can say what they are.
How Praesidia records agent decisions
Praesidia records the decisions it makes in an append-only, hash-chained audit trail, signed under the default configuration; an operator can turn signing off. A signed record shows it was not changed afterwards. It does not show that every action was captured.
What that means in practice:
- Decision receipts. A receipt reads one decision back: the agent, the authorization result, the policy version, and whether the record is signed and anchored. Receipts can be listed by AI system, agent, action or time range. See one on the AI agent evidence and proof page.
- Observe first, then enforce. New policies start in observe mode: the decision is recorded and the call goes through unchanged. Denials and holds apply once a policy is set to enforce. Approving a held call is in every plan; multi-step approval workflows are Enterprise. How decisions are made at runtime is on AI agent runtime security.
- Checked offline, without us. Organization owners and compliance officers can export a signed bundle for a chosen period. Your auditor checks it on their own machine with a verifier we provide to you and your auditor; it never calls Praesidia, and it is handed over directly rather than published as a download. What it checks, and what it does not, is set out on verify your audit trail offline.
- Anchoring is your choice. External anchoring is optional; without it, removal of the most recent records cannot be detected. With it turned on, the trail is anchored in the public Sigstore Rekor transparency log, which Praesidia does not control.
- Scope is stated. Calls that do not pass through Praesidia are not recorded. The trail is a record of the decisions Praesidia makes, not a transcript of the traffic itself, so write down which agent paths are routed through it.
- Into your SIEM. On the Advanced tier and Enterprise, decisions can be streamed to Splunk, Datadog or Microsoft Sentinel, so monitoring and evidence share one source.
Controls are mapped to the EU AI Act, SOC 2 and ISO/IEC 42001, not certified against them: the evidence supports your assessment and does not grant a certification. Evidence beyond the decision trail is described on AI audit evidence.
Common questions
What is an AI agent audit trail?
An AI agent audit trail is the record of what an AI agent tried to do, which rule or person decided whether it could, and what happened next, kept so that someone outside the team can check it later. It differs from an application log because it has to hold up for an auditor or regulator, not only help an engineer debug.
What is the difference between an audit log and an audit trail?
The terms are often used interchangeably. Where they are distinguished, an audit log is the stored list of events, and an audit trail is the connected sequence that lets a reviewer follow one action from request to decision to outcome. For AI agents the useful unit is the decision record: who acted, what they tried, the target, the policy, the decision, the result and the proof.
Does the EU AI Act require an audit trail for AI agents?
Article 12 requires high-risk AI systems to technically allow the automatic recording of events over their lifetime, and Articles 19 and 26(6) require providers and deployers to keep those logs, to the extent they control them, for an appropriate period of at least six months unless other law provides otherwise. An agent is covered when it is, or is part of, a high-risk system, and many internal agents are not. The article-by-article mapping and the dates are on EU AI Act compliance for AI agents.
Is a tamper-evident log the same as an immutable log?
No. A tamper-evident log does not stop someone with access from changing records; it makes an edit visible to anyone who checks the hashes and signatures. "Immutable" is a claim about storage that a reviewer usually cannot test, while tamper evidence is a property a reviewer can check for themselves.
Does a signed audit trail prove that every agent action was recorded?
No. A signature shows that a record was not changed after it was signed. It says nothing about actions that were never recorded. Completeness depends on coverage, meaning which agent paths pass through the control point that writes the trail, and a vendor should state that limit plainly.
How long should AI agent audit logs be kept?
As long as the strictest rule that applies to you. For high-risk systems under the EU AI Act that is a period appropriate to the intended purpose and at least six months, unless other law provides otherwise. For SOC 2, keep them long enough to cover the audit period and the fieldwork after it. For personal data in the logs, no longer than data-protection law allows. Write the retention period down.
Can an auditor verify an AI agent audit trail without access to the vendor's platform?
Only if the vendor exports records with their signatures and chain links and gives the auditor a verifier that runs offline. With Praesidia, organization owners and compliance officers export a signed bundle, and the auditor checks it on their own machine with a verifier we provide, without contacting Praesidia. A pass shows the records were not changed after signing; it does not show that every action was captured.
What does Praesidia record about AI agents?
Praesidia records the decisions it makes in an append-only, hash-chained audit trail, signed under the default configuration; an operator can turn signing off. External anchoring is optional; without it, removal of the most recent records cannot be detected. Calls that do not pass through Praesidia are not recorded.
Further reading
- Tamper-evident audit logs with cryptographic proofs: signing, chaining and anchoring in more technical depth.
- How to audit AI agent activity: a step-by-step procedure for reconstructing one agent action.
- Build an AI agent audit trail for SOC 2 evidence: the use case end to end.
- Tamper-evident log: the glossary definition.
- AI governance: the complete guide and MCP server governance: where the audit trail sits in a wider governance program.
Primary sources: Regulation (EU) 2024/1689, the EU AI Act · ISO/IEC 42001:2023 · RFC 3161 · RFC 8032 · RFC 9162 · Sigstore Rekor.