Generated, not gathered
The trail is written by the system as it runs. Nobody assembles it, nobody remembers to turn it on, and nobody can choose not to.
Fails when: a script exports logs the week before an audit.
No agent reaches production touching regulated work until six cells are solid. Deliberately shorter than the canvas, because a gate people can hold in their head is a gate that gets used. Run it per deployment, not per program.
Every item maps to something an auditor already asks for: the human oversight duty in the EU AI Act, the management system clauses in ISO 42001, and the govern and manage functions in the NIST AI risk framework, plus CMMC for anyone in the defense base.
Map the Evidence cell once and reuse it. These regimes were written separately and converge on the same handful of demands: a person who can intervene, a record of why the system did what it did, and a management system that survives an auditor rather than a demo.
| Regime | What it asks for | Cells it lands on | Artifact you must show |
|---|---|---|---|
| EU AI ActArticle 14 and Annex III | A person who can understand the system, intervene in it, and stop it. Employment uses are classed high-risk. | Authority, Mastery, Evidence | Named oversight role, intervention log, halt mechanism that is actually reachable in the runtime |
| ISO/IEC 42001AI management system, 2023 | A certifiable management system: policy, roles, risk treatment, and continual improvement around AI. | Evidence, Purpose, Learning | Documented management system, competence records, internal audit trail |
| NIST AI RMFVersion 1.0, 2023 | Practice organized around govern, map, measure, and manage. Voluntary, and widely used as the common vocabulary. | All twelve, weighted to Evidence | Risk mapping per use case, measurement plan, documented management decisions |
| CMMCFor the US defense base | Assessed cybersecurity practices, with certification levels and third-party assessment. | Evidence, Authority, Resources | Control implementation evidence, scoped boundary, assessed rather than self-declared |
Dates have moved and will move again. High-risk obligations under the EU AI Act have already been deferred once, so confirm current timing with counsel rather than with this table. The entry carries the detail and the caveat.
Evidence is the cell organizations most often score solid and are most often wrong about. These four practices separate a system that produces evidence from a team that can reconstruct it under pressure.
The trail is written by the system as it runs. Nobody assembles it, nobody remembers to turn it on, and nobody can choose not to.
Fails when: a script exports logs the week before an audit.
What the agent had in front of it, what it weighed, and what it rejected. An outcome log tells you what happened and nothing about whether it should have.
Fails when: the record is the decision without the inputs that produced it.
Which data, which version, which model, which prompt or policy was in force. Most disputes are about the inputs rather than the logic.
Fails when: the model was silently updated and nothing recorded the change.
Not a role in a policy document. A person, assessed for this workflow, rostered while the agent runs. This is what stops blame landing on whoever was nearest.
Fails when: the accountable party is a team, a committee, or a job title.

An assessment that names who may overrule the machine, and proves it.
See the ladder