Nobody Audits Your AI Policy. They Audit the Evidence.
Managing AI risk has become an evidence problem. As agentic AI moves into production, three frameworks set the terms (the NIST AI Risk Management Framework, ISO/IEC 42001, and the EU AI Act) and for all their differences they ask one question: can you prove, on demand and continuously, that your AI is operating inside the boundaries you approved? A policy states intent; evidence states what happened, and evidence is what earns trust.
This seven-part AI risk management series walks the single governance loop that produces that proof, one stage at a time: see it, design it, prove it, own it, afford it, and defend it. This opening piece makes the case that reframes the rest: the standards grade evidence, not policy, which is why the annual audit no longer works for systems that change at runtime.
Jump to:
- The one question three frameworks share
- Why the annual snapshot stopped working
- The four shifts that close the governance gap
- Govern the decisions, not the models
- The three standards, one at a time
- Diving deeper into the three frameworks
- The kill switch across all three frameworks
- For the executives in the room: platform risk and AI capital
- The platform underneath the evidence must be trustworthy too
- Evidence is a runtime property
- More from our series on AI risk management
- The ultimate test
Your AI risk program is probably not broken. If you are reading this, you likely run a mature shop: documented controls, real audits, a risk committee that meets and means it. And for the first time in years, it feels inadequate against AI. That is not a competence problem. It is a physics problem. Everything you built assumes systems that hold still, and agentic AI does not.
I have spent 25 years building and transforming security programs inside regulated industries, which is a polite way of saying I have spent 25 years being audited. Here is the lesson that never changes: when someone asks you to prove a control worked, your policy is not the answer. The evidence is. A policy states intent. Evidence states what happened. Examiners were never interested in the first one.
The good news is that the fix is not a teardown. It is an extension of the program you already run, pointed at systems that move. And the reason to bother is not the audit. It is trust: whether your board, your regulators, your customers, and your own people can trust the AI you are running. Evidence is how you earn that trust.
The one question three frameworks share
That distinction is now the whole game in AI risk. Three frameworks set the terms: the NIST AI Risk Management Framework, ISO/IEC 42001, and the EU AI Act. They come from different places and carry different weight. One is voluntary guidance. One is a standard you can be certified against. The last is law. Read them side by side and they put the same question to whoever owns risk: can you produce, on demand, defensible and continuous evidence that your AI is operating inside the boundaries you approved?
Why the annual snapshot stopped working
For most of my career, you could answer that question once a year. Model risk meant a handful of documented models reviewed on a schedule. You assembled the evidence at audit time, and it was still true when the examiner showed up.
Agentic AI broke that. An agent is not a static model. It changes behavior with no code deploy: a prompt shifts, a tool gets added, a permission expands, a model is swapped underneath the workflow. The change that matters most is the one nobody files a ticket for: an agent picks up a new tool or MCP connector, and with it, whatever access the connecting credential already carries, in the time it takes to save a config. Case in point: The Nx Console compromise this past May began with exactly that kind of exposure. A GitHub CLI token sat in a developer’s config file where any local process could read it, and the malware that found it reached the GitHub API within 74 seconds.
None of that surfaces in a framework built for systems that hold still. The org chart makes it worse. A large enterprise with more than $1 billion in revenue already runs an average of eight different governance, risk, and compliance tools, a figure Gartner expects to reach ten by 2028 (Gartner, Market Guide for AI Governance Platforms, Q4 2025), and agentic AI falls into the seams between them, often owned by no single tool and no single person. A point-in-time attestation cannot close that gap, because a snapshot says nothing about how a system behaved on the days nobody was watching.
The four shifts that close the governance gap
Organizations do not fail at AI governance because governance is absent. They fail because governance moved at the speed of paper while AI moved at the speed of software. That gap is the whole of the problem, and it is fixable, because the cure is not more paper. It is four shifts.
The first is to stop treating governance as the brake. Done right, governance is what lets you move faster, because you can approve fast when you approve continuously. The metric that matters is not “are we compliant.” It is how quickly an approved AI system reaches production, safely, without a committee reconvening.
The second is to know that there is a path, and to place yourself on it. Governance maturity runs Policy, then Controls, then Automation, then Continuous Assurance, then Predictive. Most programs are stuck at Controls, a binder of documented intentions. Maturity is not which controls you have written down. It is how continuously those controls actually run.
The third is to be honest about what governance produces. Governance does not produce compliance. It produces evidence. Compliance is only one consumer of that evidence, and it is not even the most demanding one. Your board consumes it, your auditors consume it, your customers consume it in their vendor questionnaires, your regulators consume it, and increasingly your insurers consume it when they price your coverage. Build the evidence once and many buyers pay you back.
The fourth is the reframe that makes the rest work, and it deserves its own line:
Govern the decisions, not the models
Organizations do not suffer consequences from a model’s weights. They suffer from the decisions and actions the system takes. Every action an agent takes is authorized the instant it takes it, and you cannot pre-certify a model into only safe choices. So governance has to follow the decisions themselves: what was decided, what action was taken, what data was accessed, what changed. That is the shift from model risk management, which governs a static artifact, to agentic risk management, which governs a system that acts. Hold that thought, because it is the reason runtime, not the design review, is where the evidence actually lives.
The three standards, and the one question they share
NIST AI RMF: Voluntary United States guidance, organized into four functions: govern, map, measure, and manage. It is the vocabulary most enterprises adopt first.
ISO/IEC 42001: The first AI management system standard you can actually be certified against, which means an auditor examines your evidence directly.
EU AI Act: The only one of the three that is law, with penalties, and its high-risk obligations are phasing in (probably) 2027, on a timeline currently under amendment. For high-risk systems, it requires a risk management system, data governance, technical documentation, automatic logging, human oversight, accuracy and resilience, and post-market monitoring.
A voluntary framework, a certifiable standard, a binding law. All three strive for the same operational test: continuous, attributable evidence.
If the choice of framework itself feels overwhelming, that is the correct reaction, and the answer is that you do not choose one. No single framework covers you, because each is strong where the others are silent. The callout below is the shortest, honest map.
How the frameworks vary and how they fit together
| Framework | What it is | Its role in your stack |
|---|---|---|
| NIST AI RMF | Voluntary US guidance, outcome-based | The governance backbone and shared language |
| ISO/IEC 42001 | Certifiable management-system standard | The auditable operating model you can be certified against |
| EU AI Act | Binding law, risk-tiered | The legal floor for high-risk systems |
| OWASP (LLM and Agentic Top 10) | Technical, tactical risk and control catalog | The security controls your builders implement |
| MITRE ATLAS | Adversarial-threat knowledge base for AI | The threat model your red team tests against |
The three standards are your governance and legal spine. OWASP and MITRE ATLAS are the technical and threat layer you add beneath them. Choose by intent, expect a hybrid, and revisit the mix as your AI estate changes.
Diving deeper into the three frameworks
This reads in one direction. The standard is the requirement. The middle column is the evidence an examiner or auditor will actually ask to see. The last column is the governance capability that produces it. Notice there are no product names in that last column. That is deliberate: the capability is what the standard forces on you, and which platform you use to get it is a separate conversation, and a later one.
NIST AI RMF
| Function | Evidence an examiner expects | The capability that produces it |
|---|---|---|
| Govern | A named, accountable owner for each system, and the policy behind it | Accountability by design: every system and agent tied to a responsible identity |
| Map | A current inventory of the AI you run and what it touches | Visibility. You cannot govern what you cannot see, and you cannot defend what you cannot document |
| Measure | Continuous monitoring records, including how behavior changed | Runtime verification: live behavior compared against the approved design |
| Manage | Documented response and change-approval actions | The ability to intervene: stop the run, revoke access, preserve the record |
ISO/IEC 42001:2023
| Requirement | Evidence an examiner expects | The capability that produces it |
|---|---|---|
| Clause 6.1.2 and 6.1.3, AI risk assessment and treatment | Risk records mapped to specific controls | Visibility, plus a design you can assess |
| Clause 6.1.4, AI system impact assessment | An impact record tied to the actual system | A design record of record, not a slide deck |
| Annex A, life cycle and data controls | Versioned documentation of the system and its data | A living design record, versioned, with change control |
| Clause 8, operation | Operational monitoring evidence | Runtime verification against that design |
| Clause 9.2 and Clause 10, internal audit and improvement | Audit-ready records that were not curated after the fact | Non-selective logging. Regulators accept evidence, not assumptions |
EU AI Act (Regulation (EU) 2024/1689, high-risk obligations)
| Requirement | Evidence an examiner expects | The capability that produces it |
|---|---|---|
| Article 9, risk management system | A continuous, documented risk process across the lifecycle | Risk management that runs at runtime, not only at launch |
| Article 10, data and data governance | Data lineage and governance records | Data provenance tied to the system design |
| Articles 11 and 12, technical documentation and record-keeping | Automatic logs and event traceability over the system’s life | Immutable logging, every action traceable to a responsible identity |
| Article 14, human oversight | An oversight mechanism, including the ability to stop the system | A real stop, and identity-level control over who may intervene |
| Article 15, accuracy, robustness, and cybersecurity | Evidence of consistent performance and resilience to attack | Runtime resilience and control over the tools and connectors in use |
| Article 72, post-market monitoring | Continuous, documented monitoring after deployment | Monitoring as a standing capability, not a periodic project |
Read the last column top to bottom and the same four capabilities answer all three standards: see the system, document and approve its design, verify at runtime that it still matches, and attribute every action to a responsible owner. That is not marketing symmetry. It is what “managing risk” reduces to once the systems stop holding still.
READ MORE: Is your AI governance program set up for success? Take our AI Altitude Assessment to find out. >
Here is what it looks like as one control rather than a diagram. A leaked cloud credential shows up in an AI prompt, and the system redacts it before it ever reaches the model. That single runtime action answers NIST’s security-and-resilience measure (MEASURE 2.7), the EU AI Act’s robustness-and-cybersecurity obligation (Article 15), and ISO/IEC 42001’s operational-security controls (A.6.2.6, Clause 8), all at once, and it leaves a logged, attributable record while it does it. Detecting personal data in that same prompt answers a different set (NIST MEASURE 2.10, EU Article 10, ISO A.7.2 and Clause 6.1.4). Multiply that across the hundred-plus controls a real program runs and the crosswalk stops being a slide. It becomes the day to day.
The kill switch, in all three frameworks
Take the control everyone is asking about right now: the ability to stop an AI agent. It is not a fringe demand, and it is not only ours to make; all three frameworks call for it, in their own words. NIST expects mechanisms to supersede, disengage, or deactivate a system behaving outside its intended use (MANAGE 2.4). ISO/IEC 42001 indirectly supports this capability through operational controls, monitoring, corrective action, and treatment of nonconforming AI systems (Clauses 8 and 10). And the EU AI Act, the binding one, is the most explicit: it requires that a person be able to interrupt a high-risk system and bring it to a halt in a safe state (Article 14(4)(e)).
A kill switch, in other words, is table stakes across guidance, standard, and law. And like every control in this series, it counts only if you can show you are able to use it and prove that you did. One caveat from the response seat: stopping the agent is not the same as stopping what it already set in motion. By the time you reach the switch it may have issued a token or moved data, and halting the process does not pull any of that back even if the (now disabled) LLM brain can no longer reason against that data.
The kill switch, framework by framework
| Framework | What it demands: a stop or kill capability | Provision |
|---|---|---|
| NIST AI RMF | Mechanisms to supersede, disengage, or deactivate a system operating outside its intended use | MANAGE 2.4 |
| ISO/IEC 42001 | Operational control, and the duty to correct a nonconforming system | Clauses 8 and 10 |
| EU AI Act | A person can interrupt a high-risk system and bring it to a halt in a safe state | Article 14(4)(e) |
For the executives in the room: Platform risk and AI capital
Two translations for the people who answer to a board: First, you no longer have application risk. You have platform risk. When a handful of foundation models sit underneath dozens of business processes, one model outage, one provider price change, or one compromised model becomes an enterprise event, not a project issue. That is concentration risk, and boards already know how to think about it. The first job is to see where you are concentrated. The measure a board can act on is the count of critical processes that would stall if a single model or provider went dark, and once you have that number, a vendor’s rate-limit change or a model deprecation notice reads as a continuity event rather than a technical footnote.
Second, traditional governance manages capital, and AI governance is starting to manage a new kind: model capital, data capital, agent capital, and trust capital. Framed as allocation and protection of AI capital, the board grasps the mandate immediately, because allocation and protection is the language they already speak.
The platform underneath the evidence must be trustworthy too
There is a question that comes right after “Can you produce the evidence?,” and it is “Can I trust the thing producing it?” Fair question. The evidence layer is only as trustworthy as the platform it runs on. JetStream is FedRAMP Class D (High) Certified, built on the NIST 800-53 control baseline, which is the same lineage of controls that underpins the federal government’s most sensitive workloads.
A platform that is itself examined to a federal bar is a platform whose evidence you can rely on downstream. This is vendor trust. It is not, and we never present it as, a substitute for your own AI compliance.
Evidence is a runtime property
If you are still running governance as a documentation exercise, this is the uncomfortable part. You cannot approve a design that was never written down, so the design has to live as a current record, not a binder assembled for the audit. You cannot verify what you cannot compare, so runtime behavior has to be checked against that approved design continuously.
Then there’s AI whodunnit. You cannot attribute an action you cannot trace, so every invocation by a human or another agent has to resolve to a responsible identity. That last one is the hardest in practice: an agent acts under a service account or a delegated token, so by default the record shows a machine principal, not the person who set it in motion. Automating chained attribution back to an accountable human is not trivial, but it is vital. So you have to account for how this evidence will be collected and stored.
None of that is compliance. Compliance is a determination the risk function makes and attests to. Evidence is what sits underneath the attestation. Keep the two separate, and be skeptical of anyone who tells you a tool makes you compliant. It does not. It either produces the evidence you will be asked for, or it does not.
More from our series on AI risk management
This piece is the anchor. The six installments that follow walk one governance loop, one stage at a time: see it, design it, prove it, own it, afford it, and defend it.
Each stage closes one gap where the proof breaks, and each speaks to a different seat at your AI council.
- Your security leader, who cannot see the agents already in production.
- Your builders, who fear governance will become the ticket queue that slows every release.
- Your executive sponsor, fielding board questions the org cannot yet answer.
- Your privacy and legal leader, who cannot sign off on data handling she was never shown.
- And your data and AI leader, whose policy lives in a document that no running agent has ever honored.
The series takes each of their views in turn. The thread that connects them is the one this piece has been making: you cannot manage what you cannot prove, and the point of the proof is trust.
The ultimate test
Here is the one I would apply to my own program. If an examiner walked in this afternoon and asked me to show, on demand, that every AI system in production has operated inside its approved boundaries since the day it went live, could I? Not reconstruct it next week with three people on a bridge. Show it today.
If the honest answer is no, you do not have a compliance gap yet. You have an evidence gap, and every framework in the building is about to grade you on it. The standards have already moved. They stopped accepting the promise that a control exists. They want you to show that it operated.
Next installment: you cannot prove any of this for a system you cannot see. We’ll walk through the AI you do not know you are running and how to find it.