← Blog

AI Governance and Risk Management: How the Two Programs Fit Together

Parminder Singh
Parminder Singh··12 min read
Summarize with AI

AI governance sets the policies, roles, and accountability for AI use. Risk management identifies, measures, and treats the AI-specific risks the governance framework recognizes. The two programs share inputs (data classification, use case inventory, vendor list) and produce different outputs (policies versus risk treatments). Four frameworks draw the line differently: NIST AI RMF splits GOVERN from MAP/MEASURE/MANAGE, ISO 42001 hands the risk process to ISO 23894, the EU AI Act separates Articles 17 and 26 from Article 9, and SR 26-2 (which superseded SR 11-7 on 17 April 2026) separates board oversight from validation while placing generative and agentic AI outside its scope. This piece covers all four, plus the per-request evidence both programs need to demonstrate operation.

Compliance & Regulationai-governanceai-risk-managementnist-ai-rmfiso-42001complianceaudit
AI Governance and Risk Management: How the Two Programs Fit Together

TL;DR

  • Run AI governance and risk management as two linked programs, not one blended control set: governance approves use and assigns accountability, while risk management records, treats, monitors, and accepts residual risk.
  • Build one per-request evidence layer recording the user, time, purpose, data classification, policy, and outcome, then use it to prove governance operation and measure control effectiveness.
  • Map each use case to the applicable framework’s boundary, version, scope, and obligations before choosing controls, especially for generative or agentic systems outside SR 26-2.
  • Keep policies, role assignments, sanctioned tools, risk records, control owners, metrics, remediation, and review decisions traceable to the same operational records.

AI governance sets the policies, the roles, and the accountability for how the organization uses AI. AI risk management identifies, measures, and treats the AI-specific risks the governance framework recognizes. The two programs share inputs and produce different outputs. The governance side produces policy documents, sanctioned tool lists, role assignments, and review cadences. The risk side produces a risk register, control assignments, monitoring metrics, and remediation plans.

Under NIST AI RMF, ISO 42001, EU AI Act Article 9, and US bank model risk guidance, the two programs sit alongside each other and share the operational infrastructure. The per-request audit evidence both programs depend on is the same data set. The architecture that produces it serves both purposes.

Four frameworks, and each one draws the line between the two programs in a different place:

| Framework | Governance side | Risk side | Status as of August 2026 | |---|---|---|---| | NIST AI RMF 1.0 | GOVERN | MAP, MEASURE, MANAGE | Published January 2023, currently under revision. Generative AI Profile (AI 600-1) added 26 July 2024 | | ISO/IEC 42001:2023 | Clauses 5 and 9, Annex A policy controls | Clause 6.1, elaborated by ISO/IEC 23894:2023 | Certifiable. Auditor accreditation set by ISO/IEC 42006:2025 | | EU AI Act | Articles 17 and 26 | Article 9 risk management system | High-risk obligations deferred to 2 December 2027 by Regulation (EU) 2026/1744 | | SR 26-2 | Board and senior management oversight | Validation, ongoing monitoring, effective challenge | Issued 17 April 2026, superseding SR 11-7. Generative and agentic AI explicitly out of scope |

I want to walk through how the programs fit under each of those, the shared infrastructure, and what each program needs to demonstrate operation under audit.

Governance and risk as separate functions

The governance program answers: who is accountable for AI use in the organization, what policies define acceptable use, which tools are sanctioned, what review cadence applies. The output is a set of documents and role assignments that the organization commits to. The board, the executive AI committee, the security organization, the compliance function, and the line-of-business owners each have defined responsibilities.

The risk management program answers: what could go wrong with AI use in our organization, how likely is each failure, what is the impact, what controls reduce the risk, what residual risk are we accepting. The output is a risk register, the control assignments that mitigate identified risks, the monitoring metrics that track control effectiveness, and the residual-risk acceptance decisions.

The two functions overlap in personnel (the same security and compliance officers often staff both) but produce different artifacts.

How NIST AI RMF positions them

The NIST AI RMF, published as version 1.0 in January 2023, has four functions: GOVERN, MAP, MEASURE, MANAGE. GOVERN aligns with the AI governance program. MAP, MEASURE, and MANAGE align with the risk management program. NIST states on the framework's own page that AI RMF 1.0 is being revised, with no completion date announced, so a program built on it should record which version it was built against.

GOVERN establishes the structures and processes for AI use. The output includes the AI policy, the role and responsibility assignments, the accountability structures, and the supplier management approach.

MAP identifies the context of use, the data classification, the stakeholder concerns, the third-party dependencies, and the impact of AI on individuals and groups. The output feeds the risk register.

MEASURE quantifies the risks and the performance of the AI systems. The output includes the metrics, the testing results, and the monitoring data.

MANAGE prioritizes and treats the identified risks, allocates resources, and adjusts the program based on operational results.

The two programs share data (the use case inventory feeds both, the vendor list feeds both, the data classification feeds both) and have different outputs (policy documents from GOVERN, risk register and treatments from MAP/MEASURE/MANAGE). The GOVERN function in detail is where most of the accountability language lives.

The framework's own text is general enough to cover any AI system, which is why the Generative AI Profile matters more than the base document for anything built on an LLM. NIST AI 600-1, released 26 July 2024, names 12 risks that generative AI creates or worsens (confabulation, data privacy, information integrity, harmful bias, CBRN information, dangerous or violent recommendations, human-AI configuration, information security, intellectual property, obscene content, value chain and component integration, environmental impacts) and lists over 200 suggested actions against them, each tagged to a GOVERN, MAP, MEASURE or MANAGE subcategory. A risk register that says "AI risk: model produces incorrect output" is doing less work than a register whose rows carry the AI 600-1 identifiers, because the second one tells the auditor which suggested actions you considered and declined.

NIST added a second 2026 signal worth tracking: on 7 April 2026 it released a concept note for an AI RMF Profile on Trustworthy AI in Critical Infrastructure. If you operate in one of those sectors, the profile is the document your regulator will read first once it lands.

How ISO 42001 positions them

ISO 42001 specifies an AI management system. The management system includes the policy framework (governance), the risk management process (risk), the operational controls (the implementation layer both depend on), and the continual improvement cycle.

The split most implementation teams miss is that ISO wrote the risk half as a separate document. ISO/IEC 42001:2023 requires an AI risk assessment and treatment process in Clause 6.1 but says little about how to run one. ISO/IEC 23894:2023 is that how. It mirrors the clause structure of ISO 31000:2018, the general enterprise risk management guideline, and adds AI-specific guidance wherever ISO 31000's text needs extending. So the stack reads: ISO 31000 gives the risk process any organization uses, ISO 23894 adapts it for AI, and ISO 42001 is the certifiable management system that the adapted process runs inside. ISO 23894 carries no certification of its own, which is why it gets skipped, and skipping it is how organizations end up with an AI risk register that looks nothing like their enterprise one and cannot roll up into it.

The certification audit tests both the policy framework and the risk management process. The auditor reads the policies, reviews the risk register, samples the Annex A controls, and verifies the management system is operating as specified.

The shared infrastructure between the governance and risk arms is what the auditor relies on for the operational evidence. The policies describe what the organization commits to do. The risk treatments describe how the controls reduce identified risks. The evidence layer demonstrates that both are operating.

How EU AI Act Article 9 positions them

Article 9 of the EU AI Act requires a risk management system for high-risk AI systems, and it is prescriptive in a way the voluntary frameworks are not. Article 9(1) and 9(2) require a "continuous iterative process planned and run throughout the entire lifecycle" with regular systematic review, comprising identification and analysis of known and reasonably foreseeable risks, estimation and evaluation of risks under normal use and under reasonably foreseeable misuse, evaluation of emerging risks from the Article 72 post-market monitoring data, and adoption of appropriate and targeted risk measures. Article 9(6) to 9(8) require testing against prior-defined metrics and probabilistic thresholds, both during development and before the system goes on the market. Article 9(9) requires providers to consider adverse impacts on minors and other vulnerable groups.

The governance half sits elsewhere in the regulation: Article 17 for the provider's quality management system, Article 26 for deployer obligations. That separation is deliberate and it is the cleanest statutory expression of the split this article is about.

The date changed this summer, and a lot of programs are working from the old one. Regulation (EU) 2026/1744, the digital omnibus on AI, was published in the Official Journal on 24 July 2026 and entered into force on 27 July 2026, six days before the original 2 August 2026 deadline. It defers the high-risk obligations for standalone Annex III systems to 2 December 2027, and for AI embedded in products already covered by EU product safety law under Annex I to 2 August 2028. The Commission's stated reasons were the absence of harmonised standards and the incomplete designation of national competent authorities and conformity assessment bodies. The obligations were deferred rather than dropped, and Article 9's substance is unchanged. Whether any of it applies to you turns on the high-risk classification question.

How US bank model risk guidance positions them

SR 11-7 is the document most AI governance material still cites here, and it was superseded on 17 April 2026. SR 26-2, issued jointly by the Federal Reserve, the OCC and the FDIC, replaces both SR 11-7 (April 2011) and SR 21-8 (April 2021) on model risk management for BSA/AML systems. It is expected to be most relevant to banking organizations with over $30 billion in total assets, though the Fed notes it may still apply below that threshold where model risk exposure is significant.

The structure survives. Governance means board and senior management oversight, the model inventory, the policy framework, and periodic reporting. Risk means validation of conceptual soundness, ongoing monitoring, outcomes analysis, and "effective challenge," which SR 26-2 defines as critical analysis by objective experts with sufficient independence to maintain objectivity. What changed is the calibration: validation timing, nature and frequency now vary by model purpose, methodology, rate of change and materiality, rather than defaulting to an annual cycle for everything in the inventory.

The part that matters for anyone reading this for AI purposes is footnote 3. In the agencies' own words: "Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance. Nonetheless, a banking organization's risk management and governance practices should guide the determination of appropriate governance and controls for any tools, processes, or systems not covered in this document. However, the principles described in this guidance apply to traditional statistical and quantitative models and non-generative, non-agentic AI models."

Read that carefully, because it creates the exact gap this article is about. A bank's credit scorecard sits inside a mature, examined model risk framework. The LLM that same bank's relationship managers use to draft client correspondence sits outside it, with the agencies saying only that the bank's own practices should govern. Nobody is coming to hand you a validation standard for that surface. The governance program has to claim it, define what evidence counts, and produce that evidence at examination without a supervisory template telling it what to collect. That is a harder problem than compliance with a prescriptive rule, and it is the one on the table right now.

The shared infrastructure

Both programs depend on the same operational data: who used the AI, when, for what purpose, with what data classification, under what policy, with what outcome. The governance program needs the data to demonstrate the policy is operating. The risk program needs the data to monitor the controls and measure the residual risk.

The data has to be:

Per-request. The aggregated counts from the application layer are insufficient for either program. Both need to trace specific requests.

Identity-bound. The user identity has to be on the record. Both programs build assurance on the verified identity.

Classification-aware. The data classification applied at the moment of the request has to be captured. Both programs evaluate against the classification.

Policy-versioned. The policy in effect at the moment has to be recorded. The governance program reviews the policy evolution; the risk program correlates incidents to policy changes.

Tamper-evident. The records have to be admissible as evidence. Both programs may need to produce the records under regulatory or legal review.

The architecture that produces records with these properties serves both programs. Separate infrastructure for governance and risk creates duplicated effort and gaps where the two data sets diverge.

What each program demonstrates under audit

The governance program demonstrates that the policies are documented, current, and approved by the appropriate roles; that the sanctioned tool list is maintained; that the review cadence is operating; that incidents are escalated according to the policy; that the board-level reporting is happening on schedule.

The evidence is a combination of documents, meeting minutes, and the operational records the policies translate to.

The risk management program demonstrates that the risk register is maintained, that the controls assigned to each risk are operating, that the monitoring metrics are produced, that the residual-risk decisions are documented and approved, that incidents trigger the appropriate response.

The evidence is the risk register itself, the control test results, the monitoring outputs, and the per-request records the controls operate against.

Both programs produce the management cycle outputs: the periodic reviews, the updates to the artifacts, the corrective actions.

What goes wrong when the programs are not aligned

The failure has a signature: policy documents the risk program never references, and risk treatments the policy does not enable. The governance side produces a policy naming a control. The risk side leaves that control out of the assessment because the operational data to test it is not available. The audit then finds a gap between what the policy describes and what the program can demonstrate, and the finding is written against the governance program even though the missing piece was infrastructure.

The fix is the shared infrastructure. The same operational data set drives both. The same per-request records that demonstrate the policy is operating also demonstrate the risk controls are firing. The two programs read from the same source of truth.

DeepInspect

The shared infrastructure both programs depend on is what DeepInspect produces. DeepInspect sits at the AI request boundary as a stateless proxy between users and agents and any LLM. Every request produces a per-decision record under the deployer's control: identity, classification, policy version, decision, outcome, timestamp. The record is tamper-evident.

For the governance program, the records are the operational evidence that the AI usage policy is in effect at the request layer. The sanctioned tool list, the data classification matrix, the role-based access rules - each of them has a per-request representation in the record set.

For the risk management program, the records are the monitoring data and the control test inputs. The metrics on policy violation rates, the analysis of which use cases are producing the most policy-blocked requests, the trend lines on data classification incidents - each draws from the same record set.

For the audit functions of both programs - SOC 2 Type II, ISO 42001 certification, EU AI Act Article 12 disclosure - the records are the evidence the auditor samples against.

If your governance and risk programs are reading from different sources of truth, the alignment work goes on indefinitely. The shared infrastructure removes the gap. Let's talk today.

Frequently asked questions

Should the same person own both AI governance and AI risk management?

In smaller organizations, yes, often the same role covers both. In larger organizations, the governance function typically sits with a chief data officer, a chief AI officer, or a chief compliance officer. The risk function sits with the CISO or the model risk management head depending on the regulated context. The accountability split has to be clear so the audit can identify the responsible role for each artifact. The shared infrastructure operates regardless of who owns each program.

What is the difference between ISO 42001 and ISO 23894?

ISO/IEC 42001:2023 is the certifiable management system standard: policy, roles, internal audit, continual improvement, and the Annex A control set. ISO/IEC 23894:2023 is guidance on the AI risk management process that runs inside it, and it carries no certification. ISO 23894 mirrors the clause structure of ISO 31000:2018 and extends it with AI-specific guidance, so an organization already running ISO 31000 enterprise risk management can adopt it without rebuilding its risk process. Use 42001 for the management system your customers ask to see a certificate for, and 23894 for the method that populates the risk register that system audits.

Is SR 11-7 still current?

No. SR 11-7, issued 4 April 2011, was superseded on 17 April 2026 by SR 26-2, issued jointly by the Federal Reserve, the OCC and the FDIC. SR 26-2 also supersedes SR 21-8. It keeps validation, ongoing monitoring, governance and effective challenge, and replaces uniform validation cadence with an approach calibrated to model materiality. Its footnote 3 places generative AI and agentic AI models outside its scope while keeping non-generative, non-agentic AI models inside. Any AI governance document still citing SR 11-7 as current guidance needs updating.

Does SR 26-2 cover our LLM deployment?

Not directly. SR 26-2 states that generative and agentic AI models are not within the scope of the guidance, while adding that the organization's own risk management and governance practices should determine appropriate governance and controls for systems the document does not cover. The supervisory expectation therefore still exists; the template for meeting it does not. Practically, banks are extending the SR 26-2 principles (materiality assessment, ongoing monitoring, effective challenge, independent review) to LLM deployments by analogy, and documenting that decision so an examiner can see the reasoning rather than an omission.

How does the AI risk register differ from the broader enterprise risk register?

The AI risk register is typically a sub-register of the enterprise risk register or a feeder to it. The AI-specific risks (model drift, prompt injection, unauthorized action by agents, third-party AI vendor failure) sit at a level of detail the enterprise register may not capture. The roll-up to enterprise risk happens through the standard enterprise risk management process. The AI register has the operational detail; the enterprise register has the strategic view.

Does the per-request infrastructure replace the model risk management lifecycle?

No. The per-request infrastructure produces the audit and monitoring data the lifecycle needs. The lifecycle (development, validation, deployment, monitoring, retirement) is a process that the risk management program operates. The per-request data is one input to the monitoring stage. SR 26-2 expects validation and lifecycle management to be present in addition to the operational data, and for generative or agentic systems that sit outside its scope, the bank still has to decide what its own equivalent looks like.

How often should the AI governance program review the policies?

Quarterly is where most regulated programs settle. The AI surface changes faster than annual policy reviews can keep up with, and 2026 is the case in point: SR 11-7 was replaced in April and the EU AI Act's high-risk dates moved in July. An annual review cycle would have carried a stale framework reference for most of a year. Quarterly review with the option for emergency updates when a new regulation lands or a new vendor capability shifts the surface is the pattern. The review covers the policy text, the sanctioned tool list, the role assignments, and the cadence itself.

What does the audit conclude when the programs produce inconsistent evidence?

The auditor flags the inconsistency. If the governance program says the policy includes a control and the risk program says the control is not in the register, that is a finding. If the risk program says a control is operating and the per-request records do not show the control firing, that is a finding. The shared infrastructure prevents both findings because both programs read the same data and the operational data is what the audit ultimately tests against.