NIST GenAI Profile (NIST AI 600-1): Risks, Actions, and Evidence
NIST AI 600-1 is a voluntary Generative AI Profile with 12 risk categories and suggested actions under GOVERN, MAP, MEASURE, and MANAGE. This current guide explains confabulation, correct action IDs, the relationship with OMB M-25-21, and the evidence an organization can collect without pretending one runtime log satisfies the whole framework.

On July 26, 2024, NIST published the Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, designated NIST AI 600-1. The 64-page document gives organizations a GenAI-specific companion to AI RMF 1.0. It identifies 12 risk categories, then places suggested actions under selected GOVERN, MAP, MEASURE, and MANAGE subcategories.
That last sentence carries two details that matter. NIST describes the AI RMF and this cross-sectoral Profile as voluntary. Section 3 calls its entries “suggested actions,” and organizations choose actions based on their role and context. The document is useful as a risk and control map. It is a poor candidate for checkbox theatre.
I want to correct a second source of confusion at the outset. OMB M-24-10 has been obsolete since April 3, 2025. Its replacement, M-25-21, sets current federal agency policy for high-impact AI. The GenAI Profile remains useful guidance, while OMB’s requirements arise under a separate memorandum.
The official NIST AI 600-1 structure
The official NIST AI 600-1 publication has three working layers.
Section 1 begins by defining scope. NIST designed the Profile for organizations that design, develop, deploy, use, or evaluate generative AI. The Profile extends the NIST AI Risk Management Framework, so its actions sit under the RMF’s four functions instead of forming a separate certification system.
Section 2 describes 12 risks:
- CBRN Information or Capabilities: access to information or design capabilities connected with chemical, biological, radiological, or nuclear weapons and dangerous materials.
- Confabulation: confidently presented output that is erroneous or false.
- Dangerous, Violent, or Hateful Content: generated content that can incite harm, aid illegal activity, or expose people to hateful material.
- Data Privacy: leakage, unauthorized use, disclosure, or de-anonymization of personal and sensitive data.
- Environmental Impacts: compute, energy, water, and related effects associated with model training and operation.
- Harmful Bias or Homogenization: amplified bias, performance disparities, and output homogeneity that can produce harmful decisions.
- Human-AI Configuration: automation bias, over-reliance, anthropomorphism, algorithmic aversion, and emotional entanglement.
- Information Integrity: lowered barriers to misinformation, disinformation, and content that obscures uncertainty or origin.
- Information Security: offensive cyber capability, attacks against AI components, and compromise of data, code, systems, or model weights.
- Intellectual Property: unauthorized reproduction, exposure of trade secrets, plagiarism, and related rights issues.
- Obscene, Degrading, and/or Abusive Content: generated abusive imagery, synthetic child sexual abuse material, and nonconsensual intimate imagery.
- Value Chain and Component Integration: opaque third-party components, weak supplier vetting, and untraceable upstream data or software.
Section 3 contains the suggested actions. Each action has a precise identifier such as GV-3.2-003 or MG-4.3-002. The prefix identifies the RMF function: GV for GOVERN, MP for MAP, MS for MEASURE, and MG for MANAGE. The following numbers identify the RMF subcategory and the specific Profile action.
A search phrase such as “GenAI classification B, GenAI is permitted as an assistive tool” describes an organization’s possible internal usage tier. It is absent from the official NIST taxonomy. Teams can create an assistive-use tier, but the policy should label it as an internal classification and map it to applicable NIST actions.
Confabulation is a production risk
NIST uses confabulation for confidently stated erroneous or false content, also commonly called hallucination or fabrication. The Profile notes that a generated answer can contradict its prompt, earlier statements in the same context, or known facts. It can also invent logic and citations that make the answer look more credible.
Picture a case reviewer with a model-generated summary on the left monitor and a 300-page source file on the right. The summary gives the wrong medication, contract date, or eligibility fact in polished prose. Confidence in the wording becomes part of the failure mechanism.
The Profile ties confabulation to several concrete actions. MP-2.3-001 suggests comparing output with known ground truth through human and automated evaluation. MP-2.3-003 suggests documented fact-checking techniques, especially when output draws on multiple or unknown sources. MG-4.3-002 covers policies and procedures to record and track reported errors, near-misses, and negative impacts. A practical implementation needs an evaluation set, reviewer instructions, an error intake path, and a retained record connecting the faulty output to its model version and use context. Our guide to LLM hallucination controls covers that operational layer in more detail.
My opinion is blunt: a generic instruction telling staff to “verify AI output” is paperwork dressed as a control. A reviewer needs an authoritative source, a defined decision threshold, and a place to record disagreement. High-consequence output should remain outside automated decision paths until that test exists.
Correct action IDs change the control map
A subcategory label summarizes an objective. A full action ID points to a suggested activity. Mixing them produces bad audit requests, and the original version of this article did exactly that.
These examples come directly from Section 3 of NIST AI 600-1:
- GV-3.2-003 suggests acceptable-use policies for GenAI interfaces, modalities, and human-AI configurations, including criteria for queries an application should refuse. GV-3.2 broadly concerns roles, responsibilities, human-AI configurations, and oversight.
- GV-6.2-002 covers documenting incidents involving third-party GenAI data and systems. GV-6.2-003 addresses third-party incident-response plans, ownership, rehearsals, retrospective learning, communication, and review against applicable law.
- MP-2.3-001 addresses output accuracy, quality, reliability, and authenticity against known ground truth. MP-2.3-003 focuses on fact-checking. MP-2.3 is the scientific-integrity and testing, evaluation, verification, and validation subcategory.
- MP-4.1-001 suggests periodic monitoring of generated content for privacy exposure. MP-4.1-006 addresses policies for third-party intellectual property and training data. MP-4.1 spans technology and legal risks in components, including third-party data or software.
- MP-5.1-002 calls for identifying and ranking potential content-provenance harms by likelihood and impact. MP-5.1-003 addresses disclosure of GenAI use in relevant contexts after considering purpose, audience, risk, and frequency.
- MS-2.7-001 covers established security measures for threats including compromised dependencies, data breaches, eavesdropping, model theft, extraction, and inference attacks. MS-2.7 evaluates security and resilience. It is broader than privacy testing.
- MS-2.12-002 suggests documenting anticipated environmental impacts in product design decisions. MS-2.12-003 addresses measurement or estimation of energy and water use for training, fine-tuning, and deployment. MS-2.12 concerns environmental impact and sustainability.
- MG-2.2-001 suggests comparing output against organizational risk tolerance and reviewing generated content against defined guidelines. MG-2.2-007 suggests real-time auditing where it can help track and validate lineage and authenticity. MG-2.2 concerns sustaining value in deployed systems.
- MG-4.3-001 addresses after-action assessments and incident communication. MG-4.3-002 covers error and near-miss tracking. MG-4.3-003 addresses legally required incident reporting.
This precision affects ownership. A sustainability lead might own MS-2.12 evidence. Security engineering can own part of MS-2.7. Legal and data governance have central work under MP-4.1. Incident response owns much of MG-4.3, supported by records supplied by application and infrastructure teams.
GOVERN, MAP, MEASURE, and MANAGE need different evidence
GOVERN evidence establishes authority and accountability. Useful artifacts include an approved acceptable-use policy, named owners, review records, third-party incident plans, and proof that staff received role-specific instructions. A policy PDF sitting untouched since 2024 offers weak evidence of active governance.
MAP evidence connects a system to its context. Record the intended purpose, users, affected people, upstream data and components, external connections, reasonably foreseeable misuse, and the likelihood and magnitude of identified impacts. An AI inventory provides the index. The system card, data records, supplier documentation, and impact analysis supply the detail.
MEASURE evidence demonstrates that somebody tested a claim. Preserve the evaluation dataset version, model and configuration, test procedure, result, reviewer, date, threshold, and disposition. A production monitoring alert and a pre-deployment benchmark answer different questions. Our NIST AI RMF MEASURE controls guide gives teams a more detailed testing structure.
MANAGE evidence shows what happened after a risk became visible. Retain risk acceptance, mitigation changes, release decisions, monitoring results, user feedback, incidents, corrective actions, and after-action review. The record should connect the decision to the policy and test result that informed it. A useful evidence package lets an assessor trace one sampled output without opening six dashboards and a spreadsheet called final-v7.
M-25-21 replaced M-24-10
OMB Memorandum M-25-21, issued April 3, 2025, expressly rescinded and replaced M-24-10. Current federal content should use M-25-21’s term high-impact AI. The earlier “safety-impacting” and “rights-impacting” categories belong to superseded policy.
M-25-21 defines high-impact AI around output serving as a principal basis for decisions or actions with legal, material, binding, or significant effects on rights or safety. For covered agency uses, Section 4 requires minimum practices including pre-deployment testing, a documented AI impact assessment, ongoing monitoring and periodic human review where feasible, operator training, suitable human oversight, timely human review and appeal when appropriate, and public or end-user feedback where appropriate.
Those are OMB requirements for covered agency use. NIST AI 600-1 remains voluntary guidance, and M-25-21 does not convert every suggested Profile action into a federal mandate. An agency can use the Profile to organize risks and candidate mitigations while maintaining a separate M-25-21 compliance matrix tied to the memorandum’s operative text.
Federal contractor scope also needs care. OMB Memorandum M-25-22 governs agency acquisition practices and excludes AI used incidentally at a contractor’s option when the contract neither directs nor requires that use. Contract language, the acquired system, agency direction, and actual performance determine the contractor’s duties. A federal customer logo on a slide cannot answer that analysis.
A defensible implementation sequence
Start with one GenAI use case rather than a framework-wide spreadsheet. Name its owner, intended purpose, users, model route, data classes, affected people, upstream components, and decision consequence. Select the applicable risks among NIST’s 12 categories and explain each exclusion.
Select full action IDs during the next working session, then assign an owner and evidence artifact to every chosen action. An acceptable-use policy plus a test of prohibited request handling can support action GV-3.2-003. A versioned ground-truth evaluation provides evidence for MP-2.3-001. The error register should connect sampled requests and corrective actions under MG-4.3-002.
Testing comes after the owners and artifacts are clear. Record the threshold before running the evaluation. Capture exceptions, reviewer disagreement, and failed cases. The AI compliance logging guide explains the fields needed to reconstruct individual runtime decisions.
Finally, review the map after a model, prompt template, retrieval source, integration, user population, or decision purpose changes. NIST’s Profile repeatedly emphasizes context and lifecycle activity. A static annual attestation misses the point.
DeepInspect
DeepInspect sits on a narrow enforcement boundary: authenticated HTTP requests routed through the gateway between users or agents and LLM endpoints. At that point it can apply identity-aware policy and record the principal, route, model, timestamp, policy version, classification result, permit or deny decision, and related response metadata.
Those records can contribute evidence to GV-3.2-003 acceptable-use enforcement, MG-2.2 output monitoring, MS-2.7 security measurement for observable routed traffic, and MG-4.3 incident investigation. They also make a sampled event retrievable: an investigator can connect a person or agent, a model route, the policy active at that moment, and the gateway decision.
The wider Profile remains a multi-owner program. DeepInspect leaves training-data provenance, model weights, benchmark design, environmental measurement, IP review, supplier contracts, interface design, model red-teaming, local execution, and non-HTTP activity to the teams and systems that can actually observe them. A gateway record contributes evidence. It never satisfies the whole NIST GenAI Profile.
If you are mapping routed GenAI traffic to NIST AI 600-1, book a technical deep dive at deepinspect.ai.
Frequently asked questions
- Is the NIST GenAI Profile mandatory?
NIST describes the AI RMF and its Generative AI Profile as voluntary resources. A law, regulator, contract, or agency policy can separately require practices that overlap with the Profile. Teams should identify that separate authority instead of relabeling NIST’s suggested actions as mandates.
- Is “GenAI classification B” an official NIST category?
No official “classification B” appears in NIST AI 600-1. An organization may define an internal tier in which GenAI is permitted as an assistive tool, then map its restrictions and evidence to applicable Profile actions. The tier name remains internal taxonomy.
- What does NIST mean by confabulation?
NIST defines confabulation as confidently presented erroneous or false content. The term includes output that diverges from the prompt, contradicts earlier statements, or invents supporting logic and citations. The production control set can include ground-truth evaluation, documented fact-checking, human review, error intake, and near-miss tracking.
- How does M-25-21 relate to NIST AI 600-1?
M-25-21 is current OMB policy for federal agency AI use and contains minimum practices for covered high-impact AI. NIST AI 600-1 is voluntary cross-sectoral guidance. Agencies may use the Profile to structure risk analysis, but the source of a federal duty remains M-25-21 or another applicable authority.
- Can one audit log prove implementation of the Profile?
A runtime log can support selected actions by showing identity, model route, policy state, classifications, decisions, and timestamps for traffic within its capture boundary. Other actions require evaluations, supplier records, training-data documentation, environmental measurements, legal analysis, human-factors work, and incident procedures. A credible evidence package keeps those boundaries visible.