Shadow AI in Energy and Utilities: Grid Data, Outage Plans, and Vendor Risk
Shadow AI in energy and utilities can transmit grid operations material, outage plans, engineering records, and vendor data to unapproved LLM services. This article maps that exposure to NERC CIP information-protection and supply-chain requirements, distinguishes regulated BES Cyber System Information from other sensitive utility data, and defines enforceable controls for authenticated HTTP AI traffic.

At 5:18 a.m., a transmission planner pastes a switching sequence and next week's outage window into an AI assistant to tighten the handoff note. A one-line diagram is open on the left monitor, with a coffee ring touching the printed contingency sheet below it. The prompt may also contain station names, equipment identifiers, loading assumptions, and a vendor ticket. One HTTPS request can move that operational context to an unapproved model before the grid security or engineering governance teams see the route. Vendor management misses it too.
Shadow AI in energy and utilities concentrates risk in the prompt. The sensitive object may be a few copied rows rather than a complete operating plan. Its model destination and user identity still determine what left the utility. Provider terms determine how that information may be handled and what evidence remains.
TL;DR
- Unauthorized LLM traffic can carry switching sequences and outage windows outside approved workflows. Engineering findings and customer records can leave through the same route. Vendor access details can too.
- NERC CIP-011-3 requires applicable responsible entities to identify and securely handle BES Cyber System Information. Utility information outside that defined scope can remain operationally sensitive.
- NERC CIP-013-2 addresses supply-chain cyber risk for BES Cyber Systems. An unknown model provider or vendor-controlled AI feature creates a separate review and evidence path.
- DeepInspect enforces policy on authenticated HTTP AI traffic routed through it. Browser traffic outside that routing is unseen and needs separate browser or endpoint safeguards, along with egress controls.
Operational context survives copy and paste
An outage plan carries meaning through relationships. A station name beside a planned date, affected elements, expected loading, and a switching order may disclose more than any field alone. An operator who removes the document title can still send the useful operational content. Engineering prompts create the same problem. They may contain relay settings and protection studies. Equipment-health findings and network diagrams can appear alongside vulnerability remediation notes.
The generic shadow AI article explains how unapproved model routes bypass procurement and policy. Utilities add grid topology and restoration context. They also add field procedures and customer information. Vendor access arrangements create another source. Some of that material may meet the entity's definition of BES Cyber System Information. Some sits outside NERC CIP applicability while remaining sensitive under company policy or contracts. State requirements and operational security practice can apply too.
I would reject a shadow AI register that records only provider domains. A hostname cannot tell a reliability manager that the prompt contained a blackstart reference or an outage constraint. It also misses a contractor's remote-access procedure. The useful record binds the actual content to a person and purpose, then records the route and decision.
NERC CIP-011 makes information handling specific
NERC CIP-011-3 is the mandatory information-protection standard effective January 1, 2024. Its purpose is to prevent unauthorized access to BES Cyber System Information by specifying protection requirements that support the security of BES Cyber Systems. Requirement R1 calls for documented information-protection programs for applicable high- and medium-impact systems. The standard's measures include methods to identify BCSI and records showing that it was handled according to documented procedures, including protection during transit.
That scope needs careful wording. CIP-011 applies through the standard's named responsible entities and facilities, as well as its named systems and impact categories. It never turns every utility spreadsheet into BCSI. The entity's own identification method and the standard's applicability determine the regulated set.
Shadow AI can break the handling path even when the source repository is well controlled. Once a user copies identified BCSI into an unapproved prompt, labels may vanish. The transmission may also fall outside the documented route. For other engineering and outage information, utility policy should create additional classifications rather than stretching the NERC definition.
Outage work creates concentrated prompts
Operations planning teams work with outage dates and affected facilities. They also use contingency assumptions and transfer limits. Control-room support teams hold switching orders and operator logs. They also hold restoration procedures and voice-transcript summaries. A request to condense a turnover note can assemble several pieces into a compact prompt that explains the grid's expected state.
Engineering teams add protection settings and disturbance findings. They also add inspection images and asset-health data. A model used for code assistance may receive scripts that query historians or generate study inputs. The code alone can expose naming conventions and data sources. It can reveal internal interfaces too. An approved public-document assistant should never inherit blanket permission for those materials.
Customer and emergency communications create another route. Draft outage notices can contain addresses and medical-need indicators. Account status and crew locations can appear before public release, along with estimated restoration times. The control decision needs both a content classification and a business purpose rather than a broad label such as "communications."
The existing AI governance for energy and utilities article covers the full use-case register and governance boundary. Shadow AI is the unauthorized traffic branch inside that larger program. The model call occurs before the registered owner and approved destination can govern it, and before the evidence plan can take effect.
Vendor data expands the request boundary
Utilities exchange architecture diagrams and maintenance records with vendors. They also exchange patch plans and remote-access procedures. Equipment defect reports move between the same parties. An employee may paste a support case into an LLM. A field-service company may summarize the same case inside its own tenant. An enterprise asset or outage platform may also add a model feature whose inference call runs entirely inside the vendor's environment.
NERC CIP-013-2, effective October 1, 2022, requires applicable responsible entities to implement supply-chain cyber risk management controls for BES Cyber Systems. Its stated purpose is to reduce cyber risk to reliable BES operation. The precise applicability and required plan govern the compliance conclusion. The standard should never be cited as a universal AI-vendor rule.
It still exposes the right architectural questions. The utility needs the provider and subprocessor. It needs data use and retention terms, along with the access model. Incident duties and the change process need review. Available logs matter too. It also needs to know who can send which information through that route. A vendor assessment supplies relationship evidence. Request enforcement supplies event evidence when the model traffic passes through infrastructure the utility controls.
DOE separates AI benefit from infrastructure risk
The Department of Energy's April 2024 risk assessment for AI in critical energy infrastructure identifies four broad risk categories: unintentional failure modes, adversarial attacks against AI, hostile applications of AI, and compromise of the AI software supply chain. DOE also called for regularly updated, risk-aware guidance and further work on the public availability of energy-sector data.
Shadow AI intersects that analysis through uncontrolled data movement and unknown services. A planner may expose system context to a provider that security never assessed. An engineering agent may include retrieved vendor documentation and operating data in the same request. A SaaS vendor may change its model subprocessor while the utility's original contract review remains static.
The risk assessment covers a much wider field than LLM traffic. It includes AI applications across critical energy infrastructure, along with security concerns beyond a proxy's reach. Request controls address one concrete part: the authenticated HTTP exchange between a utility-controlled user or agent and an LLM endpoint. Grid safety analysis, model validation, OT security, and supply-chain assurance remain separate disciplines.
Identity and operational purpose drive policy
A managed request should carry the authenticated human or agent identity and source application. It should also carry the role and approved use case. Operational context may include the control center or business unit and the asset class. It may also include the data owner and sensitivity label. The change or outage identifier supplies further context. The application owns the accuracy of those attributes. A shared API credential reduces the decision to a shared principal. That weakens the investigation record.
Prompt classification can look for station and circuit naming patterns or topology descriptions. It can detect switching language and equipment identifiers. Relay settings and outage windows need coverage too, as does restoration terminology. Vendor case numbers and remote-access instructions deserve separate detectors. Customer identifiers and account fields need their own treatment. The policy combines identity and purpose with the destination before deciding what happens next. It can allow or deny the request. It can also redact it.
HTTP-layer AI policy enforcement describes this decision point in more detail. For a utility, a permitted request might send public tariff language to an approved model. A prompt containing restricted outage details can receive a denial on the same route under the effective policy version.
Browser and embedded traffic define the blind spots
DeepInspect sees authenticated HTTP AI traffic deliberately routed through its proxy. A consumer AI session opened in a browser can use another path. Browser traffic outside that routing is unseen by DeepInspect, so a utility needs secure web gateway and enterprise-browser controls. Endpoint telemetry and DNS monitoring can also discover or stop un-routed use. Egress analysis and managed-device restrictions provide further coverage. The shadow AI detection guide explains how those discovery layers complement inline enforcement.
Embedded vendor AI sets a second boundary. When an outage-management or asset-management platform makes the model call inside the vendor's environment, the utility's proxy never receives that exchange. The same applies to field-service and customer platforms. Contract terms and technical documentation need to cover it. So do vendor attestations and audit exports. Ongoing monitoring is required as well.
Local models and non-HTTP transports sit outside the proxy as well. So do OT commands and protective-relay actions. Industrial protocols remain outside it too. Keeping these exclusions on the architecture diagram prevents a managed AI route from becoming a claim of universal visibility.
Evidence needs the request-level event
A useful shadow AI evidence package starts with the approved use-case inventory and data-flow diagram. It includes provider review and handling rules. Technical routing and training belong in the package too. So do exceptions and incident procedures. Operational records then show that the managed control ran for a particular exchange.
For each routed request, the record should identify the person or agent and source application. It should capture the purpose and detected information category. The model destination and policy version belong in the same record. So do the timestamp and outcome. Response events should link back to the request. Denials matter. They prove the control acted before transmission and can reveal a recurring workflow that needs a sanctioned alternative.
Retention deserves a separate decision. Storing every prompt in full may create a new repository of grid and customer data. A utility can retain classifications and fingerprints according to its evidence needs and records rules. It can also retain references and selected content. The gateway record complements operations and engineering logs. It also complements vendor and provider logs. It supplies an independent policy event for the HTTP route rather than replacing the broader NERC compliance file or the utility's incident record.
DeepInspect
DeepInspect sits inline between authenticated utility applications or agents and HTTP-based LLM endpoints. The application supplies identity and operational context. DeepInspect classifies routed prompts and responses, then applies per-role and per-route policy before traffic proceeds. It can permit or deny a request. It can also redact content according to the utility's versioned rules.
Every routed decision produces an identity-bound audit record containing the detected information category and destination. It also contains the policy version and timestamp, along with the outcome. Those records support the managed AI traffic portion of a utility's governance and evidence program. Consumer browser sessions, vendor-controlled model calls, local inference, and OT traffic that bypass the proxy remain outside DeepInspect's boundary.
Book a demo today.
Frequently asked questions
- Does every outage plan count as BES Cyber System Information?
No. NERC CIP-011-3 has defined applicability, and each responsible entity needs documented methods to identify BCSI for the applicable systems. An outage artifact may contain BCSI. It may instead sit outside that definition while remaining sensitive under utility policy or contract. State requirements and operational security practice can apply too. Security and compliance teams should classify the actual information and record the basis. Calling every outage document BCSI makes the program inaccurate and can obscure the materials that require the standard's specific handling controls.
- Can a utility block all shadow AI with an LLM gateway?
A gateway can block or shape authenticated HTTP model calls routed through it. It can evaluate identity and content before transmission. It also evaluates purpose and destination. Consumer browser traffic that avoids the configured route remains unseen. Vendor-internal inference and local models also bypass it. Utilities need browser and endpoint controls, along with DNS and egress monitoring. Procurement discovery and vendor assurance cover the other surfaces. The architecture works when each control owns a named path rather than one product claiming every path.
- What should an energy company record about a denied prompt?
The denial record should preserve the caller or agent and source application. It should capture the purpose and detected category. The intended model endpoint and policy version belong in the same record. So do the timestamp and decision. It may also retain a content fingerprint or controlled reference according to the utility's records design. That evidence shows that the payload stopped before the LLM received it. It also gives security and operations enough context to determine if the event was a mistaken workflow or a policy-tuning issue. Some events will instead require an incident investigation.