← All posts

Platform & Architecture

229 posts on platform & architecture.

NIST AI agent identity Pillars 2 and 3: authorization and audit at the request layer

NIST has framed AI agent identity and authorization around three pillars. Pillar 1 is identification at the request boundary. Pillar 2 is authorization tied to the resolved identity. Pillar 3 is audit and accountability across the request lifecycle. The public comment window on the NIST draft closed April 2, 2026. This piece walks through what Pillars 2 and 3 actually require at the architecture layer and where most enterprise AI deployments fall short.

nistai-agent-identityauthorizationauditzero-trust-ai
Read post →

22-Second Breach Windows: Why AI Enforcement Has to Be Inline

Google Mandiant M-Trends 2026 found median attacker handoff time collapsed from over 8 hours in 2022 to 22 seconds in 2025. Detect-and-respond runs after damage has occurred. For AI traffic specifically, an exfiltrated prompt is one-shot. Inline enforcement at under 50ms overhead is the architectural answer.

inline-enforcementai-securitymachine-speedmandiantprevention
Read post →

LLM Gateway Benchmarks: What to Measure, How to Measure It, and Where Most Vendor Numbers Mislead

Most vendor LLM gateway benchmarks publish a median latency figure under synthetic load and stop there. The numbers a platform team actually needs are the policy-decision tail latency, the policy-evaluation throughput under contention, the cold-cache impact, and the audit-write durability cost. This walkthrough shows the four measurement axes, the workload profiles that produce comparable numbers, and the failure modes that surface only at production traffic shape.

ai-gatewaybenchmarkslatencyperformancepolicy-enforcementengineering
Read post →

AI Rate Limiting by Identity: Why Per-Key Quotas Miss the Actual Risk

A per-API-key rate limit lets one runaway service consume the whole quota for its tenant. An identity-bound rate limit accumulates against the verified caller and produces a defensible refusal at the request layer. This walkthrough covers the four identity dimensions a useful rate limit accumulates against, the algorithms that hold under burst traffic, and the audit-record fields that make a refusal admissible.

rate-limitingai-gatewayidentityengineeringcost-governance
Read post →

AI Gateway Multi-Cloud: The Single Control Plane Across OpenAI, Anthropic, Bedrock, and Vertex

Enterprise AI traffic now spans OpenAI direct, Anthropic direct, AWS Bedrock, Azure OpenAI, and Google Vertex in the same week, often in the same application. Each provider has its own auth, its own request shape, its own error semantics, and its own audit emission. A multi-cloud AI gateway is the single control plane that normalizes identity, classification, policy, and audit across all of them. This walkthrough covers the normalization layer, the per-provider adapters, and the audit record that survives the regulator regardless of which provider the request hit.

ai-gatewaymulti-cloudopenaianthropicbedrockvertex
Read post →

Fail-Closed vs Fail-Open AI Gateway: The Decision That Survives the Post-Incident Review

A gateway whose policy plane is unreachable has to decide whether to forward the request anyway or refuse it. The decision is architectural and the cost of getting it wrong shows up as either a regulatory finding or a production outage. Fail-closed for policy decisions and fail-open for cached bundles is the pairing that survives the post-incident review. This walkthrough covers the failure modes, the per-route defaults, and the operational runbook for the brief window where the policy plane is unavailable.

fail-closedai-gatewayavailabilitypolicy-enforcementarchitecture
Read post →

Fail-closed AI gateway design: why the default failure mode is the security mode

A fail-closed AI gateway returns HTTP 503 when the policy decision point cannot reach a verdict, blocking the request rather than forwarding it. A fail-open gateway returns HTTP 200 with the upstream model response, treating the policy outage as a pass. The choice between the two postures determines whether a policy outage produces a security incident or a contemporaneous deny record. EU AI Act Article 12 and Article 26 expect the deny record. The four failure categories that test the design are policy timeout, identity provider outage, redaction engine outage, and audit write outage.

fail-closedai-gatewaypolicy-enforcementeu-ai-actaudit-logsavailability
Read post →

AI gateway observability: the metrics, traces, and logs a policy decision point should emit

An AI gateway emits four signal categories that serve four different audiences. Per-decision audit logs serve the regulator under EU AI Act Article 12. Per-request traces serve the engineering team debugging a request. Per-policy metrics serve the operations team measuring policy effects. Per-model latency histograms serve the capacity-planning team sizing the LLM provider relationship. OpenTelemetry alignment lets the four signal categories share a transport without conflating their consumers.

ai-gateway-observabilityopentelemetryaudit-logseu-ai-act-article-12policy-metricslatency-histograms
Read post →

AI policy version control: how to treat gateway policy like code

AI gateway policy that governs which users can call which models with which data lives in YAML, evolves with the organization, and carries the same regression risk as application code. Treating the policy as code means git-backed storage, semantic versioning of policy bundles, audit-log tagging of decisions with the policy version hash, blue/green policy rollout, and shadow-mode evaluation before promotion. The NIST AI RMF MAP and MANAGE functions ask the questions the version-control discipline answers.

ai-policy-version-controlpolicy-as-codenist-ai-rmfblue-green-deploymentshadow-modegitops
Read post →

AI prompt classification taxonomy: building the label set your gateway enforces against

AI prompt classification is the labelling step that produces the inputs a policy engine evaluates. The label set has to cover four dimensions: data sensitivity (PII, PHI, PCI, IP, public), intent (query, generation, code execution, agent action), risk surface (egress, lateral, instruction injection), and regulatory scope (EU AI Act high-risk, HIPAA PHI, GDPR Article 22). The policy decision joins the four dimensions against the per-user role and the per-route rule. The taxonomy is the artefact the regulator inspects when the gateway answers an audit question about why a given prompt was redacted, blocked or allowed.

ai-prompt-classificationtaxonomypolicy-engineowasp-llmdata-classificationai-policy
Read post →

AI provider rotation strategy: how to swap OpenAI for Anthropic without breaking policy or audit

AI provider rotation is the operational mechanism that lets a deployer move traffic between OpenAI, Anthropic, Google and other endpoints without breaking the policy decision or the audit trail. The mechanism requires a provider-agnostic policy model, the model identity recorded on every per-decision log, per-route routing rules that decouple the policy from the endpoint, fail-over semantics that hold the policy invariant under provider rate-limit or outage, and a documented concentration-risk posture against DORA Article 28. Rotation is the operational expression of the regulatory expectation that the deployer remains accountable across provider changes.

ai-provider-rotationmulti-providerdora-article-28concentration-riskai-gatewaypolicy-portability
Read post →

AI gateway circuit breakers: limiting blast radius when an LLM provider degrades

An AI gateway circuit breaker adapts the microservice resiliency pattern to LLM traffic. The per-provider state machine moves between closed, open, and half-open based on error rate, latency p99, and token-cost spikes. The trip thresholds, the half-open probe budget, and the breaker telemetry tie to DORA Article 19 incident reporting and produce a recovery audit trail.

ai-gatewaycircuit-breakerresiliencydora-article-19incident-responsellm-traffic
Read post →