← All posts

Platform & Architecture

229 posts on platform & architecture.

Claude Connectors Security: Controlling Remote MCP Server Access

Claude connectors give an application a route to approved remote MCP servers. Secure deployments need a defined server inventory, authorization at the MCP server, narrow tool permissions, trusted connector configuration, and request-level policy for the LLM traffic that begins the flow. This article maps those controls to their actual boundaries.

ai-securityagentic-aillm-securityprompt-injectionidentity-and-authorizationpolicy-enforcement
Read post →

Claude Connectors Audit Logs: Evidence for Remote MCP Tool Calls

Claude connectors let Claude call approved remote MCP servers. Audit evidence for that flow must distinguish the Claude request, the connector authorization, the remote tool action, and the application record. This article shows what each system can prove and where an independent request-level policy record fits.

ai-securityagentic-aiauditforensic-auditidentity-and-authorizationpolicy-enforcement
Read post →

AI Agent Observability: Traces, Decisions, and Debuggable Behavior

AI agent observability needs a trace for each model call, tool handoff, retrieval result, and policy decision. This guide defines the signals, OpenTelemetry fields, correlation IDs, and retention split that make autonomous behavior reviewable.

ai-agent-observabilityai-agent-securityopentelemetryai-securityagentic-aitraces
Read post →

AI Gateway Architecture: The Components That Sit Between an Enterprise Caller and an LLM Endpoint

An AI gateway architecture has six core components: TLS termination, identity binding, request inspection, policy evaluation, the model router, and the audit record emitter. Each component is a placement decision that ties to a regulatory obligation or an operational property. This piece walks through the components, the placement decisions, and how the gateway integrates with the corporate IdP and the SIEM.

ai-gateway-architectureai-gatewayai-securityinline-enforcementaudit-logs
Read post →

OpenAI Usage Tier Controls: How an Enterprise Enforces Per-Team Budgets on the Same API Key

OpenAI's account usage tiers describe the account-level rate ceiling. The tier is a single number the account holds, and the tier does not describe the enterprise's per-team, per-application, or per-user budgets. An enterprise that runs OpenAI at scale has to enforce a set of budget controls that sit above the account tier. This piece walks through the pattern set: per-team token budgets, per-application spend caps, per-user rate ceilings, and the audit records that tie every request back to the team accountable for it.

openaiusage-tierai-costai-gatewaybudget-controlai-security
Read post →

AI Tenant Isolation: How Multi-Tenant SaaS Enforces Per-Customer Boundaries on LLM Traffic

Multi-tenant SaaS applications that add LLM features carry a new isolation obligation on top of the database and storage isolation the platform already enforces. Prompts flow through the LLM provider carrying tenant-specific data. Retrieval-augmented generation queries the vector store where tenant data lives. Agent tools call downstream systems that hold tenant data. Each of these paths introduces a way for tenant A's data to reach tenant B's context without a database join between them. This piece walks through the four isolation domains (prompt, retrieval, tool call, response), the enforcement patterns at the AI gateway, and the audit records that demonstrate the isolation held across the audit period.

multi-tenantai-securitytenant-isolationsaasragai-gateway
Read post →

LLM fallback routing: the retry chain that survives provider outages without leaking policy

LLM fallback routing chains a primary model to a secondary and tertiary so provider outages, rate-limit errors, and quality regressions do not cause user-visible failures. The failure modes are usually not the fallback logic itself but the boundary between the fallback chain and the policy decision that authorized the request. This piece walks through the four common triggers for fallback, the retry semantics per trigger, the authorized-endpoint constraint, and the idempotency requirements for tool-calling workloads.

llm-routerllm-fallbackai-architectureai-gatewayreliability
Read post →

AI Agent Identity: NIST Pillar 1 in Production Deployments

NIST Pillar 1 names verified agent identity as the foundation of the AI agent identity and authorization framework. Per-agent identifiers, delegated authority from the authorizing user, and structured propagation to the model API call are the production requirements. Static service credentials fail the test.

agentic-aiidentity-and-authorizationnist-ai-rmfai-securityarchitecturezero-trust
Read post →

Zero Trust AI: Per-Request Evaluation at the Model Boundary

Zero trust applied to AI means evaluating every model request against verified identity, current policy, and prompt-level classification. The architectural pattern is an enforcement proxy at the HTTP AI request boundary. The post-authentication gap is the most common failure mode in current deployments.

zero-trustai-securityidentity-and-authorizationpolicy-enforcementinline-enforcementarchitecture
Read post →

Model Guardrails Are Probabilistic, Not Enforceable Controls

Model guardrails are trained behaviors inside the inference process. They degrade under fine-tuning, adversarial prompting, and role-play framing. External enforcement at the AI request boundary produces deterministic controls and identity-bound audit records that guardrails alone cannot.

ai-securityllm-securityprompt-injectionpolicy-enforcementarchitectureinline-enforcement
Read post →

22-Second Breach Windows: Why AI Enforcement Must Be Inline

Mandiant M-Trends 2026 measured median attack handoff at 22 seconds. At that tempo, log-and-alert fails as a control. Inline enforcement at the AI request boundary makes the policy decision before the request reaches the model. Under 50 ms enforcement overhead is invisible against 500 ms to 5 second model inference.

ai-securityinline-enforcementpolicy-enforcementcybersecurityarchitecturezero-trust
Read post →

Identity-Aware AI Gateway Architecture: How Inline Enforcement Binds Decisions to Users and Agents

An identity-aware AI gateway sits at the AI request boundary, attaches verified identity context to every model API call, evaluates per-route and per-role policies, and commits a per-decision audit record before the model response returns to the calling application. The architecture closes the post-authentication gap that most enterprise AI deployments have inherited from the credential-pooling pattern used by SDKs and proxy frameworks. This piece walks through the architectural building blocks, the call path, the audit primitives, and where the identity-aware gateway sits relative to existing IAM, API gateway, and DLP infrastructure.

ai-gatewayidentity-awareai-architectureenforcementauditzero-trust
Read post →