← Blog

DeepInspect vs CalypsoAI: Runtime Enforcement vs Model Evaluation

CalypsoAI, now part of F5 after a September 2025 acquisition, evaluates and red-teams AI models before and during deployment. DeepInspect enforces identity-bound policy on every live production request across any LLM provider. This comparison lays out where model evaluation ends and per-request runtime enforcement begins, and when a security team needs both.

ByParminder Singh· Founder & CEO, DeepInspect Inc.
Comparisons & Alternativesai-securityllm-securityai-governancezero-trustpolicy-enforcementarchitecture
DeepInspect vs CalypsoAI: Runtime Enforcement vs Model Evaluation

On September 26, 2025, F5 closed its acquisition of CalypsoAI and folded the platform into its Application Delivery and Security Platform, where the technology now ships as F5 AI Guardrails. The product itself did not change shape: model evaluation, red-teaming, and adversarial testing that scores whether a given model and configuration are safe to deploy. That is a pre-deployment question. DeepInspect answers a runtime one: for this specific request, right now, who is calling, what are they permitted to do, and can the decision be proven afterward. Buyers researching a CalypsoAI alternative are usually asking about the second question.

TL;DR

  • CalypsoAI, acquired by F5 on September 26, 2025, tests and red-teams candidate models before and during deployment. It ships today as F5 AI Guardrails.
  • DeepInspect enforces identity-bound policy on every live AI request in production, across any LLM provider, and signs the audit record for each decision.
  • Choose CalypsoAI or F5 AI Guardrails to decide which models are safe to approve. Choose DeepInspect to control and prove what happens on each live call.
  • The two sit at different points in the AI security stack and often run together, not instead of each other.

CalypsoAI

CalypsoAI built its platform around one question: is this model, in this configuration, safe to put in front of users? The product runs adversarial testing against known jailbreak and prompt-injection techniques, then scores the result so a security or model-risk team can decide which models and configurations clear the bar for production use.

F5 closed its acquisition of CalypsoAI on September 26, 2025, according to F5's own announcement and Cooley's deal coverage. The technology now operates inside F5's Application Delivery and Security Platform under the F5 AI Guardrails name. The evaluation cadence has not changed with the new owner. Red-teaming and adversarial testing run before a model ships and again on a recurring schedule afterward, checking behavior against a representative set of attack patterns. That produces a snapshot, repeated on a schedule, of how a model behaves under test conditions. Watching the traffic that actually reaches the model once it is live is a separate exercise, one the evaluation cycle was never built to do.

DeepInspect

DeepInspect answers a different question: for this specific request, right now, who is calling, what are they authorized to do, and can the decision be proven after the fact. DeepInspect is a stateless proxy that sits inline between authenticated users and agents and any LLM. It is model-agnostic. The same enforcement point works in front of OpenAI, Anthropic, Amazon Bedrock, Azure OpenAI, Google Vertex, and self-hosted models running Llama or Mistral. Every request is evaluated against identity and policy before it reaches the model, and every decision produces a signed, tamper-evident audit record that persists independent of the calling application.

I read the F5 press release the week it landed, sitting next to two customer security questionnaires that still listed CalypsoAI as a named line item. Neither questionnaire asked who made the call at 2 a.m., or whether the decision was ever logged somewhere the application itself could not touch. My opinion: that gap will surface in an incident postmortem long before it surfaces in an RFP, and it is the one DeepInspect closes.

Book a demo today.

Feature comparison

The two products rarely show up on the same procurement line item, because they sit at different points in an AI deployment's lifecycle.

  • Evaluation cadence. CalypsoAI, now F5 AI Guardrails, tests candidate models before deployment and on a recurring schedule after. DeepInspect evaluates every request as it arrives, in production.
  • Enforcement point. CalypsoAI's testing runs against representative attack patterns, separate from the live request path. DeepInspect sits inline on the live request path itself.
  • Identity binding. CalypsoAI's trust score attaches to a model and configuration, not to the person or agent making a specific call. DeepInspect binds every policy decision to the authenticated identity behind the request.
  • Audit output. CalypsoAI produces evaluation and red-teaming reports on the models tested. DeepInspect produces a signed, tamper-evident record for each individual decision.
  • Model coverage. F5 AI Guardrails ships as part of the F5 Application Delivery and Security Platform. DeepInspect is model-agnostic and works in front of any HTTP-based LLM endpoint, regardless of infrastructure vendor.
  • Primary buyer question answered. CalypsoAI, or F5 AI Guardrails, answers which models and configurations are safe to approve. DeepInspect answers whether this specific request, by this specific caller, was permitted, with a record to prove it.
  • Deployment layer. CalypsoAI sits in the model risk and approval process, upstream of deployment. DeepInspect sits at the AI request boundary, between callers and the model API, on traffic that has already been approved to run.

Pick CalypsoAI if...

CalypsoAI, now F5 AI Guardrails, is the right call when the open question is still about which models clear evaluation, not which requests get enforced live.

  • Your immediate priority is red-teaming candidate models before they reach production, ahead of enforcing policy on traffic that is already live.
  • You need a trust score that tells a model-risk committee which models and configurations cleared adversarial testing.
  • You want ongoing behavioral validation against known jailbreak and prompt-injection techniques as new model versions ship.
  • You are already standardized on the F5 Application Delivery and Security Platform and want the evaluation layer to live inside it.

Pick DeepInspect if...

  • You need every live AI request evaluated against identity and policy, regardless of which model handled it or how it scored in testing.
  • An auditor or regulator will ask for a signed record of one specific decision, not a summary report on model behavior.
  • You run more than one LLM provider in production and want a single enforcement layer across all of them.
  • The question you cannot currently answer is who made this exact call and whether they were permitted to, no matter which model was approved for use.

Frequently asked questions

Are DeepInspect and CalypsoAI (now F5 AI Guardrails) direct competitors?

They compete for budget more than they compete for the same architectural slot. CalypsoAI, now part of F5's Application Delivery and Security Platform as F5 AI Guardrails, sells into the model-risk and AI-governance budget: which models get approved, how they hold up against adversarial testing, and what trust score a committee can point to. DeepInspect sells into the runtime security and compliance budget: enforcing identity-bound policy on every live request and producing a signed audit record for each decision. A CISO evaluating both is often filling two different gaps in the same program, not choosing one over the other. The overlap shows up mainly in messaging, since both vendors talk about AI security in their marketing, and buyers researching a CalypsoAI alternative sometimes assume the categories are interchangeable. Once you trace where each product actually sits, the categories separate cleanly: one evaluates models before and during deployment, the other enforces policy on the traffic those models handle once they are live.

What happened to CalypsoAI after the F5 acquisition?

F5 announced its agreement to acquire CalypsoAI in mid-September 2025 and closed the deal on September 26, 2025, per F5's own announcement and Cooley's deal coverage. CalypsoAI became a wholly owned subsidiary, and its red-teaming and model-evaluation technology was integrated into F5's Application Delivery and Security Platform, where F5 now markets it as part of F5 AI Guardrails. The founding team and product line carried over into F5's security portfolio rather than being shut down or sold off in pieces. For buyers, the practical effect is that CalypsoAI as a standalone vendor name is fading from new contracts even as the underlying evaluation and red-teaming capability continues to ship, now bundled with F5's broader application delivery and security stack. Procurement teams still running security questionnaires that list CalypsoAI as an independent line item should expect that line to route through F5 going forward. The acquisition changed who owns the product and where it sits on a buyer's vendor list. The evaluation cadence the product runs on is the part that stayed the same.

Can I run DeepInspect alongside F5 AI Guardrails?

Yes, and the two are built to answer different questions rather than duplicate the same control. F5 AI Guardrails, the former CalypsoAI product, tells a model-risk committee which models and configurations passed adversarial testing before deployment. DeepInspect sits at the AI request boundary and enforces identity-bound policy on every request that actually reaches an approved model in production, then signs the audit record for each decision. An organization using F5 AI Guardrails to gate which models get approved can still put DeepInspect inline in front of those same approved models, to control who is allowed to call them, with what data, under what policy, and to produce evidence that a specific request was handled correctly. Running both closes the pre-deployment evaluation gap and the runtime enforcement gap at the same time. The two sit at different points in the request path: one upstream of the deployment decision, one inline on live traffic.

Does model red-teaming replace the need for runtime enforcement?

Red-teaming tells you how a model behaved against a specific set of attack patterns on the day it was tested. It leaves open the question of who is calling that model at 3 p.m. on a Tuesday six months later, what data they are sending, and whether their role permits the request they just made. A model that scored well on adversarial testing in January can still receive a request in July from an authenticated user who should not have access to the data in that specific prompt. That gap is identity and authorization, not model behavior, and pre-deployment red-teaming runs on a testing cycle rather than on the live request itself, so it never sees the July request at all. Enterprises that treat a strong evaluation score as the whole security program tend to discover the gap during an incident, when the postmortem shows the model performed exactly as tested and the failure sat in who was allowed to ask it what.

Which one produces evidence for an EU AI Act Article 12 audit?

Article 12 of the EU AI Act requires automatic recording of events over the lifetime of a high-risk AI system, including timestamps, input data, and identification of the natural persons involved in a specific decision. That is a per-decision, per-request evidence requirement. DeepInspect produces exactly that record: a signed, tamper-evident log entry for every request, tied to the identity that made it, the policy that governed it, and the outcome, committed before the response returns to the calling application. CalypsoAI's evaluation and red-teaming reports, now part of F5 AI Guardrails, document how a model performed against adversarial testing at a point in time. That documentation is useful for a model-risk file and for showing a regulator the organization tests its models, but it stops short of reconstructing what happened on a specific request from a specific user on a specific date, which is what Article 12 actually asks for. The two kinds of evidence serve different sections of an audit response.

Does DeepInspect do its own model evaluation or red-teaming?

No. DeepInspect's architecture is built around enforcement rather than evaluation. DeepInspect sits inline on live traffic and makes a pass, redact, or block decision on each request based on identity, role, data classification, and policy, then records the outcome. Running adversarial prompts against a model in a test environment and producing a trust score for a candidate model before deployment sit outside that job. Organizations that need pre-deployment red-teaming and adversarial testing, the kind CalypsoAI built its platform around and F5 now ships as F5 AI Guardrails, need a product built for that job. DeepInspect's job starts once a model is approved and live: controlling and recording who is allowed to call it, with what data, under what policy, for every request that follows. Buyers sometimes expect a single platform to cover pre-deployment evaluation and runtime enforcement together. Today that usually means pairing a model-evaluation product with an enforcement layer, because the two operate on different cadences and answer different questions.

See also: CalypsoAI alternatives, compared, automated red-teaming, LLM guardrail bypass, policy as code, and DeepInspect vs Cisco AI Defense.