Cerberus blocks the lethal trifecta at the tool boundary — see the 525-run evidence set.

Solution

Any agent that can read your data, read the web and send a message is already exploitable

The attack needs no zero-day and no malware — three tool calls the agent was designed to make, and an instruction hidden in content it was told to read. Defending it means judging the action, not the text.

Autonomy removed the human checkpoint

Most security controls quietly assume a person reviews the consequential action. An agent does not wait, and the exfiltration looks exactly like the work it was asked to do.

  1. Prompt filtering guesses at intentText classification loses across languages, encodings and paraphrase. The outbound call is the thing worth judging.
  2. Model hardening is not a controlModel-level resistance shifts which payloads work; it does not remove the condition that makes them possible.
  3. Poisoned memory outlives the pageAn instruction written into memory keeps working long after the injected content is gone, and the transcript looks clean.

With Odingard, you can

Block the trifecta at the tool boundary

Cerberus correlates sensitive data access, untrusted content and outbound intent across the session, and holds the guarded call when they close.

Attack your own agents first

Argus red-teams autonomous systems the way an adversary would, so exposure is found by you and not disclosed to you.

Detect memory poisoning

Mimir seeds decoys that reveal when an agent's memory has been tampered with. It is in preview and labeled as such.

Replay any verdict

Each decision carries the signals and contamination graph that produced it, so an incident review is a query rather than an excavation.

What delivers it

Start with the MIT-licensed Cerberus core in your own environment; the rest layers on top.

Warden · By Odingard

Deploy it with the engineers who built it

Warden stands the runtime up alongside your team, reviews the boundaries you are enforcing, and red-teams the agents behind them.

Go deeper

Give agents real access without the risk

The core is MIT licensed and installs in your own environment. The evidence set behind it is published in full.