Cerberus blocks the lethal trifecta at the tool boundary — see the 525-run evidence set.

Warden · Agentic Assurance

The useful control on an agent is what it is allowed to do next

A focused engagement on the tool layer: every function your agents can call, the credentials behind it, and whether the restriction you believe is in place is actually enforced anywhere.

A documented restriction is not a restriction

Most agent guardrails live in the system prompt, which is a request rather than a control. The boundary that survives an adversary is the one evaluated at the point the action is taken.

  1. Prompt-level rules are advisoryAn instruction not to send data externally is negotiable by anyone who can influence the context window.
  2. Credentials outlive their purposeThe token issued for a prototype is usually still attached when the agent reaches production.
  3. Read and write are rarely separatedAn agent that needs to read a repository is commonly given a scope that also lets it write to one.

How the review runs

1

Enumerate the tool surface

Every callable function, its parameters, its credential and the systems it reaches.

2

Compare intent to enforcement

For each stated restriction, identify where it is actually evaluated — prompt, application code, gateway or nowhere.

3

Design the constraint

Least-privilege scopes, and runtime conditions for the calls that cannot simply be removed.

4

Verify

Attempt the actions that should now be impossible, and record the result.

What you hold at the end

  • A tool register: every function, its credential, its reach and its owner
  • A gap list of restrictions that exist in documentation but not in enforcement
  • Recommended scopes and runtime conditions, expressed as policy you can deploy
  • Verification evidence for the constraints applied during the engagement

What enforces it

Constrain the action, keep the capability

The goal is not a smaller agent. It is an agent that can still do the work and cannot do the damage.