Cerberus blocks the lethal trifecta at the tool boundary — see the 525-run evidence set.

ArgusAutonomous Red-Team Assessment

A finding only counts when the agent's behavior measurably changes — and the verdict is never issued by another model.

Argus is an autonomous red team for AI systems. It engages MCP servers, agent frameworks and live endpoints with generic attack techniques, then confirms each finding against a baseline transcript using deterministic detectors.

Argus Core is MIT-licensed and installable today. The full eleven-agent kit is delivered as a Warden red-team engagement rather than as installable software.

Your existing scanners were built for software that waits to be told what to do

SAST, DAST and network scanners test the infrastructure underneath an agent. None of them test the surface the agent itself introduces — the tools it can call, the content it will trust, and the instructions it will follow. Argus is complementary to those tools, not a replacement for them.

  1. 1 · The surface is the agent, not the portPrompt surfaces, tool schemas, MCP servers and memory are attack surface that no traditional scanner enumerates.
  2. 2 · Manual red teaming does not scaleEvery model update, prompt change and new tool re-opens the question. An assessment that takes weeks is stale before it ships.
  3. 3 · Most 'findings' are theaterA model echoing your payload back proves nothing. Without a behavioural delta you are handing a CISO a screenshot, not a vulnerability.

Attack like an adversary, report like an auditor

Argus attacks with generic techniques and records what actually happened, so a finding survives contact with the engineer who has to fix it.

One input, auto-routed

Point Argus at an MCP URL, a GitHub repository, an npm or PyPI package, a local script or a framework fixture, and it dispatches to the right target factory.

Sandboxed engagement

Untrusted target processes run under Docker with capabilities dropped, no network, a read-only filesystem, an unprivileged user and a process cap.

Provider failover built in

A single gateway dispatches to OpenAI, Anthropic and Gemini with a dead-provider blacklist, so an engagement survives a provider outage mid-run.

Deterministic verdicts

Mutation may use a model. Scoring never does. Every verdict comes from a regex, shape or counter detector that you can read and re-run.

Reproducible evidence bundles

Findings ship as signed, reproducible bundles with the exact trigger, source and reason — suitable for a vulnerability disclosure submission.

Runs in the pipeline

A GitHub Action, a pre-commit hook and a webhook receiver mean the same engagement can gate a merge rather than sit in a quarterly report.

The no-cheating contract

Every Argus finding is a claim an operator hands to someone's CISO or regulator, so the rules that keep findings honest are published rather than implied.

Zeromodels in the validation pathPayload mutation may use an LLM. Scoring never does — every verdict is emitted by a deterministic detector.
Baseline vs. post-attackis what defines a findingThe observation engine compares transcripts; a measurable, reproducible behavioural delta is required. An echoed payload is not a finding.
Notarget-aware branching or hardcoded canariesAgents attack with generic techniques only. Canary tokens are injected and detected at runtime, so a technique lands because it works.
2 of 11agents ship in the MIT corePI-01 (prompt injection) and EP-11 (environment pivoting) are public. The remaining nine run in a Warden engagement.

Methodology. These are engineering constraints in the Argus repository, not measured outcomes — the integrity contract is published in full at docs/NO_CHEATING.md and acceptance tests run real agents against real target fixtures rather than constructing findings by hand. Argus reports what a technique achieved against a specific target; it does not publish a portfolio-wide vulnerability rate, because that number would say more about the sample than about your estate.

Argus on GitHub

Core and engagement

Run the public core yourself, or have the team that wrote it run the full kit against your deployment.

Core

MIT licensed

The public CLI. Self-sufficient for a focused engagement you run yourself.

  • PI-01 prompt-injection and EP-11 environment-pivoting agents
  • Universal target dispatch and sandboxed runtime
  • Deterministic scoring and offline HTML reports
  • pip install argus-core

Warden engagement

The eleven-agent kit, the swarm correlation layer and a findings report, delivered by Odingard engineers.

  • Nine additional offensive agents across the assessment phases
  • Multi-agent slate execution and cross-agent correlation
  • Reproducible evidence bundles and remediation review
  • Scoped to your frameworks and reporting obligations
Request an engagement

Warden · By Odingard

Have it run against your estate

Warden runs the full kit, interprets the findings against your obligations, and re-tests once the fix lands.

Go deeper

Find it before someone else does

Install the core and test your own agents this afternoon, or have Warden run the full kit and hand you findings with the evidence attached.