Resources
The vocabulary, defined once
Agentic security has more terminology than agreement. These are the definitions this site uses, with attribution where the term came from somewhere else.
- Agent red teaming
Adversarial testing aimed at an agent's actions rather than its answers: can it be induced to call a tool it should not, reach data outside its scope, or persist an instruction for later. Model-level jailbreak testing does not cover it.
- Agentic AI
A system where a model does not just produce text but calls tools, reads and writes state, and takes actions with external effect. The security consequence is that a wrong output is no longer a wrong answer — it is an executed action.
- AI inventory
The register of AI systems an organization actually operates, including the ones adopted without going through procurement. Nearly every AI governance obligation is per system, so nothing else can be assessed until this exists.
- Audit-grade evidence
Evidence an auditor can verify independently rather than accept on the strength of who produced it — timestamped, tamper-evident, and traceable back to the control evaluation that generated it.
- Blast radius, B(p)
The set of records reachable from a poisoned write p through the dependency graph — everything that would have to be quarantined to contain it. Computing it is what turns “something was poisoned” into a bounded, actionable remediation.
- Continuous control monitoring
Evaluating security controls against live systems on a schedule instead of sampling them once a year for an audit. It changes compliance from a periodic reconstruction into a current statement of posture.
- Control crosswalk
A mapping from one internal control to the requirements it satisfies across several frameworks, so a single evaluation can be reported against all of them rather than repeated per audit.
- EU AI Act
Regulation (EU) 2024/1689, which classifies AI systems by risk and attaches obligations to each class. The obligations land per system, which is why an inventory is the prerequisite to answering anything under it.
- Governed autonomy
Letting agents act without supervision on each action, while the boundary they act within is defined, enforced and evidenced. The alternative positions are refusing autonomy or granting it unbounded, and neither survives contact with production.
- Indirect prompt injection
An instruction planted in content the agent reads — a web page, a document, a tool result — rather than typed by the user. The model has no reliable way to tell retrieved data from instruction, which is why filtering the output is the wrong control point.
- ISO/IEC 42001
The management-system standard for AI. Certification against it is issued by an accredited certification body following an audit — a consultancy can prepare you for that audit but cannot confer the certificate.
- Lethal trifecta
The co-occurrence of three properties in one agent context: access to private data, exposure to untrusted content, and the ability to communicate externally. Any two are usually survivable; all three together mean an injected instruction can read secrets and send them out.
Coined by Simon Willison, “The lethal trifecta for AI agents”, June 2025.
- Memory poisoning
Planting content in an agent's persistent memory so it influences a later, unrelated session. The delay is the point: the harmful action happens long after the injection, when nobody is looking at the session that caused it.
- NIST AI RMF
A voluntary risk-management framework organized around four functions: govern, map, measure and manage. NIST does not issue certification against it, so alignment is a claim about practice rather than a credential.
- Observe-only mode
A measurement condition in which a runtime guard records what it would have done without intervening. Detection figures gathered this way say nothing about whether an attack would have been blocked, and must not be reported as a blocking success rate.
- Provenance ledger
A record of where each piece of data in an agent's context came from and what was derived from it. It is what makes contamination traceable after the fact rather than merely detectable at the moment of use.
- Source independence
Whether agreeing sources are distinct observations or echoes of a single origin. Consensus between dependent sources is not corroboration, and the difference is measurable from provenance alone.
- Tool boundary
The set of tools an agent can actually invoke and the arguments it is allowed to pass them. Reviewing it is usually the cheapest risk reduction available, because most agent deployments grant far more reach than the task needs.
- Transitive taint propagation
Tracking contamination through derived writes: an honest agent that reads a poisoned value and writes a conclusion based on it has propagated the taint in good faith. Without transitivity, containment stops at the first hop and the poison launders itself.
Missing a term?
The glossary tracks the vocabulary this site actually uses. If something here is defined wrongly, or a term you needed is absent, tell us — it is a page we would rather have correct than complete.