🛡️ Regex catches 0%, Meta's Prompt Guard 2 catches 1% of 629 realistic AgentDojo injection attacks when they're buried in tool output. Reproducible benchmark.
-
Updated
Oct 4, 2026 - Python
🛡️ Regex catches 0%, Meta's Prompt Guard 2 catches 1% of 629 realistic AgentDojo injection attacks when they're buried in tool output. Reproducible benchmark.
Open, privacy-bounded assurance for AI agents: containment provenance, identity passports, authorization twins, OTel evidence, CI gates, and OSCAL.
Action-graded severity scoring (L0-L6) for tool-using AI agents, computed from red-team execution traces.
Reproducible evaluation of deterministic sequence rules for AI-agent tool-call traces.
Reference-monitor security gateway for AI agents — YAML policies + information-flow control over MCP / OpenAI / Anthropic tool calls. Blocks indirect prompt-injection exfiltration by policy.
Evaluation harness that scores agent-authorization defenses on injection block rate and utility retention, using corpora derived from the AgentDojo benchmark, with label-sensitivity sweeps and a prompt-injection classifier adapter.
Authorization gateway for AI agents that act on data and money: the agent proposes, the gateway holds the keys and decides. Benchmarked on AgentDojo.
LLM agents read untrusted text, then call tools. We built a detector, measured it failing live (0/125), and replaced it with a pre-execution lock on money calls. Same 144 episodes, 0.00% attack success.
Benchmarking schema-valid false tool observations and defense baselines for tool-using LLM agents.
Formal runtime verification for tool-using LLM agents: MFOTL/MonPoly replayed offline on AgentDojo, STAC and R-Judge. Paper, MFOTL specifications, experiments, and the benchmark-readiness audit (CPSIoTSec 2026).
Deterministic, text-blind policy gate that stops prompt injection in tool-using LLM agents — measured on AgentDojo
Japanese localization of AgentDojo (prompt-injection benchmark for LLM agents). Unofficial community extension. 日本語化(非公式)
Security audit of LLM-based multi-agent systems with indirect prompt-injection PoCs and mitigations.
Information-flow control runtime for LLM agents: label data, track flows through tool calls, and block prompt-injection exfiltration at the sink. Python + TypeScript, CaMeL-style strict mode, OPA, OCSF/SARIF, AgentDojo-benchmarked.
Measure how well an agent tool-call authorization layer separates legitimate actions from injected ones, on AgentDojo ground truth. No agent runs, no model calls, seconds to run. Ships the controls that make a block rate meaningful.
Reproducibility artifact for auditing prompt-injection defense evaluation in AgentDojo.
A local-first research scaffold for evaluating models, agent harnesses, and complete agent products on realistic, stateful tasks.
Measuring prompt-injection defences against policy-conformant attacks
Evaluating provenance-gated tool calls as a prompt injection defense on AgentDojo. Reproducible runs, per-case analysis, and published results.
Prompt-injection defense layer and benchmark for tool-using LLM agents: 3 agents, 100 attacks, composable safeguards, DistilBERT detector, AgentDojo comparison
To associate your repository with the agentdojo topic, visit your repo's landing page and select "manage topics."