New Archestra's OpenAPPA Saturates Two Major Security Benchmarks with a 0% Attack Success Rate

Archestra released OpenAPPA, an open-source security engine designed to stop data exfiltration caused by prompt injection or model hallucination, according to InfoQ. The engine runs outside the agent's prompt and execution loop.
OpenAPPA implements an Agentic Permissions Policy Algebra (APPA), described in a paper by Arseny Kravchenko, Vadim Liventsev, Innokentii Konstantinov, Ildar Iskhakov, and Matvey Kukuy. Policies are configured in a single appa.toml file covering data sources, audiences, trust levels, and authorities. Each tool contract defines requires, delta, and effects attributes. Labels compose monotonically using lattice algebra, so they can only become more restrictive.
The team reports zero successful attacks on the Bench-Corp benchmark (20 multi-step enterprise workflows) and AgentThreatBench, with an 89% task completion rate. By comparison, Claude Code's native auto mode yielded a 10% attack success rate with 90% completion, and Microsoft FIDES permitted 31% of attacks and completed 41% of tasks. Bench-Corp and AgentThreatBench test explicit policy breaches covering sensitive data sharing, prompt injection, approval and ordering, and tenant isolation. AgentThreatBench operationalizes the OWASP Top 10 for Agentic Applications (2026) and was recently merged into the UK AI Safety Institute's inspect_evals repository.
OpenAPPA adds recovery semantics: sanitizers may edit payloads to expand permitted audiences, authorities route requests to human operators or internal verification APIs, and disposable child branches isolate reads of untrusted data in transient subagent branches. The team reports that disabling remedy plans in ablation experiments reduced task completion to 35.0%.
The paper, APPA: Recoverable Information-Flow Control for Real-World LLM Agents, is published on arXiv. OpenAPPA is currently a preview.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Archestra released OpenAPPA, an open-source security engine designed to stop data exfiltration caused by prompt injection or model hallucination. The team reports zero successful attacks when running security benchmarks Bench-Corp (20 multi-step enterprise workflows) and AgentThreatBench, versus 10% for Claude Code’s auto mode and 31% for Microsoft FIDES. By Bruno Couriol