Project Simurgh: Privacy-Preserving Device Integrity Proofs for Capture-Resistant High-Stakes Sessions
12-page defensive follow-up to The Invisible Window, replacing visual surveillance with metadata-only integrity proofs.
Project Simurgh is a provider-agnostic verifiable containment-attestation framework for agentic AI. It began as the defensive counterpart to The Invisible Window research, a metadata-only integrity layer for high-stakes sessions, and evolved into a general receipt for what a deployed AI agent was allowed to do after its first line of defense fails. Capability evaluations show what a model can do; Simurgh produces signed, offline-reproducible evidence of what a system let it do once it was connected to tools, files, context, and external systems. The classifier governs what a model may say; Simurgh attests what the agent was allowed to do, to a hostile reviewer, with no producer access.
Most AI defences optimise the first line: stopping the bad input. Anthropic's own Redeploying Fable 5 post concedes classifiers can be jailbroken, safety margins cost false positives, and full robustness is probably impossible. When that line fails, a prompt injection, a jailbreak, a tool-authority slip, there is neither a downstream layer that limits what the failure can do nor an evidence standard that lets an operator prove, to a skeptic, what actually happened. The threat model is a dishonest producer: an operator who wants to look contained.
Stack Tecnológico
Afirmaciones Verificadas
3 Artículos
12-page defensive follow-up to The Invisible Window, replacing visual surveillance with metadata-only integrity proofs.
5-page voting-adjacent pilot reporting 31 consented sessions alongside a Macquarie student-society event, with ballot-choice exclusion, HMAC audit chaining, forbidden-field rejection, and 5/5 collection-closure gates.
Fictional, non-bank research prototype that turns privacy and overclaim boundaries into machine-checkable evidence: a 46-name forbidden-field firewall whose rejections become audit events, a deterministic offline AI privacy firewall, and per-response privacy receipts anchored in per-session HMAC audit chains. At the evidence freeze all 417/417 unit tests, 43/43 end-to-end checks, and 27/27 security checks passed across three privacy audits and a no-egress static gate, with a formative five-tester dry run (30 sessions) recording zero sensitive values in evidence and 5/5 non-claim checklist comprehension.
Citar este trabajo
Abedini, M. R. (2026). Project Simurgh: Privacy-Preserving Device Integrity Proofs for Capture-Resistant High-Stakes Sessions. Zenodo. https://doi.org/10.5281/zenodo.20374849