الانتقال إلى المحتوى الرئيسي
رجوع
Containment AttestationAgentic AI SafetyPrompt-InjectionEd25519LLM Red-TeamingAGPL-3.0

Project Simurgh

Project Simurgh is a provider-agnostic verifiable containment-attestation framework for agentic AI. It began as the defensive counterpart to The Invisible Window research, a metadata-only integrity layer for high-stakes sessions, and evolved into a general receipt for what a deployed AI agent was allowed to do after its first line of defense fails. Capability evaluations show what a model can do; Simurgh produces signed, offline-reproducible evidence of what a system let it do once it was connected to tools, files, context, and external systems. The classifier governs what a model may say; Simurgh attests what the agent was allowed to do, to a hostile reviewer, with no producer access.

01. المشكلة

Most AI defences optimise the first line: stopping the bad input. Anthropic's own Redeploying Fable 5 post concedes classifiers can be jailbroken, safety margins cost false positives, and full robustness is probably impossible. When that line fails, a prompt injection, a jailbreak, a tool-authority slip, there is neither a downstream layer that limits what the failure can do nor an evidence standard that lets an operator prove, to a skeptic, what actually happened. The threat model is a dishonest producer: an operator who wants to look contained.

02. نظرة عامة على الحل

  • Wraps an agent in four containment boundaries and seals each run into an Ed25519-signed, metadata-only evidence pack that re-derives byte-for-byte offline
  • Assumes the evidence producer lies — decision replay and emission-completeness checks let an outside reviewer confirm a real containment claim and falsify a dishonest one
  • Publishes a containment-utility Pareto frontier and a 12-rung reproduction ladder that an outsider recomputes with no model and no producer access
  • Positions as the defense-in-depth layer complementary to inline classifiers — the agent-authority cell a content classifier never touches — with a signed non-claim that it would not have caught the June 2026 content bypass
  • Preserves its published lineage: the Invisible Window integrity origin and the machine-checked absence-claims work (Banking Shield) both feed the current attestation contract

البناء

مكدس التقنيات

Node.js / JavaScript containment gatewayEd25519-signed, offline-reproducible evidence packsJS ↔ Python byte-parity verifierLean 4 machine-checked theoremsGitHub OIDC second provenance root
  • Four containment boundaries — input firewall, context-provenance guard, tool-invocation gate, output-leakage firewall — each emitting signed evidence
  • Dishonest-producer threat model: decision-replay and emission-completeness checks catch falsified or dropped evidence, verifiable offline by a hostile reviewer
  • Multi-class escape taxonomy (containment escape, verifier deception, out-of-scope, gate-boundary evasion) with reproducible adaptive red-team campaigns
  • Verifiable friction receipts prove an approval checkpoint preceded every protected authority crossing via a distinct-key pincer that defeats self-approval and backdating

الأمان

  • Producer-independent verifier trusts only the signer's public key and runs with no network — confirms a real containment claim and falsifies a dishonest one
  • Self-red-teamed the attestation core across eight attack classes (tamper, key-swap, canonical-laundering, digest-collision, cross-stage replay, self-proof mutation, policy drift); trust root held, two detector weaknesses versioned into detector-v2 and re-tested
  • Producer-independent witness cross-checks every signed receipt against an independent consequence oracle: zero false accusations, zero missed lies across the fixtures
  • Five machine-checked Lean theorems (fail-closed, friction precedence, same-key-fails, friction coverage, no-silent-exemption)
  • Honest non-claims are signed and explicit: not a jailbreak detector, complementary post-filter, and it would not have caught the June 2026 content-generation bypass itself

03. الإثبات والتحقق

الادعاءات المُتحقق منها

  • >Real Llama Guard 4 12B input classifier over a 180-case run-set: contained 138/138 malicious cases the classifier missed (120 downstream-injection cases an input-only classifier structurally cannot see, 18 direct-input misses); combined targeted attack-success 0/150; zero unsafe tool executions or exports
  • >Live agent (self-hosted Llama-3.3-70B) on AgentDojo's workspace suite, 140 pre-registered injection cases: authority gate cut targeted attack-success from 9/140 to 0/140 with benign utility held
  • >AgentDojo full four-suite deterministic run (v2.49.0): benign 97/97, unattacked-utility 949/949, attack-success 0/949
  • >Second independent provenance root signs the release verdict with GitHub's OIDC identity, not the developer's key; the workflow fails closed before signing if reality diverges from the committed verdict by a single byte
  • >3,057 automated tests; one-command offline reproduction of the signed release ladder; AGPL-3.0
  • >Published lineage on Zenodo (CC BY 4.0): the original integrity preprint (DOI 10.5281/zenodo.20374849), a voting-adjacent Phase C pilot, and the Banking Shield absence-claims prototype

الأوراق البحثية

3 أوراق

IEEE-format preprintCC BY 4.0 preprint

Project Simurgh: Privacy-Preserving Device Integrity Proofs for Capture-Resistant High-Stakes Sessions

12-page defensive follow-up to The Invisible Window, replacing visual surveillance with metadata-only integrity proofs.

Zenodo2026DOI 10.5281/zenodo.20374849
تحميل الورقة
Supplement preprintPhase C preprint

Privacy-Preserving Integrity Evidence for Student-Society Voting-Adjacent Workflows: A Phase C Pilot of Project Simurgh at Macquarie University

5-page voting-adjacent pilot reporting 31 consented sessions alongside a Macquarie student-society event, with ballot-choice exclusion, HMAC audit chaining, forbidden-field rejection, and 5/5 collection-closure gates.

Zenodo2026DOI 10.5281/zenodo.20549736
تحميل الورقة
Banking-adjacent preprintAuthor-prepared preprint

Banking Shield: Machine-Checked Absence Claims for Privacy-Sensitive AI Explanations

Fictional, non-bank research prototype that turns privacy and overclaim boundaries into machine-checkable evidence: a 46-name forbidden-field firewall whose rejections become audit events, a deterministic offline AI privacy firewall, and per-response privacy receipts anchored in per-session HMAC audit chains. At the evidence freeze all 417/417 unit tests, 43/43 end-to-end checks, and 27/27 security checks passed across three privacy audits and a no-egress static gate, with a formative five-tester dry run (30 sessions) recording zero sensitive values in evidence and 5/5 non-claim checklist comprehension.

Zenodo2026DOI 10.5281/zenodo.20675513
تحميل الورقة

استشهد بهذا العمل

Abedini, M. R. (2026). Project Simurgh: Privacy-Preserving Device Integrity Proofs for Capture-Resistant High-Stakes Sessions. Zenodo. https://doi.org/10.5281/zenodo.20374849