We are not asking you to take our word for it.

The code that enforces the rules is public. Your technical team can read it before you sign, and check that what follows is true.

97%of organisations hit by an AI-related incident had no access control on the tools involvedIBM, 600 organisations studied · B
73%of audited production AI deployments carry a prompt-injection weaknessCisco, 2026 · B
109 to 1machine identities per human identity, 97% of them carrying excessive privilegePalo Alto Networks, 2026 · B

The mechanism, in three sentences

01

Boundaries live outside the model

They are not written in the prompt. They are enforced by a layer the agent passes through to act and cannot reach. A successful prompt injection changes what the agent wants to do, not what it is allowed to do.

02

Every call is signed and verified

The agent signs the method, the path, the body, the timestamp, a one-use nonce and the digest of its own configuration. A replayed call, an altered body or a drifted configuration is refused.

03

Refused by default

What is not explicitly granted does not happen. The refusal carries a code, a reason, and the rule that decided it.

What your CISO will ask

Three columns: their question, what the platform does, and what they can go and verify.

The questionThe controlThe verification
How do we revoke a right from an agent going wrong?Revocation on the fly, with no redeploy and no restart.The next call is refused, and the refusal is in the trail.
What stops an agent exceeding a spend cap?A consumption ledger per agent, in tokens and in euros.Going over produces a refusal, not a warning.
How do we know what actually happened?Chained receipts, each action linked to the previous one.An after-the-fact edit breaks the chain and shows.
What if someone quietly changes a rule?Rule and right changes enter the same chain as the actions.When the rule changed and through which access is in the trail.
Where does our data go?The platform runs on your infrastructure. The agent's network egress is bounded.An undeclared domain is refused, not logged after the fact.
How do we prove all this to an auditor?Export of the full trail, actions and refusals.The file verifies without us and without the platform.

Sixteen modelled attacks. Sixteen refused, enforcement on.

Every row is a scenario attempting what this page says is impossible. The results file is in the repository, with the commit and the date of the run.

The attackWhat it targets
A1Forge an agent's identity tokenIdentityBlocked
A2Replay a token already spentIdentityBlocked
A3Replay a call signature with the same nonceIntegrityBlocked
A4Reuse a captured token on another endpointIntegrityBlocked
A5Alter a call's body after it was signedIntegrityBlocked
A6Present a call timestamped outside the windowIntegrityBlocked
A7Sign with another agent's key than the one claimedIdentityBlocked
A8Reach the admin surface with no valid keyRightsBlocked
A9Act with a revoked agent's tokenRightsBlocked
A10Ask, as a delegated agent, for more than the mandate grantedRightsBlocked
A11Spend the same payment authorisation twice, concurrentlyCapsBlocked
A12Fire 125 registrations against a limit of 60CapsBlocked
A13Call from an undeclared web originEgressBlocked
A14Edit a receipt already written to the trailProofBlocked
A15Guess a signature by timing the response, over 2,000 attemptsIdentityBlocked
A16Act with an agent configuration that drifted since enrolmentIntegrityBlocked
Last published runredteam/empirical-results.json · 8d0573a · 19 August 2026
The command that re-runs itnpm run redteam

Twenty-three milliseconds, database and receipts included.

The question in an architecture review is not whether the control is good, it is what it slows down. Measured at 100 concurrent calls, on a machine whose specification is published.

SauronID: policies, PostgreSQL, receipt chain
23 ms median
60 ms 95th percentile
Reference: signature verification only, no database, no policy, no receipt
33 ms median
59 ms 95th percentile

Machine and date : AMD Ryzen 7 7735HS, 16 cores, PostgreSQL · 20 August 2026 · redteam/benchmarks/results-summary.md

Lines to write to put an agent behind it
25on the agent side
0on the server side

What is in the public repository, and what you can run tonight

Open the repository
  • The gateway that enrols agents, evaluates policies, bounds network egress and produces receipts.
  • A console that lists agents, shows activity and refusals, and revokes a right.
  • A modelled attack suite, one scenario per rejection type, with its results and the enforcement mode used.
  • Python, TypeScript and Go clients, plus an MCP server.
  • The deployment files: Docker, Helm, and a native path without Docker.

The limits we write down ourselves, before you find them.

  • It is not a certification. No platform makes a company compliant with anything, and we will not pretend otherwise.
  • It replaces neither your network isolation nor your human identity management. The platform governs agent identity, not your employees'.
  • It does not make a model reliable. It bounds what its mistake can cost, which is a different promise.
  • The proofs make execution verifiable. That is not a confidentiality feature, and we will not call it one.
  • We have no published external audit, and we will not write "audited" before there is one. What we publish instead is more checkable: the threat model, the attack suite and the results of the last run, which your team can re-run.

Have your technical team read this page.

And meanwhile, tell us which process costs you the most. Both conversations can move in parallel.