Date: 2026-07-11
Scope: the runtime LLM judge of the ## Wording-Pattern Advisory section in
deep-research/agents/socratic_mentor_agent.md (the academic-paper twin differs
only in scenario prose; substance identical, verified by diff).
Deliverables: evals/heldout/rq_framing_offlist/ (set + measurement JSON + protocol README), this report.
Provenance chain: PR #468 review thread (@brycewang-stanford) → issue #501 → PR #503 (Part 1 prompt edit) → this measurement (Part 2).
Issue #501: does the runtime judge miss real AI-typical RQ phrasings that sit outside the twenty WP01–WP20 surface forms? If the miss rate is low, close with no further action; if high, the held-out set becomes the acceptance test for any future change.
gpt-5.6-sol via Codex CLI 0.144.1); the
shell items are filtered against the shipped regex detector and the four
in-prompt examples (four negatives intentionally carry listed surface substrings
as hard-negative material — see the set README),
dual-annotated (generator + maintainer; 8 borderline/disagreement items dropped),
elicited-rewrite labels inherited by construction under a no-new-specifics
constraint. English-only per the #468 language/model-drift caveat.claude-sonnet-5 sub-agents, given only the verbatim advisory
section (variant under test) + 6 shuffled items each, no labels, no repo access.evals/gold/rq_framing_patterns/manifest.yaml:
FNR < 0.30, FPR < 0.20.| variant | overall miss | family_variant (n=23) | off_list (n=9) | false-fire (n=16) |
|---|---|---|---|---|
| pre-#503 | 12/32 = 0.375 | 6/23 = 0.261 | 6/9 = 0.667 | 0/16 = 0.000 |
| post-#503 rep1 | 12/32 = 0.375 | 5/23 = 0.217 | 7/9 = 0.778 | 0/16 = 0.000 |
| post-#503 rep2 | 11/32 = 0.344 | 4/23 = 0.174 | 7/9 = 0.778 | 0/16 = 0.000 |
Replicate stability: the two post-#503 runs disagree on exactly one item
(nat-049); the seven off-list misses are identical across both replicates
(nat-044, ti-002, ti-004, ti-008, ti-010, ti-012, ti-013).
academic-paper twin covers thesis sentences and chapter framings,
and users paste working titles as research directions), and the interrogative
off-list sample here is too small (n=1, nat-044, missed in all runs) to
support a separate question-form claim.evals/heldout/rq_framing_offlist/judge_reasoning_excerpts.md) argues the
title-shell misses as "names a specific mechanism/population", applied to
generic topical noun pairs like "cybersecurity training → behavior". Not every
judgment produced prose (some agents returned bare JSON), so this is the best
available evidence, not a per-item demonstration. On this evidence, a future
fix should sharpen the exemption (e.g., a named instrument/scale/site test)
rather than extend the pattern table — to be re-tested against this set.ti-001
"Rethinking …") was caught post but not pre. The design cannot support a causal
claim: pre-#503 has one run, the difference is 1–2 items, and that is the same
magnitude as the observed between-replicate flip (nat-049). What the data do
support: post-#503 family-variant misses sat under the 0.30 line in both
replicates, and the decorated-title shape was unaffected (none of the
paragraph's four examples resemble it).claude-sonnet-5), single generator model (gpt-5.6-sol),
English-only, 2026-07 snapshot. Both AI-typical phrasing and judge behavior are
model- and time-specific (PR #468 discussion); re-run the protocol rather than
reusing numbers.gpt-5.6-sol + claude-fable-5, the
maintainer-session agent) plus construction-inheritance with per-rewrite
noun-swap re-verification; documented drops included; borderline items were
removed rather than adjudicated. Annotator 2 shares a model family with the
measured judge — the cross-family generator annotation and the mechanical regex
filter are the independence anchors.