| Status | Design lesson — NOT a proposed feature |
| Parent epic | #255 — Kong et al. 2026 auto-research survey implications for ARS |
| Paper anchor | Kong et al. (2026), AI for Auto-Research: Roadmap & User Guide, arXiv:2605.18661 |
| Verified | 2026-06-08 — verified against the tracked repo (see Verification note) |
This is a recorded boundary, not a roadmap item. It exists so that future work and reviews apply the same line consistently.
POSITIONING.md — "the human decides at every gate" (Design philosophy);
ARS is a human-led copilot, not an autonomous paper-writing system.ARS and an auto-research system can look superficially alike. Both use multi-agent orchestration, a layered architecture, and simulated critique (ARS runs a five- reviewer panel plus a Devil's Advocate). A naive cross-walk could read these surface features as "ARS is most of the way to auto-research" — or, in the other direction, a reviewer could wave through a new agent on the grounds that "we already have many agents, one more is fine."
Both reads miss the actual line. The line is not how many agents there are or how layered the architecture is. It is who controls the next research-state transition — who gets to create, select, execute, or advance a research object of record, and whether the scholar saw and confirmed the relevant state first.
Any PR that adds or extends an agent, mode, or automation is reviewed against this question before merge:
Does this change let ARS create, select, execute, or advance a research object of record — hypothesis, RQ, evidence set, experiment result, claim, manuscript section, or dissemination artifact — without an explicit scholar-authored seed or scholar confirmation after inspecting the relevant state?
These are the specific surfaces where ARS could drift across the line if a change were not tested against the question above:
Three first-class architectural commitments, all already present:
Surface-level rejection ("multi-agent systems are auto-research, avoid them") would throw away orchestration, layering, and simulated critique — all of which are useful when the scholar owns every state transition. State-authority rejection ("no AI- owned transition of a research object of record") keeps those primitives while holding the line.
The reverse is equally true. Permitting a new agent on the grounds that "we already have many" silently re-introduces the risk: the question is never the count, it is whether the new agent can advance state the scholar has not seen and confirmed.
The Co-Scientist analysis recorded a related boundary from a different angle — who may rank and whether the user sees the full candidate set. Those docs and this one share the same underlying principle (the scholar must see and own the state the system acts on): Co-Scientist L1 (hidden ranking), L2 (feedback propagation), L3 (mechanism transfer), L4 (who may write / rank / route). The Kong L2 sharpens the specific advisory-vs-generation seam for research questions.
POSITIONING.md; the mandatory
checkpoints are real (academic-pipeline/SKILL.md). The Kong-derived features that
did ship — wording-pattern advisory (#257) and experiment provenance intake (#260)
— are advisory / provenance gates, not autonomous generation or execution layers. A
targeted review of the first-party implementation, agent prompts, and schemas found
no autonomous mechanism; the only wet_lab / slides / experiment references are
domain-evidence-profile labels or separate, non-ARS companion workflows, not
automation surfaces.