← Về thư mục
Review Criteria Framework — Structured Review Criteria Framework
This document defines universal criteria for academic paper review and type-specific criteria differentiated by paper type. All reviewer agents share this framework.
1. Universal Review Dimensions
Seven core dimensions applicable to all paper types.
Numerical weights and the aggregation formula are single-sourced in quality_rubrics.md — this document deliberately does not restate them (the two copies drifted once; issue #396). The level tables below define qualitative descriptors only.
Dimension 1: Originality
| Level |
Score |
Description |
| Outstanding |
5 |
Proposes entirely new theory/method/evidence that could change the field's direction |
| Strong |
4 |
Has clear new insights or novel combinations, fills a specific research gap |
| Adequate |
3 |
Incremental contribution, reasonable extension of existing knowledge |
| Weak |
2 |
Highly overlapping with existing literature, new contribution unclear |
| None |
1 |
Essentially repeats what is already known |
Dimension 2: Methodological Rigor
| Level |
Score |
Description |
| Outstanding |
5 |
Impeccable research design, innovative methods executed flawlessly |
| Strong |
4 |
Sound design, appropriate methods, minor room for improvement in execution |
| Adequate |
3 |
Methods basically acceptable, but with some design or execution limitations |
| Weak |
2 |
Methods have significant flaws affecting the credibility of conclusions |
| Unacceptable |
1 |
Methods fundamentally unsuitable for answering the research question, or contain serious errors |
Dimension 3: Evidence Sufficiency
| Level |
Score |
Description |
| Outstanding |
5 |
Rich, diverse, and persuasive evidence that exceeds expectations |
| Strong |
4 |
Evidence sufficiently supports all major arguments |
| Adequate |
3 |
Most arguments supported by evidence, a few need supplementation |
| Weak |
2 |
Key arguments lack sufficient evidence |
| Unacceptable |
1 |
Serious disconnect between arguments and evidence |
Dimension 4: Argument Coherence
| Level |
Score |
Description |
| Outstanding |
5 |
Clear arguments, rigorous logic, elegant structure |
| Strong |
4 |
Smooth argumentation, occasional minor logical leaps |
| Adequate |
3 |
Basically coherent, but some inter-paragraph connections are unclear |
| Weak |
2 |
Multiple logical breaks, readers have difficulty following the argument |
| Unacceptable |
1 |
Confused argumentation, core claims cannot be identified |
Dimension 5: Writing Quality
| Level |
Score |
Description |
| Outstanding |
5 |
Precise and fluent academic English/Chinese, a model of scholarly writing |
| Strong |
4 |
Clear language, occasional minor imperfections that don't affect understanding |
| Adequate |
3 |
Generally readable, with some grammar or word choice issues |
| Weak |
2 |
Frequent language issues that affect understanding |
| Unacceptable |
1 |
Language quality does not meet reviewable standards |
Dimension 6: Literature Integration
| Level |
Score |
Description |
| Outstanding |
5 |
Comprehensive, contemporary, critically integrated literature with a compelling research gap argument |
| Strong |
4 |
Covers major literature, with good integration and positioning |
| Adequate |
3 |
Basic coverage, but with omissions or insufficient integration |
| Weak |
2 |
Literature is outdated, incomplete, or merely enumerated |
| Unacceptable |
1 |
Seriously insufficient literature review or irrelevant to the topic |
Dimension 7: Significance & Impact
| Level |
Score |
Description |
| Outstanding |
5 |
Could change policy, practice, or theoretical direction |
| Strong |
4 |
Clear impact on a specific field or practice |
| Adequate |
3 |
Has some academic or practical value |
| Weak |
2 |
Limited scope of impact, mainly academic interest |
| Marginal |
1 |
Difficult to see the significance of the research |
2. Paper Type-Specific Criteria
2.1 Empirical Research
Beyond universal dimensions, specifically focus on:
| Additional Dimension |
Review Focus |
| Research hypothesis clarity |
Are hypotheses testable and consistent with theory |
| Variable operational definitions |
Are independent/dependent/control variable definitions precise |
| Internal validity |
Are confounding variables controlled |
| External validity |
Generalizability of results |
| Statistical reporting completeness |
Effect sizes, confidence intervals, assumption testing |
| Conclusion conservatism |
Do conclusions exceed what the data supports |
2.2 Theoretical/Conceptual Paper
| Additional Dimension |
Review Focus |
| Conceptual definition precision |
Are core concepts clearly delineated |
| Argument logic structure |
Is the premise -> inference -> conclusion logic chain complete |
| Counterargument handling |
Are possible opposing viewpoints considered and addressed |
| Theoretical novelty |
Does it truly advance theoretical development |
| Testability |
Can the theory generate testable propositions |
| Additional Dimension |
Review Focus |
| Search strategy |
Is it comprehensive and reproducible (PRISMA compliance) |
| Inclusion/exclusion criteria |
Are criteria clear, reasonable, and consistently applied |
| Bias risk assessment |
Is bias risk of included studies assessed |
| Heterogeneity handling |
Is statistical and conceptual heterogeneity appropriately handled |
| Synthesis method |
Goes beyond simple vote counting to achieve critical synthesis |
| Publication Bias |
Is publication bias assessed and discussed |
2.4 Case Study
| Additional Dimension |
Review Focus |
| Case selection justification |
Why was this case chosen? What does it represent? |
| Theoretical vs convenience sampling |
Is case selection theoretically grounded |
| Triangulation |
Are multiple data sources used |
| Context description thickness |
Is thick description sufficient |
| Analysis transferability |
Are analysis results transferable to other contexts |
| Researcher reflexivity |
Is the researcher's relationship with the case reflected upon |
2.5 Policy Analysis / Policy Brief
| Additional Dimension |
Review Focus |
| Policy problem definition |
Is the problem clearly defined and evidence-supported |
| Stakeholder analysis |
Are key stakeholders identified |
| Policy option analysis |
Are multiple options proposed and compared |
| Feasibility assessment |
Are policy recommendations practically feasible |
| Evidence quality |
Are policy recommendations based on reliable evidence |
| Unintended consequences |
Are unintended policy impacts considered |
3. Common Review Pitfalls
Biases Reviewers Should Avoid
| Pitfall |
Description |
How to Avoid |
| Hypercriticism |
Overblowing minor issues, ignoring the paper's overall contribution |
Affirm strengths first, then point out issues; distinguish major from minor |
| Confirmation Bias |
Only finding evidence supporting pre-existing views |
Deliberately seek the paper's merits and counterexamples to your own views |
| Preference Projection |
Requiring authors to use "my method" rather than evaluating "the author's method" |
Ask "can this method answer the question" rather than "what would I do" |
| Paradigm Bias |
Using quantitative standards to judge qualitative research (or vice versa) |
Use evaluation criteria matching the paper's research paradigm |
| Prestige Bias |
Relaxing standards because of the author's institution or past achievements |
Focus on the quality of the paper itself |
| Novelty Bias |
Only valuing novel research, undervaluing replication studies |
Acknowledge the important role of replication in science |
| Length Bias |
Long paper = good paper, short paper = sloppy |
Evaluate content density, not page count |
| Language Discrimination |
Undervaluing research quality due to non-native language imperfections |
Distinguish "language needs polishing" from "research quality is poor" |
Principles of Constructive Feedback
- Specific, not vague: "The causal inference in Section 3, paragraph 2 lacks control variables" is better than "methodology has problems"
- Problem + reason + suggestion: Every criticism should include "what," "why," and "how to fix"
- Distinguish required from suggested: Which changes are mandatory, which are "nice to have"
- Acknowledge uncertainty: "I'm not sure whether this analysis accounts for X" is more accurate than "the author ignored X"
- Respect the author: Even if paper quality is poor, the author still invested time and effort
4. Scoring Aggregation
The aggregation weights, the weighted-score formula, and the score-to-decision mapping are defined in quality_rubrics.md (single source — do not restate the numbers here). Two structural facts from that source:
- Five dimensions carry aggregation weight: Originality, Methodological Rigor, Evidence Sufficiency, Argument Coherence, Writing Quality.
- Literature Integration and Significance & Impact are reviewer-specific optional dimensions (R2 / R3 focus): scored and reported separately, factored into the editorial synthesis narrative, but not part of the numerical aggregate.
Important reminder: Scores are only reference. The final decision also needs to consider:
- Whether any single dimension sits at the bottom descriptor (e.g., methodology at "Unacceptable"), which may lead to Reject even if the overall score is passable
- Specific content of reviewer comments is more important than numbers
- Special considerations of the journal (special issue, field development needs, etc.)