Do LLM Reviewers Respect Scope? ScopeBench-PR: A Benchmark for Scope Fairness in Peer Review

Rishi Ahuja, Kumar Prateek, Simranjit Singh

ScopeBench-PR asks whether LLM reviewers evaluate a paper against the claims it actually makes. When research is explicitly regional or low-resource, do reviewers respect that scope, or penalize it for omitting English-language, global, or multilingual evaluation it never promised?

  • 1st place · Best Paper Presentation
  • Runner-up · Three Minute Thesis (3MT)

GlobalSouthAI · IJCAI–ECAI 2026

The question

LLMs are increasingly used to assist with drafting, summarizing, and auditing scientific reviews. Most fairness analyses focus on explicit hostility or direct penalties toward Global South and non-English research.

We examine a subtler failure mode: scope expansion. A reviewer may accept the underlying science while evaluating the work against a broader claim than the paper makes. Research on Hindi dialogue, for example, may be criticized for not including English or multilingual evaluation, while a regional deployment may be treated as incomplete because it is not global.

These are critiques of an expanded scope, not necessarily of the work as written.

How the benchmark is built

ScopeBench-PR keeps the scientific content of each paper fixed while varying two contextual factors: institutional prestige and language of study.

The benchmark contains 30 papers and 4,204 review runs, enabling paired comparisons of how the same work is evaluated under different framing. A 500-item human audit evaluates the weak-label detector, so its outputs should be interpreted as audit signals rather than definitive ground truth.

What we found

We observe penalties associated with regional language settings, although their magnitude varies substantially across models. Scope-aware prompting reduces these penalties for some reviewers but has little effect on others.

The most concerning result appears after rebuttal: a meta-review may acknowledge that an extra-scope demand was unfair while the paper’s score still fails to recover fully. We refer to this persistence as a sticky penalty.

I presented ScopeBench-PR at GlobalSouthAI, IJCAI–ECAI 2026. It received the Best Paper Presentation Award, and its Three Minute Thesis presentation placed runner-up.

More research