How we calibrated Peer review so it is neither too harsh nor too lenient

How we calibrated Peer review so it is neither too harsh nor too lenient

Neutropic's Peer review puts your manuscript in front of five reviewers — journal fit, methodology, domain expert, broader perspectives and a devil's advocate — and an editor who issues one of four decisions: Accept, Minor Revision, Major Revision or Reject. A panel like that has an obvious way to fail. Reviewers are there to find problems, so left alone they find them everywhere, and every paper comes back Major Revision or Reject. A decision that never changes tells you nothing.

Our first real runs were all Reject. That was defensible — two were papers that had already been published, and four were a single-participant, two-minute HRV manuscript from our example project — but none of them had a known right answer. So we built a test where the answer is known, and ran the panel against it until it stopped leaning.

A corpus with known answers

35 manuscripts, all open access, blinded before review: journal name, DOI, received and accepted dates, citation lines, licences and affiliations were stripped, so "this is already published" is not a clue the reviewers can use.

  • A — published papers (8). Two recent research articles each from Applied Psychophysiology and Biofeedback, Frontiers in Psychology, PLOS ONE and Scientific Reports, reviewed for the journal that published them. Expected: Accept or Minor Revision, never Reject.
  • B — preprints that were later published (6). The first preprint version; the journal that eventually published it is the answer. Expected: Minor or Major, with the real journal among the recommendations.
  • C — retracted papers (4). Retracted for methodological or data reasons. Expected: Major or Reject, never Accept or Minor.
  • D — the wrong journal (3). Three of the A papers, submitted instead to Frontiers in Marine Science. Expected: the journal-fit score collapses while the other scores barely move.
  • E — injected flaws (14). Copies of A papers with one deliberate defect: the Methods section removed, p-values flipped while the test statistic stays the same, the ethics statement deleted, the abstract rewritten into a causal overclaim, or references swapped for DOIs that do not exist. Expected: the matching score drops and the right reviewer raises it.

Fourteen checks, with thresholds set in advance

Before the first run we wrote down what counts as a pass:

  • False rejection — published (A) papers rejected: at most 15%.
  • Leniency — retracted (C) papers given Accept or Minor: zero.
  • Decision order — average overall score: published papers above flawed variants above retracted papers.
  • Flaw sensitivity — E variants where the panel reacted to the injected defect: at least 80%.
  • Axis independence — D papers where fit dropped by 2 or more while every other score moved 0.5 or less: at least 5/6.
  • Reproducibility (two checks) — the same manuscript three times: the same decision at least 80% of the time, and a per-score standard deviation of 0.5 or less.
  • Reviewer spread — on published papers, no reviewer more than 1 point below the others (0.5 for the devil's advocate).
  • Major inflation — major issues per reviewer per published paper: at most 1.0.
  • Guard intervention — how often the decision guard has to override the editor: at most 20%.
  • Quote fabrication — reviewer comments dropped because the quoted text is not in the manuscript: at most 10%.
  • Language — the same papers reviewed in Korean and English: the same decision, overall scores within 0.5.
  • Effort — Basic vs Deep on four published papers: the same decision at least 75% of the time.
  • Journal recommendation — the journal that actually published a B paper appears in the recommendations: at least 50%.

The baseline was harsh

The first run (r1, Basic effort, English) confirmed the worry. Half of the published papers were rejected (4 of 8) and the other half got Major Revision. Each reviewer raised 3.27 major issues per published paper. Some things were already right — no retracted paper got off lightly, only 2% of comments had to be dropped for unverifiable quotes, and the real journal was among the recommendations for all six preprints — but the panel clearly leaned toward harshness, and it caught only half of the injected flaws.

Layer A across seven runs: Reject disappears after r1 and Minor Revision returns; major issues per reviewer fall from 3.27 to 0.90.
Layer A across seven runs: Reject disappears after r1 and Minor Revision returns; major issues per reviewer fall from 3.27 to 0.90.

What we changed, one lever at a time

Seven runs, each change followed by a re-measurement:

  • A calibration paragraph shared by every reviewer (r2, r5, r6). A major issue is one that invalidates a conclusion or blocks submission; a lapse in reporting conventions is minor. At most three majors, each with a one-sentence reason. Score anchors, with "an average publishable paper scores 3–4". Later additions: severity is what it takes to fix; a scope mismatch lowers only the fit score; the overall score is the average of the five axes unless the reviewer says why. After r2 no published paper was rejected and majors per reviewer fell to 1.23 — but now all eight got Major Revision.
  • Editor decision rules (r4). Without a corroborated defect, the editor goes to Minor; Reject is reserved for defects that cannot be fixed without new data, or for papers outside the journal's scope.
  • The devil's-advocate anchor (r5, r6). The devil's advocate is supposed to push back, not to sink every paper. It got its own score anchors and one rule: score the manuscript as written; an alternative explanation costs at most one point on evidence. Its average overall score on published papers went from 1.75 in r1 to 2.88 in r7.
  • Decision guards (r2 → r7). Code, not the prompt, limits which decisions the editor may pick. Accept is blocked only by a corroborated major issue — two reviewers agreeing, or a deterministic check such as a p-value that does not match its recomputed value or a reference DOI that is not registered. In r7 we closed a hole the r6 data exposed: when the average fit is 1.5 or lower, or most reviewers recommend Reject, only Major or Reject are allowed (in r6 an out-of-scope paper had slipped through at Minor).

One fix was not about wording at all. Flipped p-values went unnoticed in r3 and r4 because recomputing p-values from the reported test statistics only ran at Deep effort. Once it also ran at Basic (r5), both flipped-p variants were caught.

Progress was not a straight line. Flaw sensitivity dropped to 0.36 in r3 before it climbed, and axis independence swung between 0 and 2 of 3 from run to run.

Injected-flaw detection and scope-only drop across the seven runs, against the pass thresholds set before the first run.
Injected-flaw detection and scope-only drop across the seven runs, against the pass thresholds set before the first run.

The result: 13 of 14

The final run, r7 (Basic, English, all 35 manuscripts), together with the repeat, language and effort runs:

  • False rejection 0 of 8. Published papers: 3 Minor Revision, 5 Major Revision, no Reject. Majors per reviewer: 0.90.
  • Leniency zero. Retracted papers: 2 Major Revision, 2 Reject.
  • Decision order holds. Average overall score: published 3.17 (A) and 3.67 (B), flawed variants 2.97, retracted 2.15.
  • Flaw sensitivity 13 of 14 (0.93). The miss: an ethics statement deleted from one paper went unnoticed.
  • Wrong journal. All three were rejected; fit fell from 4.0–4.6 to 1.0.
  • Reproducibility. Every manuscript run three times (r6): the same decision 90% of the time, per-score standard deviation 0.08.
  • Language. Six papers in Korean and English: 6 of 6 the same decision, overall scores 0.2 apart.
  • Effort. Four papers at Basic and Deep: 4 of 4 the same decision.
  • Quotes and guards. 16 of 983 comments (2%) dropped for quotes not found in the manuscript; the guard never had to override the editor.
  • Recommendations. The journal that actually published each preprint was in the list 6 of 6 times.

The one miss: axis independence, 2 of 3

On one of the three wrong-journal papers, novelty moved by 0.6 alongside fit, against a threshold of 0.5. Reviewers tend to judge novelty relative to the journal's readership, which is not unreasonable. The decision was right (Reject) and fit dropped to 1.0, so we left it inside tolerance rather than adding yet another rule.

Why five published papers still got Major Revision

We could have pushed the published papers to Minor with a stricter guard. We chose not to. For four of those five, the guard allowed Minor and the editor still chose Major, because most reviewers pointed at real problems — in one paper a results-table row repeats the same values in columns that should differ, in another the omnibus statistics are missing. Published papers are not flawless, and a reviewer who spots a copy-pasted table row is doing the job. So we track the tendency with the false-rejection and major-inflation checks, and leave the decision itself to the editor.

Limits you should know

  • A small corpus. 35 manuscripts, eight of them published papers; one paper flipping moves a rate by 12.5 points. The papers are mostly psychology, physiology and HCI.
  • Synthetic flaws. Layer E injects one clean defect at a time. Real manuscripts carry several, and they are subtler.
  • Mostly Basic, mostly English. Korean and Deep were compared on six and four papers.
  • One reviewer model. Every run used the same model; another model may calibrate differently.
  • "Published" is not ground truth. It means one journal's editors said yes at one moment.

How to read your decision

A finished Peer review: the editorial decision badge and the editor's summary, next to the steps that produced it.
A finished Peer review: the editorial decision badge and the editor's summary, next to the steps that produced it.
  • The revision guide matters more than the badge. Start with the P0 items.
  • Major Revision on a good paper is normal. Five of eight published papers got it. Read it as "fix these before you submit", not "this paper is weak".
  • Take Reject seriously, and check fit first. In the final run Reject went to retracted and wrong-journal papers, never to a published one. A low fit score means the journal may be the problem, not the paper — look at the recommended journals.
  • Disagreement is information. When one reviewer scores far below the rest, read that reviewer's major issues and the editor's ruling on them.
  • Run it again after revising and compare the decisions.

Peer review is a pre-submission check. It can tell you what a careful referee is likely to ask and where your manuscript breaks a journal's rules; it cannot tell you whether a journal will accept you, and it does not replace the editor and referees who will actually read your work. Use it to walk into that review better prepared.