Why every analysis deserves a second opinion — from an agent

Why every analysis deserves a second opinion — from an agent

A single model that both produces an analysis and grades its own work has an obvious conflict of interest. Neutropic separates the two roles: a generation agent runs your analysis, and an independent review agent inspects it before you ever rely on the result. Here is what that looks like on a real manuscript.

What the review agent checks

The review agent re-derives the analysis from the study design and the data, then looks for the failure modes that most often survive into published papers:

  • Assumption violations — normality, homogeneity of variance, sphericity, independence.
  • Multiple-comparison risk that the analysis plan never corrected for.
  • Statistical power and sample size against the claimed effect.
  • Analytical inconsistencies — a mediation model whose paths don't match the reported table, a figure computed from a different subset than the text describes, or a number in the prose that the dataset does not produce.
The generation side first: the statistics turn runs its code in a sandbox and reports effect sizes with confidence intervals, achieved power and design sensitivity.
The generation side first: the statistics turn runs its code in a sandbox and reports effect sizes with confidence intervals, achieved power and design sensitivity.

Disagreement is the feature

When the two agents disagree, Neutropic doesn't silently pick a winner. The disagreement is surfaced to you with both positions and their evidence. In our internal studies, these flagged disagreements were the single best predictor of issues a human methodologist would also raise.

The reviewer at work on the manuscript: reading the full text, re-running the statistics on the dataset, verifying citations one by one.
The reviewer at work on the manuscript: reading the full text, re-running the statistics on the dataset, verifying citations one by one.

What it caught in the wild

Early users have shared cases where the review agent flagged a repeated-measures ANOVA run on data with a broken sphericity assumption, a moderation analysis whose interaction term was computed on uncentered variables, and a study where the reported power analysis used the wrong effect-size family. In the session shown here it found that the demographic means quoted in the Methods did not match the trial dataset — while the inferential test in the same paragraph, re-run on the data, matched to the third decimal.

Issue 1 (MAJOR): the location in the manuscript, the problem, and the evidence — the reviewer's own re-computation against the reported numbers.
Issue 1 (MAJOR): the location in the manuscript, the problem, and the evidence — the reviewer's own re-computation against the reported numbers.
None of these were exotic. They were ordinary mistakes under deadline pressure — exactly the kind a second reader exists to catch.

Rigor as a default, not a service

You don't schedule a review or remember to ask for one. On Deep effort every analysis in Neutropic passes through review before its results are presented as findings — the same way version control made "backing up code" stop being a decision. And when you want a referee's reading of the whole paper, /review is one message away.