Research From this edition 2 min read
The reviewer needs a reviewer.
HalluPeer studies unsupported statements in scientific peer reviews, including the difficulty of separating an error from fair criticism.

HalluPeer pairs papers with reviews and versions containing deliberately inserted errors. The authors describe experiments covering 12,000 papers and 38,000 reviews, and report that existing detectors struggle to distinguish fabricated claims from legitimate criticism. The arXiv page lists acceptance to EMNLP Findings 2026.
The limits matter: its constructed errors are synthetic, and the source material comes from computer-science conferences on OpenReview. It does not establish the error rate of all human reviewers or all AI assistants.
Our takeaway: when using AI to assess a report, require a passage supporting each factual criticism. Separate what the document actually says from what the reviewer would have preferred it to say. Confident disapproval has never been in short supply.
Read the original accounts
- Lead: Fresh Memory, Stale Plans · submitted 3 September 2026 ↗
- Market: KC-Bench · submitted 3 September 2026 ↗
- Watch: HalluPeer · submitted 3 September 2026 ↗
Published in our 05/09/2026 edition. Source dates are shown above.
The AI Street Journal · Free to read