Language models have already entered peer review, whether or not the community decided they should. Reviews are drafted with them, evaluated with them, and increasingly judged by metrics that are themselves model-based. That creates a circularity nobody has audited carefully.
We are building the instruments to audit it. That means benchmarks for verifying the claims a review actually makes about a paper, multi-faceted frameworks for scoring review quality that do not collapse into a single opaque number, and reliability analyses of the LLM-based metrics other people are already relying on.
The underlying question is not whether models are useful here. It is what evidence would let a community trust them, and whether that evidence currently exists.
Representative work
- Can LLMs Uphold Research Integrity? Evaluating the Role of LLMs in Peer Review Quality — WSDM 2026
- PeeriScope: A Multi-Faceted Framework for Evaluating Peer Review Quality — TheWebConf 2026
- Peerify: Benchmarking Peer-Review Claim Verification — EMNLP 2026
- Judging a Review by its Cover: A Reliability Analysis of LLM-based Peer Review Evaluation Metrics — CIKM 2026