CSPaper benchmark update
700 Papers, 8 Models
OpenAI, Google, and Z.AI's Flagship LLMs Compared
CSPaper
A part of Scholar7 AB · July 2026
The short version
- This snapshot reflects results obtained on 20 July 2026: 8 models, 38 public venue and track rows, and 700 expert-calibrated papers.
- GLM-5v Turbo has the best score fit among the featured models, sits almost exactly at zero bias, and earns the highest share of positive user ratings.
- Gemini 3.1 Pro is the strongest featured ranker and leans generous. It may be the reviewer most inclined to see the glass as half full.
- GPT-5.6 Terra gives the most detailed feedback but scores more strictly. Reviewer #2 has entered the chat, with footnotes.
- These dynamic results feed the model recommendations shown to users in Step 3 of the review workflow.

The result in 30 seconds
These statistics are a snapshot of the benchmark dataset and results obtained on 20 July 2026. There is no single winner, which is exactly why the benchmark uses more than one metric. The cleanest headline belongs to Z.AI: GLM-5v Turbo records the best NMAE of the featured models, the smallest mean bias by a wide margin, and a strong SRC. It also receives the highest share of 4 and 5 star user ratings.
Gemini 3.1 Pro remains the best ranker in this group. GPT-5.6 Terra is not the numerical leader here. It produces the longest reviews and lands in a competitive range, but its scores are less accurate than GLM's and its paper ordering is less consistent than Gemini's. A new model does not need a gold medal in every column to be useful. It needs a clear profile, and now Terra has one.
NMAE
Score error
NME
Score bias
SRC
Ranking agreement
Four metrics, four different questions
Let N be the number of benchmark papers, Sgt the expert-calibrated score, and Spred the CSPaper prediction. The score range is venue specific, so normalization lets an ICLR scale and an ACL scale live in the same comparison.
NMAE: score accuracy
Lower is betterThe average absolute score gap, divided by the venue's score range. A value of 0.165 means the prediction misses by 16.5% of the available rating scale on average. This is our first headline metric because it answers the simplest question: how close was the score?
NME: score bias
Zero is idealThe signed version of NMAE. Positive means the model scores too generously; negative means it scores too harshly. Errors in opposite directions can cancel, so NME diagnoses calibration but never replaces NMAE.
SRC: ranking quality
Higher is betterSpearman rank correlation compares the predicted paper order with the expert-calibrated order. Here, dᵢ is the difference between the two ranks for paper i. A score of 1 is perfect ordering, 0 means no rank correlation, and -1 means the order is reversed. This is our other headline metric because review systems must separate stronger papers from weaker ones, not just hover near the average score.
AWC: review length
Descriptive, not a quality scoreAverage Word Count measures how much text the full review contains. Longer can mean more justification, but it can also mean more repetition. We report it because useful feedback needs room to explain itself, while treating it as the least important of the four metrics.
NMAE and SRC carry the most weight in our reading. NME tells us whether errors lean high or low. AWC tells us how much was written, not whether it was right. Four gauges are better than one, but none of them can read a reviewer's mind. More on that shortly.
Benchmark results become model recommendations
Users do not have to memorize four metrics before starting a review. In Step 3, CSPaper recommends models using the latest benchmark results for the selected venue. Labels such as low error, high correlation, concise, verbose, over-estimator, and under-estimator turn the live measurements into practical choices. As the benchmark changes, the recommendations can change with it.

The featured three, side by side
| Model | NMAE ↓ | NME → 0 | SRC ↑ | AWC | 4 or 5 stars |
|---|---|---|---|---|---|
| GLM-5v TurboZ.AI | 0.165best | -0.000best | 0.621 | 17,137 | 61%best |
| Gemini 3.1 ProGoogle | 0.178 | +0.047 | 0.658best | 8,787 | 51% |
| GPT-5.6 TerraOpenAI | 0.193 | -0.054 | 0.559 | 17,821 | 54% |
GLM's case is unusually tidy. Its NMAE is 7% lower than Gemini 3.1 Pro and 14% lower than GPT-5.6 Terra in this public-row summary. Its mean NME is -0.0004, effectively centered on zero. It does not win SRC, but 0.621 is still a strong ranking result. Put plainly: GLM gets scores close, avoids a clear optimistic or pessimistic lean, and usually orders papers well.
That is excellent work from Z.AI. The model has earned a victory lap.
Gemini 3.1 Pro tells a different story. It has the best SRC at 0.658, so it is the strongest at preserving relative paper order. Its reviews are much shorter, roughly half the length of GLM's and Terra's, and its positive NME suggests a mild optimistic tilt. Gemini seems to be the nice reviewer in the room: good at sorting the stack, and a little more willing to round up.
GPT-5.6 Terra is the wordiest of the three and is mildly conservative on average. At -0.054 NME, it behaves a bit like Reviewer #2: strict score, long report, and one more requested ablation. Its NMAE and SRC do not lead this round, so we will keep testing whether that extra detail helps authors act on the feedback.
Who wins most often?
Here is the ten-second view. For each venue or track row, one win means the model posted the lowest NMAE, the NME closest to zero, or the highest SRC. The headline total combines NMAE and SRC, our two most important metrics. GLM leads with 20, followed by Gemini 3.1 Pro with 16 and Terra with 15. Ties count for every model sharing the best displayed value.
| Model | NMAE wins | NME wins | SRC wins | Headline winsNMAE + SRC |
|---|---|---|---|---|
| GLM-5v Turbo | 7 | 7 | 13 | 20 |
| Gemini 3.1 Pro | 8 | 7 | 8 | 16 |
| GPT-5.6 Terra | 8 | 7 | 7 | 15 |
| GPT-5.4 | 5 | 5 | 6 | 11 |
| Gemini 2.5 Pro | 5 | 7 | 4 | 9 |
| GPT-5.1 | 3 | 6 | 3 | 6 |
| GPT-5 | 4 | 2 | 1 | 5 |
| Gemini 3 Flash | 1 | 3 | 3 | 4 |
This keeps all eight supported models visible without repeating the detailed values already published on the benchmark page. Coverage still matters: the featured models and GPT-5.4 cover all 38 rows, while older models cover 29 to 31. Win counts for those older models therefore come from fewer opportunities.
Scores are not the whole review
The four benchmark metrics lean heavily on scores and ordering. Review quality is broader. A useful critique should be correct, specific, fair, and possible to act on. After a user receives the review report for their paper, CSPaper asks them to rate it from 1 to 5 stars. This post-delivery feedback channel was introduced as part of the operating system described in Computational Research Assessment, and it gives us a qualitative signal that score metrics alone cannot provide.
User signal
Share of reviews rated 4 or 5 stars
GLM's 61% is the best of the three. That agreement between the main score metric and the user signal is encouraging. It is not proof that GLM wins every notion of review quality, but it is a good reason to keep taking fast-moving open-model families seriously. We plan to explore support for more of them as their vertical performance improves.
What this benchmark can and cannot say
A 700-paper test is large enough to expose patterns that a few cherry-picked examples cannot. It also covers 38 public venue and track configurations, which matters because a good review is rubric dependent. The original CSPaper system paper began with a 100-paper model-selection benchmark. The current corpus is seven times larger, and the public matrix makes the venue-level variation visible.
Still, the benchmark has limits. Human scores are not scientific truth. NMAE can reward safe predictions near the middle. SRC says little about whether a critique found the right flaw. NME can hide large errors that cancel. AWC cannot tell insight from padding. User stars can reflect tone, speed, or expectations as much as correctness.
This is why CSPaper's research direction is verification first. Our work on verification-first AI argues that review tools should increase the amount of science we can check, not merely imitate historical judgments. The Computational Research Assessment agenda likewise calls for evidence-linked critiques, explicit escalation when assessors disagree, open benchmarks, and red tests for gaming and bias.
Every new agent has to earn its place
When we build a conference, journal, or workshop agent, we do more than attach a rubric to a prompt. We bake official venue criteria, score-calibration rules, evidence and tool policies, reasoning policies, review templates, expert examples, and known failure cases into a versioned agent harness. Candidates are regenerated offline and evaluated against fixed validation packs before deployment. That controlled factory loop is documented in Self-Improving Agent Factories for Paper Review.
The reason is simple: scientific feedback is a vertical task. Progress on general reasoning does not automatically transfer to a venue's evidence standards. The position paper AI Should Verify, Not Judge, Scientific Work makes the boundary clear: these tools should help authors verify claims and improve manuscripts, not replace human peer review. We benchmark the exact job, keep the test set stable, and surface the latest results so users understand why a model is recommended and when that recommendation changes.
This also fits the broader bottleneck described by Epistemic Throughput: generation is cheap, but careful verification remains scarce. Better screening can help focus that scarce attention. It cannot make verification optional.
The CSPaper Review system paper at INLG 2025 documented the first version of this loop: venue-specific rubrics, score-conditioned review candidates, selection, synthesis, and calibration. The benchmark you see today is the same habit grown up: more papers, more venues, more metrics, more models, and public results.
The takeaway
Pick models for the review job, not for the logo on the box.
GLM-5v Turbo is the standout addition in this round. Gemini 3.1 Pro remains the strongest ranker of the featured three. GPT-5.6 Terra brings more detail, but not the best score fit. The benchmark is dynamic: we will keep refreshing the evidence, expanding providers and model families, and adding fast-moving open-source LLMs when they show strong vertical performance for scientific review.
References
- Lele Cao, Lei You, Kai Xie, Weiping Ding, Yong Du, Sven Salmonsson, Yumin Zhou, and Vilhelm von Ehrenheim. CSPaper Review: Fast, Rubric-Faithful Conference Feedback. INLG 2025 System Demonstrations.
- Lei You, Lele Cao, and Iryna Gurevych. Preventing the Collapse of Peer Review Requires Verification-First AI. ICML 2026 AI for Science Workshop.
- Lele Cao, Lei You, Kai Xie, Weiping Ding, Yong Du, Sven Salmonsson, Yumin Zhou, and Vilhelm von Ehrenheim. Adopt Machine-Human Collaboration Peer-Review through Computational Research Assessment. ICML 2026 AI for Science Workshop.
- Lele Cao. Self-Improving Agent Factories for Paper Review: When Constrained Beats Open-Ended. KDD 2026 Workshop on SciSoc: Agents & LLMs.
- Prabhant Singh, Thanh Gia Hieu Khuong, Vlasta Sikimić, Benedictus Kent Rachmat, Kola Ayonrinde, Ihsan Ullah, Christina Lioma, Luis Oala, Kevin Qinghong Lin, Lele Cao, Neil F. Abernethy, Hilde Weerts, and Joaquin Vanschoren. Position: AI Should Verify, Not Judge, Scientific Work. ICML 2026 AI for Science Workshop.
- Lei You. Epistemic Throughput: Fundamental Limits of Attention-Constrained Inference. 2026.