CSPaper benchmark update

700 Papers, 8 Models OpenAI, Google, and Z.AI's Flagship LLMs Compared

C

CSPaper

A part of Scholar7 AB · July 2026

The short version

  • This snapshot reflects results obtained on 20 July 2026: 8 models, 38 public venue and track rows, and 700 expert-calibrated papers.
  • GLM-5v Turbo has the best score fit among the featured models, sits almost exactly at zero bias, and earns the highest share of positive user ratings.
  • Gemini 3.1 Pro is the strongest featured ranker and leans generous. It may be the reviewer most inclined to see the glass as half full.
  • GPT-5.6 Terra gives the most detailed feedback but scores more strictly. Reviewer #2 has entered the chat, with footnotes.
  • These dynamic results feed the model recommendations shown to users in Step 3 of the review workflow.
CSPaper benchmark update showing 700 papers, 8 models, 4 metrics, and the featured GLM-5v Turbo, GPT-5.6 Terra, and Gemini 3.1 Pro comparison

The result in 30 seconds

These statistics are a snapshot of the benchmark dataset and results obtained on 20 July 2026. There is no single winner, which is exactly why the benchmark uses more than one metric. The cleanest headline belongs to Z.AI: GLM-5v Turbo records the best NMAE of the featured models, the smallest mean bias by a wide margin, and a strong SRC. It also receives the highest share of 4 and 5 star user ratings.

Gemini 3.1 Pro remains the best ranker in this group. GPT-5.6 Terra is not the numerical leader here. It produces the longest reviews and lands in a competitive range, but its scores are less accurate than GLM's and its paper ordering is less consistent than Gemini's. A new model does not need a gold medal in every column to be useful. It needs a clear profile, and now Terra has one.

NMAE

Score error

lower is better
GLM-5v Turbo0.165
Gemini 3.1 Pro0.178
GPT-5.6 Terra0.193

NME

Score bias

zero is best
GLM-5v Turbo-0.000
Gemini 3.1 Pro+0.047
GPT-5.6 Terra-0.054
harshbalancedgenerous

SRC

Ranking agreement

higher is better
GLM-5v Turbo0.621
Gemini 3.1 Pro0.658
GPT-5.6 Terra0.559
Figure 1. Simple means across the 38 public venue and track rows shown on the CSPaper benchmark page. Snapshot taken from results obtained on 20 July 2026. Every featured model has results for all 38 rows. NME is centered on zero: left is harsher and right is more generous. These summaries make the public matrix readable; the underlying evaluation corpus contains 700 papers.

Four metrics, four different questions

Let N be the number of benchmark papers, Sgt the expert-calibrated score, and Spred the CSPaper prediction. The score range is venue specific, so normalization lets an ICLR scale and an ACL scale live in the same comparison.

NMAE: score accuracy

Lower is better
NMAE=1Ni=1NSgt(i)Spred(i)SmaxSmin\text{NMAE}=\frac{1}{N}\sum_{i=1}^{N}\frac{|S_{\text{gt}}^{(i)}-S_{\text{pred}}^{(i)}|}{S_{\text{max}}-S_{\text{min}}}

The average absolute score gap, divided by the venue's score range. A value of 0.165 means the prediction misses by 16.5% of the available rating scale on average. This is our first headline metric because it answers the simplest question: how close was the score?

NME: score bias

Zero is ideal
NME=1Ni=1NSpred(i)Sgt(i)SmaxSmin\text{NME}=\frac{1}{N}\sum_{i=1}^{N}\frac{S_{\text{pred}}^{(i)}-S_{\text{gt}}^{(i)}}{S_{\text{max}}-S_{\text{min}}}

The signed version of NMAE. Positive means the model scores too generously; negative means it scores too harshly. Errors in opposite directions can cancel, so NME diagnoses calibration but never replaces NMAE.

SRC: ranking quality

Higher is better
SRC=16i=1Ndi2N(N21)\text{SRC}=1-\frac{6\sum_{i=1}^{N}d_i^2}{N(N^2-1)}

Spearman rank correlation compares the predicted paper order with the expert-calibrated order. Here, dᵢ is the difference between the two ranks for paper i. A score of 1 is perfect ordering, 0 means no rank correlation, and -1 means the order is reversed. This is our other headline metric because review systems must separate stronger papers from weaker ones, not just hover near the average score.

AWC: review length

Descriptive, not a quality score
AWC=1Ni=1NWords(Rcspr(i))\text{AWC}=\frac{1}{N}\sum_{i=1}^{N}|\text{Words}(R_{\text{cspr}}^{(i)})|

Average Word Count measures how much text the full review contains. Longer can mean more justification, but it can also mean more repetition. We report it because useful feedback needs room to explain itself, while treating it as the least important of the four metrics.

NMAE and SRC carry the most weight in our reading. NME tells us whether errors lean high or low. AWC tells us how much was written, not whether it was right. Four gauges are better than one, but none of them can read a reviewer's mind. More on that shortly.

Benchmark results become model recommendations

Users do not have to memorize four metrics before starting a review. In Step 3, CSPaper recommends models using the latest benchmark results for the selected venue. Labels such as low error, high correlation, concise, verbose, over-estimator, and under-estimator turn the live measurements into practical choices. As the benchmark changes, the recommendations can change with it.

Step 3 of CSPaper Review showing model recommendations and venue-specific benchmark results
Figure 2. Model selection in Step 3. Recommendations are tied to dynamic, venue-specific benchmark results, so the labels shown for one venue need not match the all-venue averages in this article.

The featured three, side by side

ModelNMAE ↓NME → 0SRC ↑AWC4 or 5 stars
GLM-5v TurboZ.AI0.165best-0.000best0.62117,13761%best
Gemini 3.1 ProGoogle0.178+0.0470.658best8,78751%
GPT-5.6 TerraOpenAI0.193-0.0540.55917,82154%
Table 1. Featured model comparison using simple means across the public venue and track rows available on 20 July 2026, plus the share of delivered reviews that users rated 4 or 5 stars.

GLM's case is unusually tidy. Its NMAE is 7% lower than Gemini 3.1 Pro and 14% lower than GPT-5.6 Terra in this public-row summary. Its mean NME is -0.0004, effectively centered on zero. It does not win SRC, but 0.621 is still a strong ranking result. Put plainly: GLM gets scores close, avoids a clear optimistic or pessimistic lean, and usually orders papers well.

That is excellent work from Z.AI. The model has earned a victory lap.

Gemini 3.1 Pro tells a different story. It has the best SRC at 0.658, so it is the strongest at preserving relative paper order. Its reviews are much shorter, roughly half the length of GLM's and Terra's, and its positive NME suggests a mild optimistic tilt. Gemini seems to be the nice reviewer in the room: good at sorting the stack, and a little more willing to round up.

GPT-5.6 Terra is the wordiest of the three and is mildly conservative on average. At -0.054 NME, it behaves a bit like Reviewer #2: strict score, long report, and one more requested ablation. Its NMAE and SRC do not lead this round, so we will keep testing whether that extra detail helps authors act on the feedback.

Who wins most often?

Here is the ten-second view. For each venue or track row, one win means the model posted the lowest NMAE, the NME closest to zero, or the highest SRC. The headline total combines NMAE and SRC, our two most important metrics. GLM leads with 20, followed by Gemini 3.1 Pro with 16 and Terra with 15. Ties count for every model sharing the best displayed value.

ModelNMAE winsNME winsSRC winsHeadline winsNMAE + SRC
GLM-5v Turbo771320
Gemini 3.1 Pro87816
GPT-5.6 Terra87715
GPT-5.455611
Gemini 2.5 Pro5749
GPT-5.13636
GPT-54215
Gemini 3 Flash1334
Table 2. Number of public venue or track rows won by each supported model. Tied displayed values count as wins for every tied model. Headline wins combine NMAE and SRC.

This keeps all eight supported models visible without repeating the detailed values already published on the benchmark page. Coverage still matters: the featured models and GPT-5.4 cover all 38 rows, while older models cover 29 to 31. Win counts for those older models therefore come from fewer opportunities.

Scores are not the whole review

The four benchmark metrics lean heavily on scores and ordering. Review quality is broader. A useful critique should be correct, specific, fair, and possible to act on. After a user receives the review report for their paper, CSPaper asks them to rate it from 1 to 5 stars. This post-delivery feedback channel was introduced as part of the operating system described in Computational Research Assessment, and it gives us a qualitative signal that score metrics alone cannot provide.

User signal

Share of reviews rated 4 or 5 stars

higher is better
GLM-5v Turbo
61%
GPT-5.6 Terra
54%
Gemini 3.1 Pro
51%
Figure 3. Share of delivered paper-review reports that users rated 4 or 5 stars on CSPaper's 1-to-5-star scale. This signal is subjective and the samples are not controlled like the benchmark corpus, so we read it alongside the experimental metrics, not as a replacement for them.

GLM's 61% is the best of the three. That agreement between the main score metric and the user signal is encouraging. It is not proof that GLM wins every notion of review quality, but it is a good reason to keep taking fast-moving open-model families seriously. We plan to explore support for more of them as their vertical performance improves.

What this benchmark can and cannot say

A 700-paper test is large enough to expose patterns that a few cherry-picked examples cannot. It also covers 38 public venue and track configurations, which matters because a good review is rubric dependent. The original CSPaper system paper began with a 100-paper model-selection benchmark. The current corpus is seven times larger, and the public matrix makes the venue-level variation visible.

Still, the benchmark has limits. Human scores are not scientific truth. NMAE can reward safe predictions near the middle. SRC says little about whether a critique found the right flaw. NME can hide large errors that cancel. AWC cannot tell insight from padding. User stars can reflect tone, speed, or expectations as much as correctness.

This is why CSPaper's research direction is verification first. Our work on verification-first AI argues that review tools should increase the amount of science we can check, not merely imitate historical judgments. The Computational Research Assessment agenda likewise calls for evidence-linked critiques, explicit escalation when assessors disagree, open benchmarks, and red tests for gaming and bias.

Every new agent has to earn its place

When we build a conference, journal, or workshop agent, we do more than attach a rubric to a prompt. We bake official venue criteria, score-calibration rules, evidence and tool policies, reasoning policies, review templates, expert examples, and known failure cases into a versioned agent harness. Candidates are regenerated offline and evaluated against fixed validation packs before deployment. That controlled factory loop is documented in Self-Improving Agent Factories for Paper Review.

The reason is simple: scientific feedback is a vertical task. Progress on general reasoning does not automatically transfer to a venue's evidence standards. The position paper AI Should Verify, Not Judge, Scientific Work makes the boundary clear: these tools should help authors verify claims and improve manuscripts, not replace human peer review. We benchmark the exact job, keep the test set stable, and surface the latest results so users understand why a model is recommended and when that recommendation changes.

This also fits the broader bottleneck described by Epistemic Throughput: generation is cheap, but careful verification remains scarce. Better screening can help focus that scarce attention. It cannot make verification optional.

The CSPaper Review system paper at INLG 2025 documented the first version of this loop: venue-specific rubrics, score-conditioned review candidates, selection, synthesis, and calibration. The benchmark you see today is the same habit grown up: more papers, more venues, more metrics, more models, and public results.

The takeaway

Pick models for the review job, not for the logo on the box.

GLM-5v Turbo is the standout addition in this round. Gemini 3.1 Pro remains the strongest ranker of the featured three. GPT-5.6 Terra brings more detail, but not the best score fit. The benchmark is dynamic: we will keep refreshing the evidence, expanding providers and model families, and adding fast-moving open-source LLMs when they show strong vertical performance for scientific review.

References

  1. Lele Cao, Lei You, Kai Xie, Weiping Ding, Yong Du, Sven Salmonsson, Yumin Zhou, and Vilhelm von Ehrenheim. CSPaper Review: Fast, Rubric-Faithful Conference Feedback. INLG 2025 System Demonstrations.
  2. Lei You, Lele Cao, and Iryna Gurevych. Preventing the Collapse of Peer Review Requires Verification-First AI. ICML 2026 AI for Science Workshop.
  3. Lele Cao, Lei You, Kai Xie, Weiping Ding, Yong Du, Sven Salmonsson, Yumin Zhou, and Vilhelm von Ehrenheim. Adopt Machine-Human Collaboration Peer-Review through Computational Research Assessment. ICML 2026 AI for Science Workshop.
  4. Lele Cao. Self-Improving Agent Factories for Paper Review: When Constrained Beats Open-Ended. KDD 2026 Workshop on SciSoc: Agents & LLMs.
  5. Prabhant Singh, Thanh Gia Hieu Khuong, Vlasta Sikimić, Benedictus Kent Rachmat, Kola Ayonrinde, Ihsan Ullah, Christina Lioma, Luis Oala, Kevin Qinghong Lin, Lele Cao, Neil F. Abernethy, Hilde Weerts, and Joaquin Vanschoren. Position: AI Should Verify, Not Judge, Scientific Work. ICML 2026 AI for Science Workshop.
  6. Lei You. Epistemic Throughput: Fundamental Limits of Attention-Constrained Inference. 2026.
700 Papers, 8 Models: OpenAI, Google, and Z.AI's Flagship LLMs Compared | CSPaper — CSPaper