Benchmark on LLMs and Venues

We rigorously test leading AI models to ensure you get the most accurate, helpful reviews for your research papers. Below you'll find detailed performance metrics comparing the latest LLMs from OpenAI (GPT) and Google (Gemini) across various top-tier computer science conferences.

Try it yourself3 free reviews to start

Privacy and security first

CSPaper ensures all papers remain confidential and legally protected. Content is processed securely and safeguarded under enterprise-level agreements with LLM providers — guaranteeing it is never used to train any model.

Benchmark Metrics

Click on a metric below to view detailed benchmarks

Normalized Mean Absolute Error (NMAE) quantifies how close the predicted paper ratings are to the established ground-truth ratings. Lower NMAE values indicate better alignment between predicted and true review ratings.

Filter by venue type
Venue
OpenAIGPT-5.6-terra
GeminiGemini-3.1-pro
Z.ai (GLM)GLM-5v-turbo
GeminiGemini-2.5-pro
OpenAIGPT-5.4
OpenAIGPT-5
OpenAIGPT-5.1
GeminiGemini-3-flash
ESWAmain0.219±0.1730.156±0.1450.188±0.2110.263±0.2210.200±0.2150.200±0.2150.125±0.1150.188±0.189
FORGEintelligence0.136±0.1580.241±0.3250.074±0.0990.041±0.0550.111±0.1430.066±0.1390.066±0.0920.129±0.119
IJMLCregular0.346±0.1910.179±0.1280.238±0.1310.332±0.307
JBDregular0.214±0.1570.214±0.1870.179±0.0980.143±0.1830.214±0.1570.250±0.1910.214±0.2130.214±0.213
MLJregular0.193±0.1570.302±0.2820.239±0.2960.168±0.1340.242±0.1080.217±0.1650.217±0.1520.177±0.172
Neurocomputingmain0.270±0.1910.190±0.1380.122±0.0850.185±0.1500.275±0.1240.200±0.1420.300±0.2060.180±0.169
PAAregular0.204±0.2160.104±0.1050.121±0.1080.121±0.108
SciRepregular0.250±0.1770.100±0.1370.200±0.2090.150±0.1370.150±0.1370.200±0.1120.150±0.1370.300±0.209
TMLRregular0.250±0.2500.250±0.2240.236±0.1920.182±0.1410.136±0.1040.167±0.191
TVCregular0.125±0.1120.125±0.1580.250±0.2090.208±0.204
AAAImain technical0.186±0.1460.110±0.1020.072±0.0830.097±0.1110.164±0.1310.111±0.1130.145±0.1290.131±0.107
AAAIsafe and robust AI0.167±0.2310.230±0.2450.222±0.2710.214±0.2490.214±0.2270.159±0.2070.238±0.2410.175±0.274
AAMASmain technical0.190±0.0990.165±0.2260.130±0.0830.079±0.0980.134±0.1050.210±0.1680.183±0.1960.153±0.201
ACLmain0.158±0.1090.213±0.1470.094±0.1220.162±0.1310.154±0.1440.150±0.1500.205±0.1860.189±0.129
AISTATSmain0.306±0.2200.176±0.1850.196±0.1560.212±0.1710.266±0.2610.231±0.1780.231±0.1780.179±0.160
CoGtechnical and vision0.089±0.1250.137±0.1740.156±0.1050.163±0.2270.100±0.1240.155±0.1340.107±0.1280.211±0.220
CVPRmain0.253±0.2220.228±0.2050.204±0.1940.201±0.2130.242±0.2160.154±0.1810.249±0.2290.192±0.194
ECML-PKDDresearch0.074±0.0930.188±0.1910.164±0.1500.141±0.0460.102±0.1800.169±0.1580.141±0.1620.231±0.149
EMNLPmain0.115±0.1250.117±0.0880.097±0.1040.150±0.1270.117±0.1670.108±0.1040.183±0.1140.158±0.153
FAccTmain0.143±0.1520.250±0.2390.250±0.1910.107±0.112
GFMall0.192±0.1310.198±0.1300.170±0.1170.164±0.016
ICASSPregular paper0.264±0.1500.246±0.2400.204±0.1820.261±0.2000.275±0.2390.203±0.2190.290±0.2520.232±0.186
ICCtechnical symposia0.217±0.1750.287±0.2720.222±0.1790.260±0.2470.233±0.2410.273±0.2490.260±0.2530.353±0.256
ICLRmain0.068±0.0650.141±0.1170.094±0.0780.108±0.1070.115±0.1290.119±0.1080.104±0.0930.148±0.130
ICMEregular0.272±0.1400.164±0.1600.191±0.1320.253±0.1740.166±0.178
ICMLmain0.133±0.0860.186±0.1350.176±0.1410.196±0.1370.135±0.1150.190±0.1500.170±0.1230.266±0.183
ICMLposition0.129±0.1290.148±0.1070.139±0.0940.169±0.136
IJCAImain0.211±0.1330.174±0.1430.111±0.0960.159±0.1210.190±0.1410.198±0.1390.238±0.1620.095±0.130
IJCAIsurvey0.281±0.3640.188±0.2220.062±0.1160.094±0.1290.062±0.1160.188±0.2220.188±0.2910.188±0.222
IROSmain0.132±0.1550.125±0.0650.104±0.0760.104±0.0390.121±0.1090.093±0.0930.086±0.0430.143±0.080
KDDdatasets and benchmarks0.189±0.1790.217±0.2050.217±0.2050.214±0.2280.180±0.3110.175±0.1660.189±0.1980.256±0.208
KDDresearch0.183±0.1830.180±0.1070.148±0.1550.167±0.1800.231±0.1740.218±0.1580.205±0.1820.179±0.173
NeurIPSdatasets and benchmarks0.330±0.2240.161±0.1970.144±0.0810.193±0.2030.249±0.1960.170±0.2100.268±0.1870.174±0.202
NeurIPSmain0.194±0.1730.173±0.1680.151±0.1530.140±0.132
SIGIRfull paper0.188±0.2300.125±0.2020.264±0.3210.037±0.0840.163±0.2050.113±0.1710.237±0.2660.170±0.226
SIGIRshort paper0.132±0.1040.154±0.1920.154±0.1270.296±0.2880.132±0.1040.154±0.1270.189±0.1510.154±0.192
TheWebConfresearch0.174±0.2230.105±0.0850.115±0.0840.141±0.1370.149±0.1540.126±0.1140.116±0.0750.150±0.115
WACVmain0.160±0.1840.127±0.1350.178±0.1560.236±0.2160.255±0.1290.200±0.1260.273±0.2050.164±0.121

Understanding NMAE

Interpreting the numbers

  • Lower NMAE = better alignment. Top-performing models stay below 0.2.
  • The ± values show standard deviation, reflecting model stability across papers.

How we calculate it

NMAE=1Ni=1NSgt(i)Spred(i)SmaxSmin\text{NMAE} = \frac{1}{N} \sum_{i=1}^{N} \frac{|S_{\text{gt}}^{(i)} - S_{\text{pred}}^{(i)}|}{S_{\text{max}} - S_{\text{min}}}

NN: the total number of benchmark papers included in the evaluation.

Sgt(i)S_{\text{gt}}^{(i)}: ground-truth overall rating of the i-th benchmark paper.

Spred(i)S_{\text{pred}}^{(i)}: predicted overall rating produced by our CSPR agent.

SmaxS_{\text{max}}&SminS_{\text{min}}: venue-specific upper and lower bounds of the rating scale.

Want to learn more?

Explore the research and technical foundations behind CSPaper (a part of Scholar7):

Get high-quality reviews powered by the best AI models

We continuously evaluate and select the best-performing models to ensure you receive accurate, actionable feedback for your research.

Start your free review