Open Problems

RLHF & Preference Optimization

Quantifying Evaluation Circularity and Judge Bias in Preference Optimization

Barrier to removePartly addressed
Possible candidate · 3/5 runs9 papers report this89% from 2025+

Generated automatically from the limitations stated in 9 papers (ICML, ACL, NAACL), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.

The problem

Current preference optimization and RLHF methods evaluate policy improvements almost exclusively through proprietary LLM judges (primarily GPT-4/GPT-4o) and automated reward models that often share architectures or training data with the policies being tested. This reliance introduces severe evaluation circularity and metric mismatch, where reported win-rate gains frequently reflect alignment with a specific judge's stylistic preferences rather than genuine policy improvements. Consequently, researchers cannot determine whether novel preference optimization algorithms generalize or merely exploit the biases of automated proxies.

Why it matters

Enables researchers to validate preference optimization algorithms without proprietary API dependencies, providing a standardized way to detect reward hacking and judge-specific overfitting.

Ways to approach it

Prior-work checks are free with an account. Results someone already ran are shown to everyone.

  1. 1

    Conduct a cross-evaluation audit of existing preference-tuned checkpoints (e.g., DPO, RLOO, PPO) across disjoint open reward models, proprietary judges, and a controlled set of human annotations, measuring the rank-order correlation and win-rate variance across judge types.

  2. 2

    Systematically measure reward circularity by training policies on one family of reward signals (e.g., DPO implicit rewards, PickScore, Eurus) and evaluating them against strictly held-out, independently trained reward architectures to quantify metric inflation.

  3. 3

    Construct an open-source, reproducible evaluation protocol combining length-normalized paired comparisons, rule-verifiable sub-tasks, and diverse open reward model ensembles, measuring its alignment with existing human preference datasets versus single-judge GPT-4 baselines.

Have a different approach?

Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.

Free · 3 checks per day

Why it might fail

High costs and logistical bottlenecks in acquiring sufficiently dense, high-quality human annotations to establish ground-truth correlations could stall the project before decisive conclusions are reached.

Evidence

Each paper's own statement of the limitation, verbatim.

Show all 9 papers

Nearest existing work

Related open problems

RLHF & Preference Optimization

Barrier to removeOpen

Preference Optimization Under Unlabeled and Heterogeneous Real-World Annotator Noise

Current theoretical and algorithmic advances in noise-robust and personalized reward modeling fundamentally rely on preconditions absent in standard preference datasets: either assuming synthetic noise models generated by golden reward models or requiring persistent annotator identifiers and graphs. In practice, public preference datasets are largely anonymous, aggregated across heterogeneous annotator pools, and exhibit non-monotone human error patterns that synthetic noise models fail to capture. As a result, noise-mitigation and debiasing techniques developed under idealized assumptions remain unvalidated and often brittle when deployed on actual crowdsourced preference data.

Possible candidate · 3/5 runs6 papers report this83% from 2025+

RLHF & Preference Optimization

Scope to testOpen

Empirical Robustness of Preference Optimization Across Synthetic, Multi-Annotator, and Real-World Feedback Distributions

Current alignment algorithms (e.g., DPO, PPO, KTO) are almost exclusively evaluated on synthetic user proxies, converted scalar scores, or curated single-annotator datasets like UltraFeedback. Because individual studies restrict evaluation to narrow, semi-synthetic testbeds (such as session-similarity graph samplings or single-dimension datasets like HelpSteer2), it remains unknown whether standard preference optimization algorithms generalize reliably when exposed to genuine multi-annotator disagreement and raw binary human signals. Without a cross-setting evaluation spanning these distinct regimes, practitioners cannot determine whether reported algorithmic gains reflect true alignment robustness or artifacts of synthetic data curation.

Possible candidate · 3/5 runs4 papers report this75% from 2025+

RLHF & Preference Optimization

Barrier to removeOpen

Eliminating Matched-Pair Preconditions and Scale Sensitivity in Contrastive Preference Guidance

Current contrastive guidance techniques for preference optimization structurally depend on having access to a matched tuned/untuned small proxy model pair whose latent reward aligns with the target preference. When such pairs are unavailable or poorly matched, guidance quality collapses and caps output quality at the proxy's low baseline capability (~10% win rate), while weak-to-strong self-steering without external reward signals fails to converge. Furthermore, the guidance scale parameter $\gamma$ exhibits high variance in noisy regions and fails to transfer across differing target architectures without periodic, expensive reward re-estimation.

Strong candidate · 4/5 runs3 papers report this33% from 2025+
Generated automatically, not curated by hand. Automated prior-work checks catch about a third of existing work, so treat this problem as a lead to investigate.