Protein & Biomolecular Sequence Modeling
Benchmarking Surrogate-Trained Sequence Optimization Against Empirical Fitness Landscapes
Generated automatically from the limitations stated in 7 papers (ICML, NeurIPS, AISTATS), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Current machine learning models for protein and biomolecular sequence design are evaluated almost exclusively against learned surrogate oracles (such as CNN fitness predictors) rather than verified experimental outcomes. Because black-box optimizers can easily exploit surrogate artifacts and out-of-distribution prediction errors, reported in-silico fitness gains often fail to reflect actual biological activity. Without systematic evaluation across empirically ground-truthed landscapes, it is unknown which sequence design methods reliably find functional sequences versus which merely overfit to computational evaluator flaws.
Why it matters
Sequence designers can identify which optimization strategies produce robust improvements across protein families rather than exploiting idiosyncratic evaluator shortcuts, establishing reliable in-silico evaluation protocols before committing to wet-lab synthesis.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Benchmark existing sequence optimization algorithms (e.g., autoregressive generation, genetic algorithms, MCMC) on large-scale deep mutational scanning (DMS) datasets (such as ProteinGym or FLIP) using disjoint train/oracle/test splits to measure true rank correlation and fitness recovery when surrogates extrapolate.
- 2
Construct a multi-oracle robustness testbed that trains diverse surrogate architectures (CNNs, ESM-based regressors, GNNs) on identical assay data and evaluates the degree of oracle-divergence and score collapse when optimization trajectories leave the training distribution.
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
The study would fail if existing retrospective DMS datasets are too sparse to evaluate high-order combinatorial variants without training a new proxy model, thereby reintroducing the exact surrogate gap under investigation.
Evidence
Each paper's own statement of the limitation, verbatim.
- STRIDE: Post-Training LLMs to Reason and Refine Bio-Sequences via Edit TrajectoriesICML 2026
Evaluations rely entirely on in silico surrogate oracles (e.g., GFP oracles primarily trained on substitution data) without wet-lab experimental confirmation.
- Steering Generative Models with Experimental Data for Protein Fitness OptimizationNeurIPS 2025
Evaluation relies on computational oracles (supervised models trained on real data) rather than wet-lab measurements, so absolute performance may not reflect true fitness landscapes; authors advise only relative comparisons
- Protein Fitness Landscape: Spectral Graph Theory PerspectiveAISTATS 2025
Evaluation uses an in-silico surrogate evaluator (a trained CNN oracle) rather than real wet-lab experimental validation, so reported fitness gains may not transfer to real protein function
- Generative property enhancer: implicit guided generation through conditional density estimationNeurIPS 2025
Validation is entirely in silico with a proxy fitness predictor, not wet-lab confirmed; predictor reliability is acknowledged as questionable
- Data Distillation for extrapolative protein design through exact preference optimizationICLR 2025
Evaluation relies entirely on in-silico evaluators whose MSE and rank correlation deteriorate in the extrapolation region (especially non-smooth on GFP), so conclusions may not reflect real fitness
- A Variational Perspective on Generative Protein Fitness OptimizationICML 2025
Evaluation is purely in-silico using a trained oracle; no experimental validation of generated sequences
- Importance Weighted Expectation-Maximization for Protein Sequence DesignICML 2023
Designed sequences were never wet-lab validated; fitness is measured by oracle surrogate models trained on real data, so reported scores may reflect oracle uncertainty/exploitation
Nearest existing work
- ProSpero: Active Learning for Robust Protein Design Beyond Wild-Type NeighborhoodsNeurIPS 2025
- Proximal Exploration for Model-guided Protein Sequence DesignICML 2022
- FLIP2: Expanding Protein Fitness Landscape Benchmarks for Real-World Machine Learning ApplicationsICML 2026
- A Variational Perspective on Generative Protein Fitness OptimizationICML 2025
- Population-Based Black-Box Optimization for Biological Sequence DesignICML 2020
- Interpretability Driven Evolutionary Approach for the Design of Biological SequencesICML 2026
- Robust Optimization in Protein Fitness Landscapes Using Reinforcement Learning in Latent SpaceICML 2024
- Importance Weighted Expectation-Maximization for Protein Sequence DesignICML 2023
- Robust Model-Based Optimization for Challenging Fitness LandscapesICLR 2024
- ProteinNPT: Improving Protein Property Prediction and Design with Non-Parametric TransformersNeurIPS 2023
- Unifying Likelihood-free Inference with Black-box Optimization and BeyondICLR 2022
- DePLM: Denoising Protein Language Models for Property OptimizationNeurIPS 2024
- TadA-Bench: A Million-Variant Benchmark for Future-Round Discovery Toward Agentic Protein EngineeringICML 2026
- Data Distillation for extrapolative protein design through exact preference optimizationICLR 2025
- Multi-Scale Representation Learning for Protein Fitness PredictionNeurIPS 2024
Related open problems
Protein & Biomolecular Sequence Modeling
Benchmarking the Fidelity of In Silico Protein Design Proxies Against Experimental Biophysical Assays
Generative biomolecular modeling currently assesses sequence designs almost entirely through computational surrogates—such as AlphaFold confidence (pLDDT/pTM), ESMFold self-consistency, and Rosetta energy terms—rather than physical assay measurements. Because these proxies serve as both optimization objectives and evaluation metrics, models risk overfitting to artifacts of structure-prediction heuristics rather than true biophysical viability. Without systematic evaluation against physical assays across sequence design methods, it remains unknown whether improvements on in silico leaderboards translate to viable expression, foldability, or binding in physical systems.
Protein & Biomolecular Sequence Modeling
A Systematic Robustness Study of Genomic Foundation Models Beyond the Human Reference Genome
Current genomic sequence models are predominantly pretrained and evaluated strictly on the single human reference genome (GRCh38/hg38). Consequently, it remains unknown whether representations trained on a single reference sequence maintain their predictive capabilities across distinct chromosomes, across individual genetic variation in population-scale cohorts, or across non-human species genomes. Without cross-setting robustness benchmarks, practitioners cannot determine whether model performance transfers to non-reference genomic contexts or degrades under distribution shifts.
Protein & Biomolecular Sequence Modeling
Benchmarking and Mitigating Model Sensitivity to Predicted Structure Quality in Protein Sequence Modeling
Current structure-conditioned protein sequence and representation models routinely filter out test proteins with predicted confidence scores below pLDDT 70 or assume near-experimental coordinate fidelity. Consequently, these models cannot reliably be applied to disordered proteins, orphan sequences, or hard-to-predict folds where high-confidence static predictions do not exist. Downstream pipelines experience unquantified error propagation when forced to consume predicted structures from AlphaFold2 or ESMFold, yet no systematic evaluation exists mapping how performance degrades across the full spectrum of structural prediction quality.
Protein & Biomolecular Sequence Modeling
A Benchmark of Experimentally Validated Functional Motifs for Training and Evaluating Motif-Conditioned Generative Models
Motif-conditioned generators in both protein structure and molecular design are trained and scored against motifs that are artifacts of the pipeline — random residue crops, rigid bond-breaking heuristics, or proxy-evaluated structures — rather than motifs whose function has been experimentally verified. Because these synthetic motifs may not match the distribution of true functional motifs, reported success rates measure the ability to satisfy a surrogate objective, not the ability to produce functional designs. Until a curated set of validated motifs with defined success criteria exists, no model comparison (structure-based vs. sequence-based, heuristic vs. expert extraction) can be trusted, and designs cannot be credibly advanced toward wet-lab validation.