Protein & Biomolecular Sequence Modeling
Benchmarking the Fidelity of In Silico Protein Design Proxies Against Experimental Biophysical Assays
Generated automatically from the limitations stated in 19 papers (ICML, ICLR, NeurIPS), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Generative biomolecular modeling currently assesses sequence designs almost entirely through computational surrogates—such as AlphaFold confidence (pLDDT/pTM), ESMFold self-consistency, and Rosetta energy terms—rather than physical assay measurements. Because these proxies serve as both optimization objectives and evaluation metrics, models risk overfitting to artifacts of structure-prediction heuristics rather than true biophysical viability. Without systematic evaluation against physical assays across sequence design methods, it remains unknown whether improvements on in silico leaderboards translate to viable expression, foldability, or binding in physical systems.
Why it matters
Establishes quantitative error bounds for purely computational design pipelines and identifies which in silico metrics reliably predict physical protein expression and binding across diverse design regimes.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Benchmark existing in silico scoring proxies (AlphaFold2/3, ESMFold, Rosetta, ProGen2 likelihood) across standardized high-throughput experimental datasets (e.g., ProteinGym, Mega-scale cDNA display stability assays, and experimental binding screens), measuring Rank Correlation (Spearman $\rho$) and Top-$k$ hit enrichment rates.
- 2
Select candidate sequences generated by diverse sequence design architectures (autoregressive language models, diffusion models, inverse folding models) and measure their synthesis success, soluble expression yields, and circular dichroism stability spectra via an academic high-throughput expression pipeline or commercial cell-free expression service.
- 3
Construct an open evaluation suite and leaderboard that reports calibration error and false-positive rates for computational proxy metrics relative to ground-truth wet-lab outcomes.
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
High cost and throughput limitations of physical wet-lab assays may restrict validation scale, or existing public high-throughput assay datasets may suffer from narrow sequence-space coverage that limits generalization to de novo designs.
Evidence
Each paper's own statement of the limitation, verbatim.
- TD3B: Transition-Directed Discrete Diffusion for Allosteric Binder GenerationICML 2026
Validation relies entirely on in silico predictors (Direction Oracle, neural affinity predictors, AlphaFold3) without wet-lab experimental assays (e.g., cAMP or β-arrestin recruitment).
- SwitchCraft: A Programmatic Framework for Designing State-Switching ProteinsICML 2026
Validation is entirely in silico based on Boltz-1 predictions, lacking extensive wet-lab experimental confirmation.
- Scaling Atomistic Protein Binder Design with Generative Pretraining and Test-Time ComputeICLR 2026
Validation is entirely restricted to in-silico structure prediction metrics (AlphaFold2-Multimer and RosettaFold-3 scores) without experimental wet-lab binding validation.
- RIDER: 3D RNA Inverse Design with Reinforcement Learning-Guided DiffusionICLR 2026
Relies entirely on computational structure prediction oracles (RhoFold, AlphaFold3) with no wet-lab experimental validation of designed RNAs.
- Proteo-R1: Reasoning Foundation Models for De Novo Protein DesignICML 2026
Lacks experimental wet-lab binding affinity or expression validation, relying entirely on computational structural and Rosetta energy evaluations.
- Property-Driven Protein Inverse Folding with Multi-Objective Preference AlignmentICLR 2026
Validation relies entirely on in silico proxy predictors with no wet-lab experimental validation of designed sequences.
- Pallatom-Ligand: an All-Atom Diffusion Model for Designing Ligand-Binding ProteinsICLR 2026
Validation is entirely in silico using AlphaFold3 metrics, with no experimental wet-lab binding assays or biophysical characterization.
- PDFBench: A Benchmark for De Novo Protein Design from FunctionICML 2026
The benchmark relies entirely on computational proxies (e.g., ESMFold, ProTrek) and lacks physical wet-lab validation of the designed proteins.
Show all 19 papers
- PDAgent: An LLM-Driven Autonomous Agent Framework Towards *In Silico* Protein Design via Directed MutationICML 2026
Validation is entirely in silico using computational surrogate predictors (ESMFold, DDGun, PPTStab) without wet-lab biochemical or experimental validation.
- Multi-state Protein Design with DynamicMPNNICLR 2026
Lacks experimental wet-lab validation; all results rely entirely on in silico computational proxies (AlphaFold 3 and BioEmu refoldability).
- Multi-Objective Protein Design via Memory-Aware Test-Time Scaling in Diffusion ModelsICML 2026
Evaluations are entirely computational in silico (e.g., ESMFold/predicted pLDDT, pTM) without wet-lab experimental validation of expressibility, solubility, or biological activity.
- LineageFlow: Flow Matching for High-Fidelity Family-Aware Protein Sequence GenerationICML 2026
Relies entirely on in silico computational proxies (OmegaFold pLDDT, ESM-IF, ESM-2 property predictors) without experimental wet-lab validation of physical expression or enzymatic activity.
- Controlling Repetition in Protein Language ModelsICLR 2026
Biological plausibility is evaluated solely using in silico folding proxies (AlphaFold pLDDT/pTM) without wet-lab experimental validation of protein stability or function.
- Constrained Diffusion for Protein Design with Hard Structural ConstraintsICLR 2026
Evaluation is entirely in silico (structural metrics, DSSP, ProteinMPNN/AlphaFold2 consistency) without experimental wet-lab validation of expression or binding.
- Co-Generative De Novo Functional Protein DesignICML 2026
Relies entirely on in silico validation (ESMFold and DeepGO-SE) without wet-lab experimental verification.
- Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed SamplingICML 2026
All evaluation relies entirely on in silico AlphaFold2-Multimer metrics without experimental wet-lab binding or expression validation.
- (Be Cautious!) Bio-Foundation Models Are Not Yet Robust to Biologically Plausible Perturbations and ML TransformationsICML 2026
Findings are purely in silico and lack physical wet-lab experimental validation to directly measure downstream real-world biological impact.
- Pareto-Optimal Energy Alignment for Designing Nature-Like AntibodiesNeurIPS 2025
Evaluation relies almost entirely on in silico Rosetta energy metrics; no experimental in vitro binding validation is reported, and the authors note prior work shows these metrics can be manipulated.
- CarbonNovo: Joint Design of Protein Structure and Sequence Using a Unified Energy-based ModelICML 2024
Evaluation is purely in silico (self-consistency with ESMFold/OmegaFold, Rosetta energy, ProGen2 likelihood); no wet-lab experimental validation of foldability or function is provided.
Nearest existing work
- ProtDBench: A Unified Benchmark of Protein Binder Design and EvaluationICML 2026
- PDFBench: A Benchmark for De Novo Protein Design from FunctionICML 2026
- Learning to engineer protein flexibilityICLR 2025
- TadA-Bench: A Million-Variant Benchmark for Future-Round Discovery Toward Agentic Protein EngineeringICML 2026
- ProteinBench: A Holistic Evaluation of Protein Foundation ModelsICLR 2025
- FLIP2: Expanding Protein Fitness Landscape Benchmarks for Real-World Machine Learning ApplicationsICML 2026
- MutAtlas: A PDB-Wide Energy-Guided Atlas of Protein Mutation EffectsICML 2026
- Inference-time optimization for experiment-grounded protein ensemble generationICML 2026
- PDAgent: An LLM-Driven Autonomous Agent Framework Towards *In Silico* Protein Design via Directed MutationICML 2026
- SwitchCraft: A Programmatic Framework for Designing State-Switching ProteinsICML 2026
- Property-Driven Protein Inverse Folding with Multi-Objective Preference AlignmentICLR 2026
- Proximal Exploration for Model-guided Protein Sequence DesignICML 2022
- Data Distillation for extrapolative protein design through exact preference optimizationICLR 2025
- ProSpero: Active Learning for Robust Protein Design Beyond Wild-Type NeighborhoodsNeurIPS 2025
- SurfPro: Functional Protein Design Based on Continuous SurfaceICML 2024
Related open problems
Protein & Biomolecular Sequence Modeling
A Systematic Robustness Study of Genomic Foundation Models Beyond the Human Reference Genome
Current genomic sequence models are predominantly pretrained and evaluated strictly on the single human reference genome (GRCh38/hg38). Consequently, it remains unknown whether representations trained on a single reference sequence maintain their predictive capabilities across distinct chromosomes, across individual genetic variation in population-scale cohorts, or across non-human species genomes. Without cross-setting robustness benchmarks, practitioners cannot determine whether model performance transfers to non-reference genomic contexts or degrades under distribution shifts.
Protein & Biomolecular Sequence Modeling
Benchmarking Surrogate-Trained Sequence Optimization Against Empirical Fitness Landscapes
Current machine learning models for protein and biomolecular sequence design are evaluated almost exclusively against learned surrogate oracles (such as CNN fitness predictors) rather than verified experimental outcomes. Because black-box optimizers can easily exploit surrogate artifacts and out-of-distribution prediction errors, reported in-silico fitness gains often fail to reflect actual biological activity. Without systematic evaluation across empirically ground-truthed landscapes, it is unknown which sequence design methods reliably find functional sequences versus which merely overfit to computational evaluator flaws.
Protein & Biomolecular Sequence Modeling
Benchmarking and Mitigating Model Sensitivity to Predicted Structure Quality in Protein Sequence Modeling
Current structure-conditioned protein sequence and representation models routinely filter out test proteins with predicted confidence scores below pLDDT 70 or assume near-experimental coordinate fidelity. Consequently, these models cannot reliably be applied to disordered proteins, orphan sequences, or hard-to-predict folds where high-confidence static predictions do not exist. Downstream pipelines experience unquantified error propagation when forced to consume predicted structures from AlphaFold2 or ESMFold, yet no systematic evaluation exists mapping how performance degrades across the full spectrum of structural prediction quality.
Protein & Biomolecular Sequence Modeling
A Benchmark of Experimentally Validated Functional Motifs for Training and Evaluating Motif-Conditioned Generative Models
Motif-conditioned generators in both protein structure and molecular design are trained and scored against motifs that are artifacts of the pipeline — random residue crops, rigid bond-breaking heuristics, or proxy-evaluated structures — rather than motifs whose function has been experimentally verified. Because these synthetic motifs may not match the distribution of true functional motifs, reported success rates measure the ability to satisfy a surrogate objective, not the ability to produce functional designs. Until a curated set of validated motifs with defined success criteria exists, no model comparison (structure-based vs. sequence-based, heuristic vs. expert extraction) can be trusted, and designs cannot be credibly advanced toward wet-lab validation.