Protein & Biomolecular Sequence Modeling
Benchmarking and Mitigating Model Sensitivity to Predicted Structure Quality in Protein Sequence Modeling
Generated automatically from the limitations stated in 7 papers (ICLR, NeurIPS, CVPR), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Current structure-conditioned protein sequence and representation models routinely filter out test proteins with predicted confidence scores below pLDDT 70 or assume near-experimental coordinate fidelity. Consequently, these models cannot reliably be applied to disordered proteins, orphan sequences, or hard-to-predict folds where high-confidence static predictions do not exist. Downstream pipelines experience unquantified error propagation when forced to consume predicted structures from AlphaFold2 or ESMFold, yet no systematic evaluation exists mapping how performance degrades across the full spectrum of structural prediction quality.
Why it matters
Enables sequence-structure models to operate reliably across full proteomes—including low-confidence, flexible, or hard-to-predict targets—without ad-hoc dataset truncation or unmonitored error propagation.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Stratified structural quality benchmark: Evaluate representative inverse folding and structure-representation models across test sets stratified into granular pLDDT/pTM bins and matched experimental-predicted structure pairs, measuring sequence recovery, perplexity, and downstream property prediction degradation.
- 2
Confidence-aware structural encoding: Modify invariant geometric graph neural networks to ingest per-residue and pairwise confidence metrics (pLDDT and PAE) as spatial variance features, measuring whether soft uncertainty conditioning outperforms hard pLDDT thresholding and fallback heuristics.
- 3
Realistic error injection: Benchmark model robustness against synthetic geometric perturbations calibrated to empirical AlphaFold prediction errors across distinct secondary structure classes.
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
If low-pLDDT regions in practice correspond almost exclusively to intrinsically disordered segments where static 3D coordinate modeling is inherently ill-posed, making sequence-only baselines unsurpassable regardless of structural uncertainty modeling.
Evidence
Each paper's own statement of the limitation, verbatim.
- Multimodal Protein Language Models for Enzyme Kinetic Parameters: From Substrate Recognition to Conformational AdaptationCVPR 2026
Susceptibility to error propagation when utilizing predicted structures (OpenFold/ESMFold) instead of experimental data
- MeToken: Uniform Micro-environment Token Boosts Post-Translational Modification PredictionICLR 2025
Heavy reliance on AlphaFold-predicted structures which may introduce model-specific geometric biases.
- Immunogenicity Prediction with Dual Attention Enables Vaccine Target SelectionICLR 2025
Structures are predicted (ESMFold/AlphaFold2) rather than experimental; performance depends on structure-prediction accuracy, and scarcity of crystal structures prevented direct comparison with experimental structures
- BAnG: Bidirectional Anchored Generation for Conditional RNA DesignICML 2025
Test set requires AlphaFold2 structures with pLDDT > 70%; proteins with poorly predicted structures are excluded, and performance on disordered or hard-to-predict proteins is unknown
- Multi-Scale Representation Learning for Protein Fitness PredictionNeurIPS 2024
Performance degrades on low-quality AlphaFold2-predicted structures (below pLDDT 70 the model falls back to ESM-2), and pre-training on only ~31k CATH structures may underuse available structural data
- MAPE-PPI: Towards Effective and Efficient Protein-Protein Interaction Prediction via Microenvironment-Aware Protein EmbeddingICLR 2024
Relies on predicted (AlphaFold2) 3D structures; accuracy may degrade for proteins without reliable structure predictions
- DePLM: Denoising Protein Language Models for Property OptimizationNeurIPS 2024
Depends on AlphaFold2-predicted structures via ESM-IF, so accuracy degrades for proteins without reliable structural predictions
Nearest existing work
- AlphaFold Database Debiasing for Robust Inverse FoldingNeurIPS 2025
- Evaluating Representation Learning on the Protein Structure UniverseICLR 2024
- Inference-time optimization for experiment-grounded protein ensemble generationICML 2026
- AlphaFold Meets Flow Matching for Generating Protein EnsemblesICML 2024
- Learning inverse folding from millions of predicted structuresICML 2022
- CalPro: Prior-Aware Evidential Conformal Prediction with Structure-Aware Sensitivity Bounds for Protein StructuresICML 2026
- Distilling Structural Representations into Protein Sequence ModelsICLR 2025
- DualMPNN: Harnessing Structural Alignments for High-Recovery Inverse Protein FoldingNeurIPS 2025
- KW-Design: Pushing the Limit of Protein Design via Knowledge RefinementICLR 2024
- Learning from Protein Structure with Geometric Vector PerceptronsICLR 2021
- Learning conformational ensembles of proteins based on backbone geometryNeurIPS 2025
- Efficient Generative Modelling of Protein Structure Fragments using a Deep Markov ModelICML 2021
- SimpleFold: Folding Proteins is Simpler than You ThinkICLR 2026
- FAFE: Immune Complex Modeling with Geodesic Distance Loss on Noisy Group FramesICML 2024
- Deep Supervised and Convolutional Generative Stochastic Network for Protein Secondary Structure PredictionICML 2014
Related open problems
Protein & Biomolecular Sequence Modeling
Benchmarking the Fidelity of In Silico Protein Design Proxies Against Experimental Biophysical Assays
Generative biomolecular modeling currently assesses sequence designs almost entirely through computational surrogates—such as AlphaFold confidence (pLDDT/pTM), ESMFold self-consistency, and Rosetta energy terms—rather than physical assay measurements. Because these proxies serve as both optimization objectives and evaluation metrics, models risk overfitting to artifacts of structure-prediction heuristics rather than true biophysical viability. Without systematic evaluation against physical assays across sequence design methods, it remains unknown whether improvements on in silico leaderboards translate to viable expression, foldability, or binding in physical systems.
Protein & Biomolecular Sequence Modeling
A Systematic Robustness Study of Genomic Foundation Models Beyond the Human Reference Genome
Current genomic sequence models are predominantly pretrained and evaluated strictly on the single human reference genome (GRCh38/hg38). Consequently, it remains unknown whether representations trained on a single reference sequence maintain their predictive capabilities across distinct chromosomes, across individual genetic variation in population-scale cohorts, or across non-human species genomes. Without cross-setting robustness benchmarks, practitioners cannot determine whether model performance transfers to non-reference genomic contexts or degrades under distribution shifts.
Protein & Biomolecular Sequence Modeling
Benchmarking Surrogate-Trained Sequence Optimization Against Empirical Fitness Landscapes
Current machine learning models for protein and biomolecular sequence design are evaluated almost exclusively against learned surrogate oracles (such as CNN fitness predictors) rather than verified experimental outcomes. Because black-box optimizers can easily exploit surrogate artifacts and out-of-distribution prediction errors, reported in-silico fitness gains often fail to reflect actual biological activity. Without systematic evaluation across empirically ground-truthed landscapes, it is unknown which sequence design methods reliably find functional sequences versus which merely overfit to computational evaluator flaws.
Protein & Biomolecular Sequence Modeling
A Benchmark of Experimentally Validated Functional Motifs for Training and Evaluating Motif-Conditioned Generative Models
Motif-conditioned generators in both protein structure and molecular design are trained and scored against motifs that are artifacts of the pipeline — random residue crops, rigid bond-breaking heuristics, or proxy-evaluated structures — rather than motifs whose function has been experimentally verified. Because these synthetic motifs may not match the distribution of true functional motifs, reported success rates measure the ability to satisfy a surrogate objective, not the ability to produce functional designs. Until a curated set of validated motifs with defined success criteria exists, no model comparison (structure-based vs. sequence-based, heuristic vs. expert extraction) can be trusted, and designs cannot be credibly advanced toward wet-lab validation.