Protein & Biomolecular Sequence Modeling
A Benchmark of Experimentally Validated Functional Motifs for Training and Evaluating Motif-Conditioned Generative Models
Generated automatically from the limitations stated in 6 papers (ICML, ICLR), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Motif-conditioned generators in both protein structure and molecular design are trained and scored against motifs that are artifacts of the pipeline — random residue crops, rigid bond-breaking heuristics, or proxy-evaluated structures — rather than motifs whose function has been experimentally verified. Because these synthetic motifs may not match the distribution of true functional motifs, reported success rates measure the ability to satisfy a surrogate objective, not the ability to produce functional designs. Until a curated set of validated motifs with defined success criteria exists, no model comparison (structure-based vs. sequence-based, heuristic vs. expert extraction) can be trusted, and designs cannot be credibly advanced toward wet-lab validation.
Why it matters
Trustworthy cross-method comparison on the same real motifs, and a direct path from in silico success to experimental validation for motif-conditioned generation in both biomolecular and small-molecule settings.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Assemble a benchmark of functionally annotated motifs from curated databases (e.g., catalytic site and binding-site annotations for proteins; experimentally characterized substructures and ring systems for small molecules), each with source structure, functional role, and extraction provenance. Measure: coverage (how many distinct motif families are represented) and distributional distance between these motifs and the synthetic/heuristic motifs currently used, quantified by a trained motif classifier or divergence metrics.
- 2
Re-evaluate representative motif-scaffolding and motif-generation models on the real-motif benchmark, holding the extraction protocol fixed. Measure: conditional success rate, motif RMSD / similarity rank, and diversity of successful designs, compared head-to-head against the same models on synthetic motifs to quantify how much current rankings depend on the surrogate.
- 3
Run a small prospective wet-lab or high-confidence structural-validation study (e.g., ESMFold/AlphaFold agreement on designed scaffolds, or DFT re-optimization of decoded molecules) on a subset of real-motif designs. Measure: agreement between in silico motif-satisfaction metrics and experimentally/predicted-structure-verified motif presence.
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
The set of experimentally validated motifs with enough structural context to condition generation turns out to be too small and too biased toward well-studied families to support a meaningful benchmark, so the distribution shift measured is itself an artifact of annotation availability.
Sub-problems
Topology-Adaptive Molecular Motif Decomposition for Non-Heuristic Graph Learning
Current motif-based molecular graph architectures rely either on rigid bond-breaking heuristics (such as BRICS or bridge-bond cuts) or expensive manual expert annotations to partition molecules into substructures. When applied to uncommon rings, macrocycles, and complex bridged topologies, heuristic decomposition creates fragmented or unphysical motifs that degrade downstream property prediction and molecular generation quality. Without a topology-aware or data-driven decomposition mechanism, motif-based frameworks remain brittle and restricted to standard chemical spaces that conform to handcrafted rules.
Evidence
Each paper's own statement of the limitation, verbatim.
- ProteinBench: A Holistic Evaluation of Protein Foundation ModelsICLR 2025
Structure-based and sequence-based methods were compared using proxies (e.g., ESMFold-predicted structures for sequence-only methods in motif scaffolding), adding evaluator-induced error
- MAGNet: Motif-Agnostic Generation of Molecules from ScaffoldsICLR 2025
Uncommon rings remain the hardest scaffolds to decode to correct motifs, with motif-generation similarity rank sometimes above 0 or 1
- Diffusion on Language Model Encodings for Protein Sequence GenerationICML 2025
Structure-based methods like RFDiffusion achieve higher overall motif-scaffolding success rates; DiMA only wins on unique success rate
- Representing Molecules as Random Walks Over Interpretable GrammarsICML 2024
Performance depends on quality of motif fragmentation; expert annotation workflows were required for the three evaluation datasets, and heuristic motifs are more outlier-sensitive (expert motifs gave 27% higher MAE than heuristic on HOPV)
- Floating Anchor Diffusion Model for Multi-motif ScaffoldingICML 2024
Only evaluated in silico; no wet-lab validation of designed structures, and motif RMSD is trivially 0 so biological motif presence is untested
Training uses virtual motifs (random 20-80 residue crops) rather than real functional motifs, which may not match the distribution of true motifs
- Hierarchical Generation of Molecular Graphs using Structural MotifsICML 2020
The motif extraction process relies on rigid heuristics (breaking bridge bonds), which may not be optimal for all molecular topologies.
Nearest existing work
- GeomMotif: A Benchmark for Arbitrary Geometric Preservation in Protein GenerationICLR 2026
- Floating Anchor Diffusion Model for Multi-motif ScaffoldingICML 2024
- SwitchCraft: A Programmatic Framework for Designing State-Switching ProteinsICML 2026
- Importance Weighted Expectation-Maximization for Protein Sequence DesignICML 2023
- Co-Generative De Novo Functional Protein DesignICML 2026
- NMA-tune: Generating Highly Designable and Dynamics Aware Protein BackbonesICML 2025
- ProtDBench: A Unified Benchmark of Protein Binder Design and EvaluationICML 2026
- PDFBench: A Benchmark for De Novo Protein Design from FunctionICML 2026
- FIDIA: Function-Informed Sequence Design via Inference-Aligned Policy OptimizationICML 2026
- Proteo-R1: Reasoning Foundation Models for De Novo Protein DesignICML 2026
- Protein Design with Dynamic Protein VocabularyNeurIPS 2025
- ProteinBench: A Holistic Evaluation of Protein Foundation ModelsICLR 2025
- Interpretability Driven Evolutionary Approach for the Design of Biological SequencesICML 2026
- Inference-time optimization for experiment-grounded protein ensemble generationICML 2026
- ProSpero: Active Learning for Robust Protein Design Beyond Wild-Type NeighborhoodsNeurIPS 2025
Related open problems
Protein & Biomolecular Sequence Modeling
Benchmarking the Fidelity of In Silico Protein Design Proxies Against Experimental Biophysical Assays
Generative biomolecular modeling currently assesses sequence designs almost entirely through computational surrogates—such as AlphaFold confidence (pLDDT/pTM), ESMFold self-consistency, and Rosetta energy terms—rather than physical assay measurements. Because these proxies serve as both optimization objectives and evaluation metrics, models risk overfitting to artifacts of structure-prediction heuristics rather than true biophysical viability. Without systematic evaluation against physical assays across sequence design methods, it remains unknown whether improvements on in silico leaderboards translate to viable expression, foldability, or binding in physical systems.
Protein & Biomolecular Sequence Modeling
A Systematic Robustness Study of Genomic Foundation Models Beyond the Human Reference Genome
Current genomic sequence models are predominantly pretrained and evaluated strictly on the single human reference genome (GRCh38/hg38). Consequently, it remains unknown whether representations trained on a single reference sequence maintain their predictive capabilities across distinct chromosomes, across individual genetic variation in population-scale cohorts, or across non-human species genomes. Without cross-setting robustness benchmarks, practitioners cannot determine whether model performance transfers to non-reference genomic contexts or degrades under distribution shifts.
Protein & Biomolecular Sequence Modeling
Benchmarking Surrogate-Trained Sequence Optimization Against Empirical Fitness Landscapes
Current machine learning models for protein and biomolecular sequence design are evaluated almost exclusively against learned surrogate oracles (such as CNN fitness predictors) rather than verified experimental outcomes. Because black-box optimizers can easily exploit surrogate artifacts and out-of-distribution prediction errors, reported in-silico fitness gains often fail to reflect actual biological activity. Without systematic evaluation across empirically ground-truthed landscapes, it is unknown which sequence design methods reliably find functional sequences versus which merely overfit to computational evaluator flaws.
Protein & Biomolecular Sequence Modeling
Benchmarking and Mitigating Model Sensitivity to Predicted Structure Quality in Protein Sequence Modeling
Current structure-conditioned protein sequence and representation models routinely filter out test proteins with predicted confidence scores below pLDDT 70 or assume near-experimental coordinate fidelity. Consequently, these models cannot reliably be applied to disordered proteins, orphan sequences, or hard-to-predict folds where high-confidence static predictions do not exist. Downstream pipelines experience unquantified error propagation when forced to consume predicted structures from AlphaFold2 or ESMFold, yet no systematic evaluation exists mapping how performance degrades across the full spectrum of structural prediction quality.