Open Problems

Protein & Biomolecular Sequence Modeling

A Systematic Robustness Study of Genomic Foundation Models Beyond the Human Reference Genome

Scope to testOpen
Possible candidate · 3/5 runs9 papers report this67% from 2025+

Generated automatically from the limitations stated in 9 papers (ICLR, ICML, NeurIPS), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.

The problem

Current genomic sequence models are predominantly pretrained and evaluated strictly on the single human reference genome (GRCh38/hg38). Consequently, it remains unknown whether representations trained on a single reference sequence maintain their predictive capabilities across distinct chromosomes, across individual genetic variation in population-scale cohorts, or across non-human species genomes. Without cross-setting robustness benchmarks, practitioners cannot determine whether model performance transfers to non-reference genomic contexts or degrades under distribution shifts.

Why it matters

Provides empirical performance boundaries and explicit robustness baselines for applying genomic sequence models outside the standard reference genome. Allows researchers to make evidence-based choices on model architectures and pretraining corpora for population genetics and comparative genomics tasks.

Ways to approach it

Prior-work checks are free with an account. Results someone already ran are shown to everyone.

  1. 1

    Evaluate open-source genomic models (e.g., Nucleotide Transformer, HyenaDNA, Evo) on standard downstream functional genomics tasks across distinct chromosome splits and across diverse non-human reference genomes (e.g., mouse, dog, yeast) under standardized probing and fine-tuning protocols to measure cross-species and cross-chromosome degradation.

  2. 2

    Construct a population-variation evaluation suite using multi-individual pangenome data (e.g., 1000 Genomes, Human Pangenome Reference Consortium) and measure prediction variance and accuracy shifts across individual haplotypes versus the reference genome.

  3. 3

    Train controlled, compute-matched genomic baseline models on single-reference, multi-genome, and multi-species pretraining corpora to directly measure the empirical effect of training corpus diversity on cross-distribution robustness.

Have a different approach?

Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.

Free · 3 checks per day

Why it might fail

Large model developers may simultaneously release multi-species and pangenome-pretrained models that standardize multi-distribution evaluation before this study concludes, or pretraining compute constraints may prevent rigorous controlled ablations against leading foundation models.

Evidence

Each paper's own statement of the limitation, verbatim.

Show all 9 papers

Nearest existing work

Related open problems

Protein & Biomolecular Sequence Modeling

Scope to testPartly addressed

Benchmarking the Fidelity of In Silico Protein Design Proxies Against Experimental Biophysical Assays

Generative biomolecular modeling currently assesses sequence designs almost entirely through computational surrogates—such as AlphaFold confidence (pLDDT/pTM), ESMFold self-consistency, and Rosetta energy terms—rather than physical assay measurements. Because these proxies serve as both optimization objectives and evaluation metrics, models risk overfitting to artifacts of structure-prediction heuristics rather than true biophysical viability. Without systematic evaluation against physical assays across sequence design methods, it remains unknown whether improvements on in silico leaderboards translate to viable expression, foldability, or binding in physical systems.

Possible candidate · 2/5 runs19 papers report this95% from 2025+

Protein & Biomolecular Sequence Modeling

Scope to testPartly addressed

Benchmarking Surrogate-Trained Sequence Optimization Against Empirical Fitness Landscapes

Current machine learning models for protein and biomolecular sequence design are evaluated almost exclusively against learned surrogate oracles (such as CNN fitness predictors) rather than verified experimental outcomes. Because black-box optimizers can easily exploit surrogate artifacts and out-of-distribution prediction errors, reported in-silico fitness gains often fail to reflect actual biological activity. Without systematic evaluation across empirically ground-truthed landscapes, it is unknown which sequence design methods reliably find functional sequences versus which merely overfit to computational evaluator flaws.

Strong candidate · 4/5 runs7 papers report this86% from 2025+

Protein & Biomolecular Sequence Modeling

Effect to explainPartly addressed

Benchmarking and Mitigating Model Sensitivity to Predicted Structure Quality in Protein Sequence Modeling

Current structure-conditioned protein sequence and representation models routinely filter out test proteins with predicted confidence scores below pLDDT 70 or assume near-experimental coordinate fidelity. Consequently, these models cannot reliably be applied to disordered proteins, orphan sequences, or hard-to-predict folds where high-confidence static predictions do not exist. Downstream pipelines experience unquantified error propagation when forced to consume predicted structures from AlphaFold2 or ESMFold, yet no systematic evaluation exists mapping how performance degrades across the full spectrum of structural prediction quality.

Possible candidate · 3/5 runs7 papers report this57% from 2025+

Protein & Biomolecular Sequence Modeling

Barrier to removePartly addressed

A Benchmark of Experimentally Validated Functional Motifs for Training and Evaluating Motif-Conditioned Generative Models

Motif-conditioned generators in both protein structure and molecular design are trained and scored against motifs that are artifacts of the pipeline — random residue crops, rigid bond-breaking heuristics, or proxy-evaluated structures — rather than motifs whose function has been experimentally verified. Because these synthetic motifs may not match the distribution of true functional motifs, reported success rates measure the ability to satisfy a surrogate objective, not the ability to produce functional designs. Until a curated set of validated motifs with defined success criteria exists, no model comparison (structure-based vs. sequence-based, heuristic vs. expert extraction) can be trusted, and designs cannot be credibly advanced toward wet-lab validation.

Possible candidate · 3/5 runs6 papers report this67% from 2025+
Generated automatically, not curated by hand. Automated prior-work checks catch about a third of existing work, so treat this problem as a lead to investigate.