Protein & Biomolecular Sequence Modeling
A Systematic Robustness Study of Genomic Foundation Models Beyond the Human Reference Genome
Generated automatically from the limitations stated in 9 papers (ICLR, ICML, NeurIPS), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Current genomic sequence models are predominantly pretrained and evaluated strictly on the single human reference genome (GRCh38/hg38). Consequently, it remains unknown whether representations trained on a single reference sequence maintain their predictive capabilities across distinct chromosomes, across individual genetic variation in population-scale cohorts, or across non-human species genomes. Without cross-setting robustness benchmarks, practitioners cannot determine whether model performance transfers to non-reference genomic contexts or degrades under distribution shifts.
Why it matters
Provides empirical performance boundaries and explicit robustness baselines for applying genomic sequence models outside the standard reference genome. Allows researchers to make evidence-based choices on model architectures and pretraining corpora for population genetics and comparative genomics tasks.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Evaluate open-source genomic models (e.g., Nucleotide Transformer, HyenaDNA, Evo) on standard downstream functional genomics tasks across distinct chromosome splits and across diverse non-human reference genomes (e.g., mouse, dog, yeast) under standardized probing and fine-tuning protocols to measure cross-species and cross-chromosome degradation.
- 2
Construct a population-variation evaluation suite using multi-individual pangenome data (e.g., 1000 Genomes, Human Pangenome Reference Consortium) and measure prediction variance and accuracy shifts across individual haplotypes versus the reference genome.
- 3
Train controlled, compute-matched genomic baseline models on single-reference, multi-genome, and multi-species pretraining corpora to directly measure the empirical effect of training corpus diversity on cross-distribution robustness.
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
Large model developers may simultaneously release multi-species and pangenome-pretrained models that standardize multi-distribution evaluation before this study concludes, or pretraining compute constraints may prevent rigorous controlled ablations against leading foundation models.
Evidence
Each paper's own statement of the limitation, verbatim.
- PatchDNA: A Flexible and Biologically-Informed Alternative to Tokenization for DNAICLR 2026
Pretraining was conducted solely on the human reference genome rather than multi-species or population-scale variation datasets.
- DNACHUNKER: Learnable Tokenization for DNA Language ModelsICML 2026
Pretraining was conducted on a fixed maximum context of 8,192 bp on the human reference genome (GRCh38/hg38), requiring downstream adaptation for multi-megabase tasks.
- BioToken and BioFM – Biologically-Informed Tokenization Enables Accurate and Efficient Genomic Foundation ModelsICML 2026
Pretrained on whole genomes from only 50 individuals, leaving scaling behaviors across large biobank-scale cohorts underexplored.
- A New Paradigm for Genome-wide DNA Methylation Prediction Without Methylation InputICLR 2026
Models trained on specific chromosomes exhibit poor zero-shot cross-chromosome generalization (e.g., chromosome 1-trained model drops from 0.5455 to 0.3025 MAS-PCC on chromosomes 1–3).
- Revisiting Convolution Architecture in the Realm of DNA Foundation ModelsICLR 2025
Pretraining used only the human reference genome (HG38); multispecies pretraining, which benefits other models, was not evaluated
Pretraining data was not restricted to functional genomic regions, which might otherwise improve performance
- JanusDNA: A Powerful Bi-directional Hybrid DNA Foundation ModelNeurIPS 2025
Pretraining restricted to the human reference genome only (HG38); no variants or non-human species, limiting data diversity
- Caduceus: Bi-Directional Equivariant Long-Range DNA Sequence ModelingICML 2024
All pre-training and evaluation is restricted to the human reference genome; generalization to other species' genomes is not demonstrated
- BEND: Benchmarking DNA Language Models on Biologically Meaningful TasksICLR 2024
Evaluation is frozen-embedding + lightweight CNN only, excluding possible fine-tuning gains, and restricted to the human genome, leaving cross-species generalization untested
Show all 9 papers
- HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide ResolutionNeurIPS 2023
Pretrained only on a single human reference genome, which may limit generalizability and introduce bias compared to multi-genome pretraining
Nearest existing work
- Genomic Foundationless Models: Pretraining Does Not Promise PerformanceICLR 2026
- SPACE: Your Genomic Profile Predictor is a Powerful DNA Foundation ModelICML 2025
- GENEB: Why Genomic Models Are Hard to CompareICML 2026
- Genome-Factory: A Library for Tuning, Deploying, and Interpreting Genomic Foundation ModelsICML 2026
- Omni-DNA: A Genomic Model Supporting Sequence Understanding, Long-context, and Textual AnnotationNeurIPS 2025
- The OMG dataset: An Open MetaGenomic corpus for mixed-modality genomic language modelingICLR 2025
- Predicting evolutionary rate as a pretraining task improves genome language model representationsICML 2026
- BioToken and BioFM – Biologically-Informed Tokenization Enables Accurate and Efficient Genomic Foundation ModelsICML 2026
- Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time RetrievalICML 2022
- HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide ResolutionNeurIPS 2023
- Scaling Laws and Architectural Frontiers in Metagenomic Foundation ModelsICML 2026
- VQDNA: Unleashing the Power of Vector Quantization for Multi-Species Genomic Sequence ModelingICML 2024
- Interpreting Genomic Language Models using Sparse AutoencodersICML 2026
- PoET: A generative model of protein families as sequences-of-sequencesNeurIPS 2023
- From Likelihood to Fitness: Improving Variant Effect Prediction in Protein and Genome Language ModelsNeurIPS 2025
Related open problems
Protein & Biomolecular Sequence Modeling
Benchmarking the Fidelity of In Silico Protein Design Proxies Against Experimental Biophysical Assays
Generative biomolecular modeling currently assesses sequence designs almost entirely through computational surrogates—such as AlphaFold confidence (pLDDT/pTM), ESMFold self-consistency, and Rosetta energy terms—rather than physical assay measurements. Because these proxies serve as both optimization objectives and evaluation metrics, models risk overfitting to artifacts of structure-prediction heuristics rather than true biophysical viability. Without systematic evaluation against physical assays across sequence design methods, it remains unknown whether improvements on in silico leaderboards translate to viable expression, foldability, or binding in physical systems.
Protein & Biomolecular Sequence Modeling
Benchmarking Surrogate-Trained Sequence Optimization Against Empirical Fitness Landscapes
Current machine learning models for protein and biomolecular sequence design are evaluated almost exclusively against learned surrogate oracles (such as CNN fitness predictors) rather than verified experimental outcomes. Because black-box optimizers can easily exploit surrogate artifacts and out-of-distribution prediction errors, reported in-silico fitness gains often fail to reflect actual biological activity. Without systematic evaluation across empirically ground-truthed landscapes, it is unknown which sequence design methods reliably find functional sequences versus which merely overfit to computational evaluator flaws.
Protein & Biomolecular Sequence Modeling
Benchmarking and Mitigating Model Sensitivity to Predicted Structure Quality in Protein Sequence Modeling
Current structure-conditioned protein sequence and representation models routinely filter out test proteins with predicted confidence scores below pLDDT 70 or assume near-experimental coordinate fidelity. Consequently, these models cannot reliably be applied to disordered proteins, orphan sequences, or hard-to-predict folds where high-confidence static predictions do not exist. Downstream pipelines experience unquantified error propagation when forced to consume predicted structures from AlphaFold2 or ESMFold, yet no systematic evaluation exists mapping how performance degrades across the full spectrum of structural prediction quality.
Protein & Biomolecular Sequence Modeling
A Benchmark of Experimentally Validated Functional Motifs for Training and Evaluating Motif-Conditioned Generative Models
Motif-conditioned generators in both protein structure and molecular design are trained and scored against motifs that are artifacts of the pipeline — random residue crops, rigid bond-breaking heuristics, or proxy-evaluated structures — rather than motifs whose function has been experimentally verified. Because these synthetic motifs may not match the distribution of true functional motifs, reported success rates measure the ability to satisfy a surrogate objective, not the ability to produce functional designs. Until a curated set of validated motifs with defined success criteria exists, no model comparison (structure-based vs. sequence-based, heuristic vs. expert extraction) can be trusted, and designs cannot be credibly advanced toward wet-lab validation.