Molecular Modeling & Drug Design
Benchmarks and Training Protocols for Size-and-Structure Transfer in Molecular ML
Generated automatically from the limitations stated in 14 papers (ICML, NeurIPS, ICLR), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Molecular machine-learning models are trained and evaluated within narrow slices of chemical space — small, drug-like, scaffold-homogeneous molecules averaging ~44 atoms — so any deployment beyond that slice degrades sharply or requires full retraining per system. This blocks the field's central promise: screening large protein-ligand complexes, extrapolating to novel chemistries, and reusing models across temperatures, states, or unseen systems. Without a measurement standard for out-of-distribution molecular generalization, papers cannot even quantify how far their models actually transfer.
Why it matters
Models that a lab can apply to a new molecular system, size regime, or thermodynamic condition without retraining, plus an honest yardstick for claiming transferability.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Build a standardized size/chemistry-shift benchmark suite: split existing datasets (GEOM, MD22, PubChem) by atom count, scaffold, and functional-group novelty, with held-out regimes far from training support. Measure zero-shot accuracy curves as a function of distance from training distribution, for at least five published architectures.
- 2
Test whether unlabeled structures from the target regime help: pretrain on large/chemically diverse molecules without labels, then fine-tune on the small labeled sets, measuring how much the zero-shot error (e.g., ε_Δ on held-out sizes) drops relative to training from scratch.
- 3
Probe architectural capacity limits directly: vary the number of cross-attention tokens and replace fully-connected message passing with local neighborhoods, measuring reconstruction and prediction error as a function of molecule size to isolate whether bottleneck (information) or scaling (compute) dominates degradation.
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
Degradation may turn out to be irreducibly data-limited — no benchmark or protocol fixes it, and all architectures fail similarly, yielding a purely negative result. Or the benchmark splits may be dominated by confounds (e.g., conformer quality) rather than true distribution shift.
Sub-problems
Systematic Cross-Regime Generalization and Robustness Benchmarking in Molecular Machine Learning
Current molecular machine learning models are predominantly validated on standard scaffold or assay splits, obscuring severe performance degradation when applied out-of-distribution. Across diverse tasks, existing models fail to generalize across underrepresented functional groups, varying thermodynamic variables (temperature and pressure), temporal bioassay splits, and unseen meta-stable state transitions, frequently reverting to performance no better than trivial frequent-hitter baselines. Because no single study has systematically evaluated architectures across these separate operating regimes, practitioners cannot distinguish whether a model learns transferable chemical representations or merely memorizes dataset-specific artifacts. This forces downstream discovery pipelines to retrain models per system without clear failure criteria.
Evidence
Each paper's own statement of the limitation, verbatim.
- VecMol: Vector-Field Representations for 3D Molecule GenerationICML 2026
Autoencoder reconstruction capacity forms the primary performance bottleneck on larger, complex molecules.
- SymSpectra: Symmetric Information Bottleneck Framework for Molecular Structure Recognition under Imbalanced SettingsICML 2026
Generalization limited by training data diversity: lower accuracy on underrepresented molecular scaffolds and rare/novel functional groups
- Excited Pfaffians: Generalized Neural Wave Functions Across Structure and StateICML 2026
Accuracy may degrade for distinct, uncorrelated compounds in multi-structure training; lossless generalization across chemical compounds remains open
- CARD: Coarse-to-fine Autoregressive Modeling with Radix-based Decomposition for Transferable Free Energy EstimationICML 2026
Runtime scales quadratically with atom count due to the unoptimized transformer; evaluation is limited to drug-like small molecules, not protein-ligand complexes
- Transition Path Sampling with Improved Off-Policy Training of Diffusion Path SamplersICLR 2025
Does not generalize across unseen pairs of meta-stable states or to different molecular systems — a new model must be trained per system and state pair
- Template-Guided 3D Molecular Pose Generation via Flow Matching and Differentiable OptimizationNeurIPS 2025
Performance degrades as template similarity decreases and molecular complexity (atom count, rotatable bonds) increases, though it remains above baselines
- SemlaFlow – Efficient 3D Molecular Generation with Latent Attention and Equivariant Flow MatchingAISTATS 2025
Still uses fully-connected message passing, limiting scalability to larger molecular systems; also discards molecules with >72 atoms (~1%) from GEOM Drugs training
- Sampling 3D Molecular Conformers with Diffusion TransformersNeurIPS 2025
Attention scales quadratically in atom count, making large molecular systems computationally expensive; evaluation is restricted to small/medium molecules (training molecules average 44 atoms)
Show all 14 papers
- Enhancing the Scalability and Applicability of Kohn-Sham Hamiltonians for Molecular SystemsICLR 2025
Generalization beyond training atom counts is only demonstrated on a homogeneous series (elongated alkanes up to 182 atoms) in the test set, not diverse out-of-distribution chemistries
- Self-Consistency Training for Density-Functional-Theory Hamiltonian PredictionICML 2024
Generalization across a large molecular-size gap remains poor even with the method (zero-shot ϵ_Δ ~10^4–10^5 µEh on MD22), so unlabeled structures from the target size regime are still needed for fine-tuning
- LLaMo: Large Language Model-based Molecular Graph AssistantNeurIPS 2024
The use of fixed-size learnable tokens for cross-attention may create an information bottleneck for exceptionally large or complex molecules
- Learning Invariant Molecular Representation in Latent Discrete SpaceNeurIPS 2023
Approach is validated only on molecular graph benchmarks with scaffold/size/assay shifts; no evidence it transfers to other non-Euclidean OOD settings
- Implicit Transfer Operator Learning: Multiple Time-Resolution Models for Molecular DynamicsNeurIPS 2023
No generalization across chemical space or thermodynamic variables (temperature, pressure): each model is system-specific and must be retrained per molecule and ensemble
- Enhancing Activity Prediction Models in Drug Discovery with the Ability to Understand Human LanguageICML 2023
Predictive ability degrades outside the chemical/bioassay space of training data and inherits biases present in PubChem/ChEMBL; on the time_a_c PubChem split the frequent-hitter baseline was not significantly outperformed
Nearest existing work
- On the Scalability of GNNs for Molecular GraphsNeurIPS 2024
- Pushing the Limits of All-Atom Geometric Graph Neural Networks: Pre-Training, Scaling, and Zero-Shot TransferICLR 2025
- Exploring Molecular Pretraining Model at ScaleNeurIPS 2024
- Understanding the Limitations of Deep Models for Molecular property prediction: Insights and SolutionsNeurIPS 2023
- Transfer learning for atomistic simulations using GNNs and kernel mean embeddingsNeurIPS 2023
- The Importance of Being Scalable: Improving the Speed and Accuracy of Neural Network Interatomic Potentials Across Chemical DomainsNeurIPS 2024
- UniMatch: Universal Matching from Atom to Task for Few-Shot Drug DiscoveryICLR 2025
- Bridging the Gap Between Cross-Domain Theory and Practical Application: A Case Study on Molecular DissolutionNeurIPS 2025
- Learning Over Molecular Conformer Ensembles: Datasets and BenchmarksICLR 2024
- Enhancing Activity Prediction Models in Drug Discovery with the Ability to Understand Human LanguageICML 2023
- What Ails Generative Structure-based Drug Design: Expressivity is Too Little or Too Much?AISTATS 2025
- CIDD: Collaborative Intelligence for Structure-Based Drug Design Empowered by LLMsNeurIPS 2025
- TerraBind: Fast and Accurate Binding Affinity Prediction through Coarse Structural RepresentationsICML 2026
- Physical Consistency Bridges Heterogeneous Data in Molecular Multi-Task LearningNeurIPS 2024
- NExT-Mol: 3D Diffusion Meets 1D Language Modeling for 3D Molecule GenerationICLR 2025
Related open problems
Molecular Modeling & Drug Design
Lossless Generalization Across Molecular Scaffold Diversity: Closing the Accuracy Gap on Unseen Compounds
Molecular generative and predictive models currently trade off reconstruction fidelity, chemical validity, and generalization: models that reconstruct accurately lose validity on novel chemistry, and models that stay valid degrade sharply on scaffolds, functional groups, and complex molecules absent or rare in training. This means any model's utility on genuinely new chemistry — the cases that matter most in drug discovery — is unpredictable and unmeasured per scaffold. Structurally, the field cannot yet certify that a model generalizes to uncorrelated compounds rather than interpolating within its training distribution.
Molecular Modeling & Drug Design
Benchmarking Cross-System and Cross-Ensemble Generalization in Molecular Generative and Transition Models
Current deep learning models for molecular transitions, ensemble generation, and structure-based drug design are routinely trained and evaluated on isolated systems, fixed thermodynamic conditions, or homogeneous configuration spaces. When deployed on unseen molecular systems, new temperature/pressure regimes, novel pairs of meta-stable states, or shifted configuration spaces (such as moving from isolated ligands to pocket environments), these models require costly de novo retraining. Because prior studies only evaluate models within their narrow training distribution, the extent to which existing architectures can transfer across these varying physical and structural conditions remains unknown.