Open Problems

Molecular Modeling & Drug Design

Benchmarking Cross-System and Cross-Ensemble Generalization in Molecular Generative and Transition Models

Scope to testPartly addressed
Weak candidate · 1/3 runs3 papers report this33% from 2025+

Generated automatically from the limitations stated in 3 papers (NeurIPS, ICLR), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.

The problem

Current deep learning models for molecular transitions, ensemble generation, and structure-based drug design are routinely trained and evaluated on isolated systems, fixed thermodynamic conditions, or homogeneous configuration spaces. When deployed on unseen molecular systems, new temperature/pressure regimes, novel pairs of meta-stable states, or shifted configuration spaces (such as moving from isolated ligands to pocket environments), these models require costly de novo retraining. Because prior studies only evaluate models within their narrow training distribution, the extent to which existing architectures can transfer across these varying physical and structural conditions remains unknown.

Why it matters

Establishes the first systematic empirical boundaries for zero-shot and few-shot transfer of molecular generative and dynamic models across systems, states, and thermodynamic conditions.

Ways to approach it

Prior-work checks are free with an account. Results someone already ran are shown to everyone.

  1. 1

    Construct a standardized cross-condition benchmark using existing molecular dynamics trajectories across varying thermodynamic variables ($T, P$) and diverse molecular chemistries, measuring performance degradation (energy error, transition path likelihood, and sample validity) when models trained on one condition are evaluated zero-shot on another.

  2. 2

    Evaluate transferability across unseen pairs of meta-stable states and multi-body complexes by testing whether representation fine-tuning or zero-shot conditioning in flow/diffusion models can recover transition paths without full retraining.

  3. 3

    Systematically test whether unconditional ligand models can be guided via test-time conditioning into pocket spaces without joint configuration training, quantifying spatial clash rates and binding affinity against joint-trained baselines.

Have a different approach?

Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.

Free · 3 checks per day

Why it might fail

If the underlying physics across distinct thermodynamic ensembles and complex multi-body spaces creates fundamental distribution shifts so severe that zero-shot transfer is mathematically impossible without system-specific sampling data, rendering comparative benchmarking purely negative.

Evidence

Each paper's own statement of the limitation, verbatim.

Nearest existing work

Related open problems

Molecular Modeling & Drug Design

Scope to testOpen

Benchmarks and Training Protocols for Size-and-Structure Transfer in Molecular ML

Molecular machine-learning models are trained and evaluated within narrow slices of chemical space — small, drug-like, scaffold-homogeneous molecules averaging ~44 atoms — so any deployment beyond that slice degrades sharply or requires full retraining per system. This blocks the field's central promise: screening large protein-ligand complexes, extrapolating to novel chemistries, and reusing models across temperatures, states, or unseen systems. Without a measurement standard for out-of-distribution molecular generalization, papers cannot even quantify how far their models actually transfer.

Possible candidate · 3/5 runs14 papers report this75% from 2025+

Molecular Modeling & Drug Design

Effect to explainPartly addressed

Lossless Generalization Across Molecular Scaffold Diversity: Closing the Accuracy Gap on Unseen Compounds

Molecular generative and predictive models currently trade off reconstruction fidelity, chemical validity, and generalization: models that reconstruct accurately lose validity on novel chemistry, and models that stay valid degrade sharply on scaffolds, functional groups, and complex molecules absent or rare in training. This means any model's utility on genuinely new chemistry — the cases that matter most in drug discovery — is unpredictable and unmeasured per scaffold. Structurally, the field cannot yet certify that a model generalizes to uncorrelated compounds rather than interpolating within its training distribution.

Possible candidate · 1/2 runs6 papers report this100% from 2025+
Generated automatically, not curated by hand. Automated prior-work checks catch about a third of existing work, so treat this problem as a lead to investigate.