Open Problems

Domain Adaptation & Generalization

Domain-Invariant Learning Without Domain Labels or Target-Domain Data at Training Time

Barrier to removePartly addressed
Strong candidate · 4/5 runs10 papers report this57% from 2025+

Generated automatically from the limitations stated in 10 papers (NeurIPS, ICCV, ICLR), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.

The problem

Current methods across these settings all condition on information that disappears in deployment: explicit domain labels for training samples, knowledge of the domain count, or a pool of unlabeled target-domain data available during training. When source domain identity is unknown, only one source is available, or target data cannot be collected before the model is fixed, these methods are simply unusable — not merely degraded. This blocks the realistic scenario where a single labeled dataset is all one has, and the test distribution is only discovered at inference time.

Why it matters

Domain generalization becomes applicable when domain identity is unannotated, the domain count is unknown, or only one labeled source exists — the common case in practice rather than the benchmark case.

Ways to approach it

Prior-work checks are free with an account. Results someone already ran are shown to everyone.

  1. 1

    Infer domain structure automatically: cluster training samples using model-internal disagreement or feature-space density to produce pseudo-domain partitions, then run domain-invariant learning over the discovered partitions with no explicit labels. Measure whether discovered partitions recover ground-truth domains (adjusted mutual information) and whether invariance over inferred partitions matches invariance over true labels on standard benchmarks, including single-source splits.

  2. 2

    Train a domain-label-free objective that requires no per-domain decoders or domain count — e.g., invariance enforced via sample-level reweighting or contrasts that do not scale with the number of partitions — and measure the point at which performance degrades as the number of latent domains grows from 1 to 50.

  3. 3

    Test all of the above under a strict protocol: no unlabeled target data, no domain labels, single or unknown source count, reporting the gap to methods that are granted each privilege, to quantify exactly which piece of supervision buys how much.

Have a different approach?

Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.

Free · 3 checks per day

Why it might fail

The inferred partitions may be too noisy or orthogonal to true domain structure, so invariance over them buys nothing beyond a plain single-model baseline — in which case domain labels turn out to carry information that cannot be recovered from the data alone.

Sub-problems

  • Few-Shot and Zero-Shot Learning Under Extreme Unlabeled Target Data Scarcity and Mismatch

    Current semi-supervised and transductive few-shot methods depend on access to thousands of unlabeled in-domain target samples during training or adaptation. In practical deployment scenarios, such as rare disease diagnostics or proprietary edge deployments, collecting even moderate pools of unlabeled target data is impossible or violates privacy constraints. Consequently, existing models degrade unpredictably or fail entirely when evaluated in truly data-starved or severely mismatched target regimes.

Evidence

Each paper's own statement of the limitation, verbatim.

Show all 10 papers

Nearest existing work

Related open problems

Domain Adaptation & Generalization

Scope to testPartly addressed

Cross-Architecture Robustness and Evaluation of Domain Adaptation in Object Detection

Current domain adaptation and domain generalization methods for object detection are almost exclusively evaluated on two-stage anchor-based architectures, primarily Faster R-CNN with ResNet backbones. Because these adaptation mechanisms often couple directly to region proposal networks, RoI pooling representations, or batch normalization layers, it is completely unknown whether reported domain adaptation gains transfer to one-stage detectors (e.g., FCOS, YOLO) or query-based transformers (e.g., DETR variants). Consequently, practitioners deploying modern transformer or anchor-free architectures have no empirical guidance on whether established domain adaptation methods function outside Faster R-CNN.

Possible candidate · 2/5 runs5 papers report this60% from 2025+

Domain Adaptation & Generalization

Barrier to removePartly addressed

Systematic Evaluation of Data Augmentation Techniques Across Disparate Distribution Shift Types

Data augmentation methods developed for out-of-distribution robustness are overwhelmingly hand-engineered to simulate surface-level texture and appearance variations (e.g., color perturbations and synthetic corruptions). When these methods are deployed under other distribution shift types—such as subpopulation shifts, geometric transformations, or domain shifts in specialized fields like medical imaging—they frequently fail or degrade accuracy below standard Empirical Risk Minimization (ERM). Because augmentation techniques have not been systematically evaluated across these differing shift categories, practitioners have no reliable way to anticipate when a given augmentation strategy will harm downstream generalization.

Possible candidate · 2/5 runs4 papers report this0% from 2025+

Domain Adaptation & Generalization

Barrier to removeOpen

Cold-Start Streaming Domain Adaptation and Discovery Without Labeled Base Sessions or Offline Pre-Passes

Current continual domain adaptation and streaming open-world discovery methods structurally depend on an offline initialization phase—either requiring a fully labeled source base dataset or an offline full-dataset inference pass to compute initial prototypes and representations. In truly dynamic environments where data arrives purely as an online stream and domain shifts occur from the first observation, collecting labeled base data or executing multi-pass offline initialization is impossible. Consequently, existing frameworks cannot be deployed in pure cold-start streaming regimes.

Possible candidate · 2/5 runs3 papers report this100% from 2025+
Generated automatically, not curated by hand. Automated prior-work checks catch about a third of existing work, so treat this problem as a lead to investigate.