Open Problems

Domain Adaptation & Generalization

Systematic Evaluation of Data Augmentation Techniques Across Disparate Distribution Shift Types

Barrier to removePartly addressed
Possible candidate · 2/5 runs4 papers report this0% from 2025+

Generated automatically from the limitations stated in 4 papers (CVPR, ICML, ICCV), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.

The problem

Data augmentation methods developed for out-of-distribution robustness are overwhelmingly hand-engineered to simulate surface-level texture and appearance variations (e.g., color perturbations and synthetic corruptions). When these methods are deployed under other distribution shift types—such as subpopulation shifts, geometric transformations, or domain shifts in specialized fields like medical imaging—they frequently fail or degrade accuracy below standard Empirical Risk Minimization (ERM). Because augmentation techniques have not been systematically evaluated across these differing shift categories, practitioners have no reliable way to anticipate when a given augmentation strategy will harm downstream generalization.

Why it matters

Provides empirical boundaries identifying where existing augmentation strategies transfer or actively harm performance across distinct distribution shift regimes. This enables safer deployment in non-corruption domains and grounds future work on shift-aware augmentation search.

Ways to approach it

Prior-work checks are free with an account. Results someone already ran are shown to everyone.

  1. 1

    Conduct a cross-shift evaluation benchmarking standard augmentation families (such as AugMix, RandConv, AugMax, and Mixup variants) across standardized benchmarks representing subpopulation shift, geometric viewpoint shift, and real-world domain shift, measuring average accuracy, worst-group accuracy, and delta relative to ERM.

  2. 2

    Evaluate representation-level properties across these trained models to measure how specific augmentation classes alter model sensitivity to non-appearance features and spurious correlations across shift categories.

Have a different approach?

Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.

Free · 3 checks per day

Why it might fail

If broad distribution shift benchmark suites already incorporate exhaustive evaluations of all standard augmentation families, rendering a purely empirical robustness study redundant.

Evidence

Each paper's own statement of the limitation, verbatim.

Nearest existing work

Related open problems

Domain Adaptation & Generalization

Barrier to removePartly addressed

Domain-Invariant Learning Without Domain Labels or Target-Domain Data at Training Time

Current methods across these settings all condition on information that disappears in deployment: explicit domain labels for training samples, knowledge of the domain count, or a pool of unlabeled target-domain data available during training. When source domain identity is unknown, only one source is available, or target data cannot be collected before the model is fixed, these methods are simply unusable — not merely degraded. This blocks the realistic scenario where a single labeled dataset is all one has, and the test distribution is only discovered at inference time.

Strong candidate · 4/5 runs10 papers report this57% from 2025+

Domain Adaptation & Generalization

Scope to testPartly addressed

Cross-Architecture Robustness and Evaluation of Domain Adaptation in Object Detection

Current domain adaptation and domain generalization methods for object detection are almost exclusively evaluated on two-stage anchor-based architectures, primarily Faster R-CNN with ResNet backbones. Because these adaptation mechanisms often couple directly to region proposal networks, RoI pooling representations, or batch normalization layers, it is completely unknown whether reported domain adaptation gains transfer to one-stage detectors (e.g., FCOS, YOLO) or query-based transformers (e.g., DETR variants). Consequently, practitioners deploying modern transformer or anchor-free architectures have no empirical guidance on whether established domain adaptation methods function outside Faster R-CNN.

Possible candidate · 2/5 runs5 papers report this60% from 2025+

Domain Adaptation & Generalization

Barrier to removeOpen

Cold-Start Streaming Domain Adaptation and Discovery Without Labeled Base Sessions or Offline Pre-Passes

Current continual domain adaptation and streaming open-world discovery methods structurally depend on an offline initialization phase—either requiring a fully labeled source base dataset or an offline full-dataset inference pass to compute initial prototypes and representations. In truly dynamic environments where data arrives purely as an online stream and domain shifts occur from the first observation, collecting labeled base data or executing multi-pass offline initialization is impossible. Consequently, existing frameworks cannot be deployed in pure cold-start streaming regimes.

Possible candidate · 2/5 runs3 papers report this100% from 2025+
Generated automatically, not curated by hand. Automated prior-work checks catch about a third of existing work, so treat this problem as a lead to investigate.