Open Problems

Depth & Stereo Estimation

Systematic Cross-Domain Real-World Evaluation of Synthetic-Trained Depth and Stereo Models

Effect to explainPartly addressed
Strong candidate · 4/5 runs4 papers report this100% from 2025+

Generated automatically from the limitations stated in 4 papers (CVPR, ICML, ICCV), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.

The problem

Current depth and stereo estimation models are trained almost exclusively on synthetic datasets (such as TartanAir, VKITTI2, and OmniHouse) and are typically evaluated either on synthetic test splits or on narrow qualitative real-world samples. Because quantitative ground-truth evaluation across diverse physical environments is missing, practitioners cannot predict failure modes or transfer performance outside synthetic driving and indoor scenes. Without standardized multi-domain real-world evaluation, competing architectures cannot be fairly compared on real generalization capability.

Why it matters

Establishes the first standardized cross-domain transfer benchmark for synthetic-trained stereo and depth estimators, enabling principled model selection and diagnostic failure analysis on real-world sensor data.

Ways to approach it

Prior-work checks are free with an account. Results someone already ran are shown to everyone.

  1. 1

    Curate and standardize an evaluation harness across existing real-world datasets with ground-truth depth/disparity (e.g., Middlebury, ETH3D, KITTI raw, NYUv2, Booster), evaluating public checkpoints of synthetic-only stereo and depth models to measure absolute metric degradation against synthetic validation benchmarks.

  2. 2

    Construct a targeted real-world stress test using existing calibrated stereo rigs in unconstrained conditions (varying lighting, adverse weather, non-planar geometry, and non-driving scenes) to measure disparity error distributions (EPE, >3px error rate) across distinct environmental domains.

  3. 3

    Conduct an empirical ablation of synthetic pre-training mixtures, systematically varying synthetic scene diversity and measuring zero-shot real-world transfer metrics across all assembled real-world test domains.

Have a different approach?

Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.

Free · 3 checks per day

Why it might fail

The study would fail if existing public real-world datasets lack sufficient camera calibration or ground-truth quality to establish statistically significant performance differences among synthetic-trained models.

Evidence

Each paper's own statement of the limitation, verbatim.

Nearest existing work

Generated automatically, not curated by hand. Automated prior-work checks catch about a third of existing work, so treat this problem as a lead to investigate.