Depth & Stereo Estimation
Systematic Cross-Domain Real-World Evaluation of Synthetic-Trained Depth and Stereo Models
Generated automatically from the limitations stated in 4 papers (CVPR, ICML, ICCV), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Current depth and stereo estimation models are trained almost exclusively on synthetic datasets (such as TartanAir, VKITTI2, and OmniHouse) and are typically evaluated either on synthetic test splits or on narrow qualitative real-world samples. Because quantitative ground-truth evaluation across diverse physical environments is missing, practitioners cannot predict failure modes or transfer performance outside synthetic driving and indoor scenes. Without standardized multi-domain real-world evaluation, competing architectures cannot be fairly compared on real generalization capability.
Why it matters
Establishes the first standardized cross-domain transfer benchmark for synthetic-trained stereo and depth estimators, enabling principled model selection and diagnostic failure analysis on real-world sensor data.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Curate and standardize an evaluation harness across existing real-world datasets with ground-truth depth/disparity (e.g., Middlebury, ETH3D, KITTI raw, NYUv2, Booster), evaluating public checkpoints of synthetic-only stereo and depth models to measure absolute metric degradation against synthetic validation benchmarks.
- 2
Construct a targeted real-world stress test using existing calibrated stereo rigs in unconstrained conditions (varying lighting, adverse weather, non-planar geometry, and non-driving scenes) to measure disparity error distributions (EPE, >3px error rate) across distinct environmental domains.
- 3
Conduct an empirical ablation of synthetic pre-training mixtures, systematically varying synthetic scene diversity and measuring zero-shot real-world transfer metrics across all assembled real-world test domains.
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
The study would fail if existing public real-world datasets lack sufficient camera calibration or ground-truth quality to establish statistically significant performance differences among synthetic-trained models.
Evidence
Each paper's own statement of the limitation, verbatim.
- PointDiT: Pixel-Space Diffusion for Monocular Geometry EstimationICML 2026
Trained exclusively on synthetic data; performance scales with dataset size and diversity, so real-world coverage may be limited
- LiDAR Prompted Spatio-Temporal Multi-View Stereo for Autonomous DrivingCVPR 2026
Trained entirely on synthetic datasets (TartanAir, TartanGround, VKITTI2, MVS-Synth), so generalization to unseen real domains is zero-shot only and not guaranteed for domain types far from driving
- 240FPS Stereo Vision from Monocular Mixed SpikesCVPR 2026
Training relies on synthetic data (TartanAir/KITTI simulation) since no public dataset exists for this setup; real-world evaluation is qualitative only
- MDP-Omni: Parameter-free Multimodal Depth Prior-based Sampling for Omnidirectional Stereo MatchingICCV 2025
Only evaluated on synthetic datasets (OmniThings, OmniHouse, driving variants); real-world evaluation is qualitative only on a few samples
Nearest existing work
- Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching Even Where Either Stereo or Mono FailCVPR 2025
- Robust Synthetic-to-Real Transfer for Stereo MatchingCVPR 2024
- Learning Monocular Depth by Distilling Cross-domain Stereo NetworksECCV 2018
- Synthetic-to-Real Self-supervised Robust Depth Estimation via Learning with Motion and Structure PriorsCVPR 2025
- StereoGAN: Bridging Synthetic-to-Real Domain Gap by Joint Optimization of Domain Translation and Stereo MatchingCVPR 2020
- T2Net: Synthetic-to-Realistic Translation for Solving Single-Image Depth Estimation TasksECCV 2018
- What Makes Good Synthetic Training Data for Zero-Shot Stereo Matching?CVPR 2026
- Domain Generalized Stereo Matching via Hierarchical Visual TransformationCVPR 2023
- FoundationStereo: Zero-Shot Stereo MatchingCVPR 2025
- Depth Anything V2NeurIPS 2024
- Unsupervised Adaptation for Deep StereoICCV 2017
- Learning to Adapt for StereoCVPR 2019
- Learning Stereo from Single ImagesECCV 2020
- ZeroStereo: Zero-shot Stereo Matching from Single ImagesICCV 2025
- Depth Any Panoramas: A Foundation Model for Panoramic Depth EstimationCVPR 2026