Open Problems

3D Point Clouds & LiDAR Perception

An Outdoor-Scale Benchmark for 3D Scene Understanding Methods Currently Validated Only on Indoor ScanNet-Derived Data

Scope to testOpen
Possible candidate · 3/5 runs39 papers report this30% from 2025+

Generated automatically from the limitations stated in 39 papers (CVPR, NeurIPS, ICCV), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.

The problem

Nearly the entire 3D perception, scene reconstruction, grounding, and question-answering stack is trained and evaluated on ScanNet-family indoor scans, so claims of scene understanding hold only within bounded, static, room-scale environments with known camera poses. As a result, no one can tell whether reported architectural advances—reasoning modules, open-vocabulary grounding, generalization arguments—transfer to outdoor, unbounded, or unsegmented scenes at all, or whether they simply exploit ScanNet's statistics. Progress reports that look cumulative within this benchmark family may be incommensurable with outdoor work that cannot reuse the same evaluation. Any paper today claiming 3D scene understanding is doing so on a single environmental regime.

Why it matters

For the first time, published 3D scene-understanding results become comparable across environmental regimes, allowing the field to identify which advances are architecture-driven versus ScanNet-driven. It also opens evaluation of embodied and open-vocabulary methods in genuinely open-world settings where their current claims are untestable.

Ways to approach it

Prior-work checks are free with an account. Results someone already ran are shown to everyone.

  1. 1

    Curate and release an outdoor counterpart to ScanNet in format compatibility: use existing outdoor LiDAR + RGBD sources (e.g.,SemanticKITTI,nuScenes, ONCE) re-processed into ScanNet-like scene-level structures (meshes, instance masks, camera poses, VQA-style questions), and measure transfer of representative ScanNet-trained models zero-shot versus fine-tuned on small outdoor subsets.

  2. 2

    Build a "ScanNet-style" dense outdoor capture pipeline (mobile scanner / drone photogrammetry of building exteriors, campus-scale environments) annotated with the same schema (poses, depth, instance and semantic labels, question-answer pairs), reporting performance of current indoor SOTA methods as a function of scene scale, view density, and depth quality.

  3. 3

    Construct a stress-test suite that interpolates between indoor and outdoor regimes (indoor-outdoor transitions, large atriums, open plazas) to isolate which failure axes—scale, sparsity, GPS-frame poses, dynamic content—drive the observed degradation, measuring per-axis breakdowns rather than one aggregate number.

Have a different approach?

Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.

Free · 3 checks per day

Why it might fail

Outdoor annotation of the depth and QA quality ScanNet provides may be so expensive or noisy that the benchmark ends up testing annotation robustness rather than method generalization, leaving the transfer gap unmeasurable.

Sub-problems

  • Cross-Domain Generalization and Robustness Benchmark for Sparse and Feed-Forward 3D Gaussian Splatting

    Current sparse and feed-forward 3D Gaussian Splatting methods are almost exclusively trained and evaluated on room-scale, inside-out indoor datasets such as ScanNet and Replica. Because no unified empirical study evaluates these models across varying scene scales and camera trajectories, it is unknown whether existing architectures generalize to outdoor, unbounded, or object-centric settings. Without systematic cross-domain benchmarking, researchers cannot distinguish whether reported gains reflect robust 3D representations or domain-specific overfitting to indoor geometry and fixed capture protocols.

  • Cross-Domain Generalization and Failure Modes in Neural Radiance and Inverse Rendering

    Current neural radiance field and implicit surface reconstruction methods are almost exclusively evaluated on small, hand-picked subsets of indoor benchmarks like ScanNet and synthetic Replica scenes. Because no standardized cross-domain evaluation exists, it is unknown whether reported geometric and radiometric reconstruction gains persist when moving to outdoor, large-scale, object-centric, or uncalibrated captures. Consequently, practitioners cannot predict when neural inverse rendering pipelines will catastrophically fail outside tight indoor bounds, and algorithmic progress risks overfitting to the geometric biases of a dozen indoor rooms.

  • Cross-Domain Generalization and Sensor Robustness Benchmark for 3D Visual Reasoning

    Current 3D visual reasoning and grounding models are trained and evaluated almost exclusively on indoor scene reconstructions from ScanNet and ScanNet++. Because no standardized evaluation exists spanning outdoor environments, sparse sensor regimes, and noisy unconstrained point clouds, the degradation profile of these models outside controlled indoor rooms remains unknown. Consequently, researchers cannot verify whether architectural advances reflect robust 3D spatial reasoning or overfitting to high-quality indoor RGB-D reconstruction artifacts.

  • Cross-Domain Robustness and Out-of-Distribution Generalization of 3D Visual Reasoning Models

    Current 3D visual reasoning and embodied question-answering systems are trained and evaluated almost exclusively on static, indoor, object-segmented datasets from the ScanNet family. Because existing benchmarks provide clean bounding-box proposals and closed indoor geometry, it is unknown whether current architectures degrade catastrophically on unsegmented point clouds, outdoor LiDAR scans, or dynamic environments. Without systematic cross-domain evaluation, researchers cannot determine whether 3D visual reasoning methods have learned generalizable spatial-semantic representations or are simply overfitting to ScanNet-specific scan artifacts and object priors.

Evidence

Each paper's own statement of the limitation, verbatim.

Show all 39 papers

Nearest existing work

Related open problems

3D Point Clouds & LiDAR Perception

Scope to testOpen

Benchmarking and Adapting Indoor-Trained 3D Perception and Generation Models to Outdoor and LiDAR Point Cloud Scenes

Nearly every 3D foundation-ish model — for detection, segmentation, registration, scene generation, and novel view synthesis — is trained and validated exclusively on indoor RGB-D or synthetic benchmarks (ScanNet, SUN RGB-D, 3D-FRONT, Habitat simulators), with explicit author acknowledgments that outdoor and LiDAR data is untested. This means the field's core 3D representations are calibrated to room-scale geometry (bounded depth, dense views, flat priors like indoor depth predictors and category vocabularies), and nobody can currently deploy them on autonomous-driving or aerial scans without unquantified degradation. The consequence is a silent incompatibility: outdoor researchers must either rebuild models from scratch or trust unmeasured transfer.

Possible candidate · 3/5 runs20 papers report this53% from 2025+

3D Point Clouds & LiDAR Perception

Barrier to removeOpen

Robust Multimodal 3D Perception Under Sensor Extrinsic and Intrinsic Miscalibration

State-of-the-art multimodal 3D perception architectures strictly require accurate extrinsic and intrinsic calibration between LiDAR and camera frames to construct point-to-pixel projections, frustum cost volumes, and cross-attention neighbourhoods. In real-world deployments, mechanical vibration, thermal expansion, and mounting shifts cause calibration parameters to drift, leading to severe geometric misalignment. Because current multimodal models hardcode these spatial transformations into their feature lifting layers, minor calibration errors cause catastrophic performance degradation or complete pipeline failure.

Possible candidate · 2/5 runs13 papers report this38% from 2025+

3D Point Clouds & LiDAR Perception

Barrier to removePartly addressed

Benchmark and Training Protocols for 3D Perception from Unpaired or Calibration-Free 2D-3D Data

Nearly every current 2D-3D fusion method assumes registered multi-view imagery, known intrinsics/extrinsics, and per-frame paired data — assumptions that hold only in curated lab captures. The consequence is that point clouds collected standalone (LiDAR scans, crowdsourced reconstructions, legacy datasets) cannot benefit from image features at all, and methods cannot be deployed where a fixed calibrated rig with synchronized cameras is impractical. Progress is blocked because no established protocol or dataset exists for learning 2D-3D correspondences from unpaired, uncalibrated, or single-view inputs.

Possible candidate · 2/5 runs11 papers report this17% from 2025+
Generated automatically, not curated by hand. Automated prior-work checks catch about a third of existing work, so treat this problem as a lead to investigate.