3D Point Clouds & LiDAR Perception
An Outdoor-Scale Benchmark for 3D Scene Understanding Methods Currently Validated Only on Indoor ScanNet-Derived Data
Generated automatically from the limitations stated in 39 papers (CVPR, NeurIPS, ICCV), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Nearly the entire 3D perception, scene reconstruction, grounding, and question-answering stack is trained and evaluated on ScanNet-family indoor scans, so claims of scene understanding hold only within bounded, static, room-scale environments with known camera poses. As a result, no one can tell whether reported architectural advances—reasoning modules, open-vocabulary grounding, generalization arguments—transfer to outdoor, unbounded, or unsegmented scenes at all, or whether they simply exploit ScanNet's statistics. Progress reports that look cumulative within this benchmark family may be incommensurable with outdoor work that cannot reuse the same evaluation. Any paper today claiming 3D scene understanding is doing so on a single environmental regime.
Why it matters
For the first time, published 3D scene-understanding results become comparable across environmental regimes, allowing the field to identify which advances are architecture-driven versus ScanNet-driven. It also opens evaluation of embodied and open-vocabulary methods in genuinely open-world settings where their current claims are untestable.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Curate and release an outdoor counterpart to ScanNet in format compatibility: use existing outdoor LiDAR + RGBD sources (e.g.,SemanticKITTI,nuScenes, ONCE) re-processed into ScanNet-like scene-level structures (meshes, instance masks, camera poses, VQA-style questions), and measure transfer of representative ScanNet-trained models zero-shot versus fine-tuned on small outdoor subsets.
- 2
Build a "ScanNet-style" dense outdoor capture pipeline (mobile scanner / drone photogrammetry of building exteriors, campus-scale environments) annotated with the same schema (poses, depth, instance and semantic labels, question-answer pairs), reporting performance of current indoor SOTA methods as a function of scene scale, view density, and depth quality.
- 3
Construct a stress-test suite that interpolates between indoor and outdoor regimes (indoor-outdoor transitions, large atriums, open plazas) to isolate which failure axes—scale, sparsity, GPS-frame poses, dynamic content—drive the observed degradation, measuring per-axis breakdowns rather than one aggregate number.
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
Outdoor annotation of the depth and QA quality ScanNet provides may be so expensive or noisy that the benchmark ends up testing annotation robustness rather than method generalization, leaving the transfer gap unmeasurable.
Sub-problems
Cross-Domain Generalization and Robustness Benchmark for Sparse and Feed-Forward 3D Gaussian Splatting
Current sparse and feed-forward 3D Gaussian Splatting methods are almost exclusively trained and evaluated on room-scale, inside-out indoor datasets such as ScanNet and Replica. Because no unified empirical study evaluates these models across varying scene scales and camera trajectories, it is unknown whether existing architectures generalize to outdoor, unbounded, or object-centric settings. Without systematic cross-domain benchmarking, researchers cannot distinguish whether reported gains reflect robust 3D representations or domain-specific overfitting to indoor geometry and fixed capture protocols.
Cross-Domain Generalization and Failure Modes in Neural Radiance and Inverse Rendering
Current neural radiance field and implicit surface reconstruction methods are almost exclusively evaluated on small, hand-picked subsets of indoor benchmarks like ScanNet and synthetic Replica scenes. Because no standardized cross-domain evaluation exists, it is unknown whether reported geometric and radiometric reconstruction gains persist when moving to outdoor, large-scale, object-centric, or uncalibrated captures. Consequently, practitioners cannot predict when neural inverse rendering pipelines will catastrophically fail outside tight indoor bounds, and algorithmic progress risks overfitting to the geometric biases of a dozen indoor rooms.
Cross-Domain Generalization and Sensor Robustness Benchmark for 3D Visual Reasoning
Current 3D visual reasoning and grounding models are trained and evaluated almost exclusively on indoor scene reconstructions from ScanNet and ScanNet++. Because no standardized evaluation exists spanning outdoor environments, sparse sensor regimes, and noisy unconstrained point clouds, the degradation profile of these models outside controlled indoor rooms remains unknown. Consequently, researchers cannot verify whether architectural advances reflect robust 3D spatial reasoning or overfitting to high-quality indoor RGB-D reconstruction artifacts.
Cross-Domain Robustness and Out-of-Distribution Generalization of 3D Visual Reasoning Models
Current 3D visual reasoning and embodied question-answering systems are trained and evaluated almost exclusively on static, indoor, object-segmented datasets from the ScanNet family. Because existing benchmarks provide clean bounding-box proposals and closed indoor geometry, it is unknown whether current architectures degrade catastrophically on unsegmented point clouds, outdoor LiDAR scans, or dynamic environments. Without systematic cross-domain evaluation, researchers cannot determine whether 3D visual reasoning methods have learned generalizable spatial-semantic representations or are simply overfitting to ScanNet-specific scan artifacts and object priors.
Evidence
Each paper's own statement of the limitation, verbatim.
- VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D ReconstructionCVPR 2026
Training and evaluation data are static indoor scenes (ScanNet, ScanNet++, ARKitScenes); dynamic and extreme environments are left to future work
- SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language ModelsICLR 2026
Dependency on structured 3D reconstructions (ScanNet) for high-quality ground truth, which may limit dataset scaling to unconstrained real-world environments.
- SimRecon: SimReady Compositional Scene Reconstruction from Real VideosCVPR 2026
Evaluation is limited to 20 indoor ScanNet scenes with a fixed stack of components (2DGS, SceneSplat, Rodin, Qwen2.5-VL); generalization to other scene types or VLM errors in relation inference is untested
- Monocular Open Vocabulary Occupancy Prediction for Indoor ScenesCVPR 2026
Evaluated on a single indoor benchmark (Occ-ScanNet) restricted to the camera view frustum.
- Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language UnderstandingCVPR 2026
Evaluated only on indoor ScanNet-derived benchmarks with object-centric frameworks; relies on pretrained object detector proposals, so applicability to outdoor scenes, unsegmented scenes, or point/video-based representations is untested
- Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMsICML 2026
Evaluation is confined to indoor ScanNet-family benchmarks with pre-segmented object-centric scenes; performance on unsegmented or outdoor/open-world 3D data is not assessed
- 3D Scene Assertion VerificationICML 2026
Method and benchmark are indoor ScanNet scenes; outdoor and predicted-proposal settings are only checked as supplementary robustness experiments
- ViGiL3D: A Linguistically Diverse Dataset for 3D Visual GroundingACL 2025
Evaluation is constrained to the static indoor scenes provided by the underlying ScanNet-based data
Show all 39 papers
- Taming Video Diffusion Prior with Scene-Grounding Guidance for 3D Gaussian Splatting from Sparse InputsCVPR 2025
Benchmarks are small and indoor-only: 6 synthetic Replica scenes and 4 ScanNet++ scenes with a custom 6-view 'inside-out' protocol, so generalization to outdoor/large scenes is untested
- Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMsNeurIPS 2025
Structured only for indoor scenes (6K annotated indoor rooms from ARKitScenes/ScanNet/ScanNet++); generalization to outdoor or open-world scenes is untested
- SpatialSplat: Efficient Semantic 3D from Sparse Unposed ImagesICCV 2025
Main evaluation limited to 2-view settings at 256×256 resolution on indoor ScanNet-style scenes; only a small 3-view extension is shown, so scaling to larger scenes or higher resolution is unverified
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial IntelligenceNeurIPS 2025
All training data derives from ScanNet scenes, so generalization to out-of-domain environments is untested; evaluation is confined to ScanNet-family spatial benchmarks (VSI-Bench, ScanQA, SQA3D).
- SPC-GS: Gaussian Splatting with Semantic-Prompt Consistency for Indoor Open-World Free-view Synthesis from Sparse InputsCVPR 2025
Evaluation restricted to indoor 'inside-out' scenes (Replica, ScanNet) with a fixed sparse protocol of 12 views encompassing the whole scene; outside-in, object-centric scenarios and other view counts are not tested in the main results
- SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language ModelsNeurIPS 2025
Training data is derived only from indoor ScanNet/ScanNet++ scenes; quantitative performance relies on depth quality (performance drops to 51.4% under δ=0.7 Gaussian depth noise and 46.7% with no depth)
- Robust Cross-modal Alignment Learning for Cross-Scene Spatial Reasoning and GroundingNeurIPS 2025
Evaluation is confined to indoor ScanNet-derived scenes; scalability beyond hundreds of scenes and robustness to noisy real-world point clouds are not tested
- Robin3D: Improving 3D Large Language Model via Robust Instruction TuningICCV 2025
Training and evaluation are restricted to ScanNet indoor scenes; no outdoor, LiDAR-based, or open-vocabulary evaluation is conducted (acknowledged by the authors)
- NeRFPrior: Learning Neural Radiance Field as a Prior for Indoor Scene ReconstructionCVPR 2025
Evaluated only on indoor scenes (ScanNet, Replica, BlendSwap); performance on outdoor or object-surrounding scenes beyond the MonoSDF generalization argument is not shown
- Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction TuningCVPR 2025
Evaluated only on ScanNetV2 indoor scenes with RGB-D frames and point-level instance segmentation annotations; no outdoor or object-centric 3D data
- Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language ModelCVPR 2025
Evaluation on indoor benchmark datasets (ScanNet, ScanNet200, ScanNet++, S3DIS); no evidence for outdoor or unstructured scenes
- Dynamic Focused Masking for Autoregressive Embodied Occupancy PredictionNeurIPS 2025
Only tested on ScanNet-derived indoor benchmarks; outdoor generalization mentioned but detailed results deferred to appendix
- DSPNet: Dual-vision Scene Perception for Robust 3D Question AnsweringCVPR 2025
Benchmarked only on indoor ScanNet-derived datasets; depth of reasoning module beyond 4 layers overfits due to limited dataset scale
- 3D Question Answering via only 2D Vision-Language ModelsICML 2025
Evaluated only on indoor ScanNet-derived benchmarks (ScanQA, SQA); outdoor or large-scale scene generalization is untested
- VersatileGaussian: Real-time Neural Rendering for Versatile Tasks using Gaussian SplattingECCV 2024
Evaluated only on Replica (synthetic) and ScanNet fragments with 40-75 sampled frames per scene at 480x640; generalization to large-scale or unbounded scenes is untested
- Surface Reconstruction for 3D Gaussian Splatting via Local Structural HintsECCV 2024
Evaluated only on indoor scenes (8 Replica, 4 ScanNet); robustness to large-scale outdoor or object-centric settings is not demonstrated
- Ray-Distance Volume Rendering for Neural Scene ReconstructionECCV 2024
Evaluation uses only 4 ScanNet scenes and 4 Tanks and Temples scenes (plus 8 Replica scenes), so broad-scene generalization is not demonstrated
- PanoRecon: Real-Time Panoptic 3D Reconstruction from Monocular VideoCVPR 2024
Evaluated only on ScanNetV2 indoor scenes; requires known camera poses from a SLAM system
- Large Spatial Model: End-to-end Unposed Images to Semantic 3DNeurIPS 2024
Trained only on indoor ScanNet++/ScanNet (1565 scenes, 3 days on 8 A100s); generalization beyond indoor room-scale scenes is not demonstrated (Replica is also simulated indoor)
- H2O-SDF: Two-phase Learning for 3D Indoor Reconstruction using Object Surface FieldsICLR 2024
Evaluated on only 12 ScanNet scenes (4 from prior work plus 8 randomly selected), so scene diversity is limited
- GSNeRF: Generalizable Semantic Neural Radiance Fields with Enhanced 3D Scene UnderstandingCVPR 2024
Evaluation restricted to small numbers of indoor scenes (10 unseen ScanNet scenes, 2 unseen Replica scenes), leaving outdoor or large-scale generalization untested
- GOV-NeSF: Generalizable Open-Vocabulary Neural Semantic FieldsCVPR 2024
Evaluated only on indoor room-scale scenes (ScanNet, Replica) with relatively sparse input views
- FreeSplat: Generalizable 3D Gaussian Splatting Towards Free View Synthesis of Indoor ScenesNeurIPS 2024
Trained and evaluated on indoor scenes (ScanNet/Replica); no evidence on outdoor or object-centric settings
- 3D Weakly Supervised Semantic Segmentation with 2D Vision-Language GuidanceECCV 2024
Evaluated only on indoor RGBD-scanned point cloud datasets (S3DIS, ScanNet)
- SurfelNeRF: Neural Surfel Radiance Fields for Online Photorealistic Reconstruction of Indoor ScenesCVPR 2023
Evaluation is limited to 8 test scenes on ScanNet with known camera parameters; generalization to other domains or uncalibrated input is not demonstrated
- Prototypical Variational Autoencoder for 3D Few-shot Object DetectionNeurIPS 2023
Evaluated only on indoor benchmarks (FS-ScanNet, FS-SUNRGBD) with VoteNet-style detection backbones; no evidence on outdoor LiDAR or other tasks
- PlaneRecTR: Unified Query Learning for 3D Plane Recovery from a Single ViewICCV 2023
Requires planar ground truth derived from ScanNet's piece-wise planar annotations; trained/evaluated only on indoor scenes at 256×192 resolution
- ObjectSDF++: Improved Object-Compositional Neural Implicit SurfacesICCV 2023
Evaluation covers only 8 Replica and 4 ScanNet scenes; the MLP variant is only comparable to (not clearly better than) MonoSDF on ScanNet, so scene-level gains partly stem from the multi-res grid architecture shared with prior work
- Meta Architecture for Point Cloud AnalysisCVPR 2023
Findings and PointMetaBase are validated mainly on indoor scene benchmarks (S3DIS, ScanNet V2, ScanObjectNN, ShapeNetPart); outdoor/large-scale benchmarks are not reported
- Geometric Viewpoint Learning with Hyper-Rays and Harmonics EncodingICCV 2023
Evaluated only on indoor ScanNet scenes with scene-boundary-constrained location grids; requires a point cloud input
- 3D Instance Segmentation via Enhanced Spatial and Semantic SupervisionICCV 2023
Evaluated only on indoor benchmark datasets (ScanNet V2 / ScanNet200); no outdoor or unstructured point cloud evidence
Nearest existing work
- From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMsCVPR 2026
- LiDAR-Net: A Real-scanned 3D Point Cloud Dataset for Indoor ScenesCVPR 2024
- ScanNet++: A High-Fidelity Dataset of 3D Indoor ScenesICCV 2023
- A Unified Framework for 3D Scene UnderstandingNeurIPS 2024
- Unveiling the Mist over 3D Vision-Language Understanding: Object-centric Evaluation with Chain-of-AnalysisCVPR 2025
- Habitat-Matterport 3D Semantics DatasetCVPR 2023
- SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene UnderstandingECCV 2024
- SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language PretrainingICCV 2025
- MultiScan: Scalable RGBD scanning for 3D environments with articulated objectsNeurIPS 2022
- EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AICVPR 2024
- 3D Semantic Segmentation in the Wild: Learning Generalized Models for Adverse-Condition Point CloudsCVPR 2023
- GPT4Scene: Understand 3D Scenes from Videos with Vision-Language ModelsICLR 2026
- Hierarchical 3D Scene Graphs Construction OutdoorsICCV 2025
- MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMsICCV 2025
- Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene UnderstandingNeurIPS 2024
Related open problems
3D Point Clouds & LiDAR Perception
Benchmarking and Adapting Indoor-Trained 3D Perception and Generation Models to Outdoor and LiDAR Point Cloud Scenes
Nearly every 3D foundation-ish model — for detection, segmentation, registration, scene generation, and novel view synthesis — is trained and validated exclusively on indoor RGB-D or synthetic benchmarks (ScanNet, SUN RGB-D, 3D-FRONT, Habitat simulators), with explicit author acknowledgments that outdoor and LiDAR data is untested. This means the field's core 3D representations are calibrated to room-scale geometry (bounded depth, dense views, flat priors like indoor depth predictors and category vocabularies), and nobody can currently deploy them on autonomous-driving or aerial scans without unquantified degradation. The consequence is a silent incompatibility: outdoor researchers must either rebuild models from scratch or trust unmeasured transfer.
3D Point Clouds & LiDAR Perception
Robust Multimodal 3D Perception Under Sensor Extrinsic and Intrinsic Miscalibration
State-of-the-art multimodal 3D perception architectures strictly require accurate extrinsic and intrinsic calibration between LiDAR and camera frames to construct point-to-pixel projections, frustum cost volumes, and cross-attention neighbourhoods. In real-world deployments, mechanical vibration, thermal expansion, and mounting shifts cause calibration parameters to drift, leading to severe geometric misalignment. Because current multimodal models hardcode these spatial transformations into their feature lifting layers, minor calibration errors cause catastrophic performance degradation or complete pipeline failure.
3D Point Clouds & LiDAR Perception
Benchmark and Training Protocols for 3D Perception from Unpaired or Calibration-Free 2D-3D Data
Nearly every current 2D-3D fusion method assumes registered multi-view imagery, known intrinsics/extrinsics, and per-frame paired data — assumptions that hold only in curated lab captures. The consequence is that point clouds collected standalone (LiDAR scans, crowdsourced reconstructions, legacy datasets) cannot benefit from image features at all, and methods cannot be deployed where a fixed calibrated rig with synchronized cameras is impractical. Progress is blocked because no established protocol or dataset exists for learning 2D-3D correspondences from unpaired, uncalibrated, or single-view inputs.