Visual Reasoning & Embodied Navigation
Systematic Generalisation and Robustness Evaluation of Visual Reasoning Methods Across Distinct Perceptual Regimes
Generated automatically from the limitations stated in 4 papers (CVPR, ICML), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Current visual reasoning techniques are predominantly evaluated on static images containing localized, explicitly present visual evidence where language reasoning dominates. Consequently, it is unknown whether their performance improvements transfer to vision-dominant perceptual tasks, temporal video sequences, global visual synthesis problems, or queries involving implicit visual evidence. Practitioners cannot determine whether a method’s gains represent generalizable visual reasoning or narrow optimizations for localized static VQA. A systematic robustness evaluation across these untested dimensions is necessary to chart where existing methods hold and where they fail.
Why it matters
Provides the first comprehensive empirical map of where current visual reasoning methods generalize and where they break, identifying the exact operating boundaries that require new architectural research.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Multi-Regime Benchmark Suite: Assemble a standardized evaluation pipeline pairing static localized VQA with suites spanning vision-dominant QA (e.g., fine-grained diagram perception), temporal video QA, global synthesis tasks, and implicit-evidence queries. Measure accuracy drops and failure rates across regimes under identical model backbones.
- 2
Cross-Method Sensitivity Analysis: Evaluate leading open-weight visual reasoning architectures and reasoning-tuned multimodal models across each regime. Measure performance differentials, reliance on language priors versus visual inputs, and performance shifts when varying image resolution or temporal frame counts.
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
If rapid multimodal model scaling natively absorbs video comprehension and fine-grained perception before the study is completed, rendering the identified generalisation gaps obsolete across all evaluated architectures.
Evidence
Each paper's own statement of the limitation, verbatim.
- Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language ModelsCVPR 2026
Effectiveness is primarily demonstrated on VQA tasks with localized visual evidence and may fail on tasks requiring global image synthesis.
- DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful ReasoningICML 2026
Evaluation restricted to static images; video and temporal reasoning untested
- Interleaved-Modal Chain-of-ThoughtCVPR 2025
Restricted to reasoning tasks where the necessary visual evidence is explicitly present within the original input image.
- Bring Reason to Vision: Understanding Perception and Reasoning through Model MergingICML 2025
Fails to improve or slightly degrades performance on vision-only and vision-dominant tasks (such as figure QA) where visual recognition/perception rather than textual reasoning is the bottleneck.
Nearest existing work
- VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit CuesCVPR 2026
- VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual ReasoningICLR 2026
- Look on Demand: A Cognitive Scheduling Framework for Visual Evidence Acquisition in Multimodal ReasoningICML 2026
- VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMsCVPR 2026
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsICLR 2026
- Multimodal Learning and Reasoning for Visual Question AnsweringNeurIPS 2017
- VTQA: Visual Text Question Answering via Entity Alignment and Cross-Media ReasoningCVPR 2024
- On the General Value of Evidence, and Bilingual Scene-Text Visual Question AnsweringCVPR 2020
- VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMsNeurIPS 2025
- Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language ModelsCVPR 2026
- VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain KnowledgeICML 2026
- See Less, See Right: Bi-directional Perceptual Shaping For Multimodal ReasoningCVPR 2026
- VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?ICLR 2026
- Reasoning-Aligned Perception Decoupling for Scalable Multi-modal ReasoningICLR 2026
- R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal FormalizationICCV 2025