Video-Language Understanding & Temporal Grounding
Benchmarking and Mitigating Spatio-Temporal Grounding Degradation in Continuous Video-LLMs
Generated automatically from the limitations stated in 4 papers (ICLR, ICML, ECCV), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Current video-language models evaluate spatial grounding primarily on sparse frames (e.g., 0.5 FPS), leaving performance in continuous video unmeasured. When models attempt spatio-temporal grounding, performance degrades significantly under multi-point trajectory supervision (e.g., mIoU drops beyond 3 tracking points) and dynamic multi-resolution token sampling raises spatio-temporal error relative to plain spatial pooling. Because full spatio-temporal box annotation is prohibitively expensive, visual grounding in video remains unresolved with sub-10% IoU baselines.
Why it matters
Establishes reliable spatio-temporal localization capabilities in continuous video streams without requiring intractable manual frame-by-frame bounding box annotations.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Evaluate open-source Video-LLMs across varying evaluation frame rates (from sparse 0.5 FPS to dense continuous video) on existing spatio-temporal datasets (e.g., VidSTG, HC-STVG), measuring bounding box mIoU and trajectory stability across temporal sampling densities.
- 2
Controlled comparison of visual token reduction strategies (uniform spatial pooling vs. dynamic multi-resolution sampling) across dense frame sequences, measuring inter-frame representation consistency and spatio-temporal error.
- 3
Assess the effect of varying trajectory supervision point density ($P \in \{1, 3, 5, 10\}$) and trajectory filtering heuristics on temporal grounding accuracy under noisy pseudo-box annotations.
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
Dense spatio-temporal foundation models or general video diffusion representations might inherently resolve inter-frame spatial tracking, rendering empirical workarounds for autoregressive Video-LLM tokenizers obsolete.
Evidence
Each paper's own statement of the limitation, verbatim.
- SportR: A Benchmark for Multimodal Large Language Model Reasoning in SportsICLR 2026
Visual grounding remains essentially unsolved: even after SFT+RL, best IoU is 9.94%, and video grounding is excluded entirely because annotating spatio-temporal boxes was intractable
- TUMTraf VideoQA: Dataset and Benchmark for Unified Spatio-Temporal Video Understanding in Traffic ScenesICML 2025
Multi-resolution token sampling strategies degrade spatio-temporal grounding performance relative to plain spatial pooling (e.g., 7B ST error rises from 0.2382 to 0.2493-0.2496), suggesting dynamic resolution changes introduce inter-frame representation ambiguity
- PiTe: Pixel-Temporal Alignment for Large Video-Language ModelECCV 2024
Temporal grounding degrades with more tracking points (mIoU 19.4 at 1 point, 22.0 at 3, 19.9 at 5), showing sensitivity to noisy trajectory supervision and task-dependent optimal P
- TVQA+: Spatio-Temporal Grounding for Video Question AnsweringACL 2020
Spatial grounding is only supervised and evaluated on sparse frames (0.5 FPS) rather than continuous video
Nearest existing work
- GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal GroundingCVPR 2026
- Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video GroundingNeurIPS 2025
- VideoLoom: A Video Large Language Model for Joint Spatial-Temporal UnderstandingICML 2026
- SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMsCVPR 2026
- Divid: Disentangled Spatial-Temporal Modeling within LLMs for Temporally Grounded Video UnderstandingICLR 2026
- SARL-STG: A Spatially Aware Reinforcement Learning Framework for Refining MLLMs in Spatio-Temporal Video GroundingCVPR 2026
- VideoITG: Multimodal Video Understanding with Instructed Temporal GroundingCVPR 2026
- VideoGLaMM : A Large Multimodal Model for Pixel-Level Visual Grounding in VideosCVPR 2025
- OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal GroundingCVPR 2026
- What When and Where? Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated InstructionsCVPR 2024
- ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long VideosCVPR 2025
- Enhancing Visual Token Representations for Video Large Language Models via Training-free Spatial-Temporal Pooling and GriddingICLR 2026
- On the Consistency of Video Large Language Models in Temporal ComprehensionCVPR 2025
- DisTime: Distribution-based Time Representation for Video Large Language ModelsICCV 2025
- TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMsCVPR 2026