Open Problems

Vision-Language Understanding & Captioning

Benchmarking the Cross-Model and Cross-Failure Generalization of Vision-Language Hallucination Mitigation

Scope to testPartly addressed
Possible candidate · 2/5 runs7 papers report this86% from 2025+

Generated automatically from the limitations stated in 7 papers (ICLR, EMNLP, ACL), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.

The problem

Current vision-language hallucination mitigations are evaluated under fragmented and narrow experimental conditions—often restricted to nominal entity metrics like CHAIR, small image sets, and single model checkpoints. Because existing techniques target specific decoding-time attention dynamics without assessing transfers across diverse architectures, hyperparameter shifts, or non-nominal hallucination types (such as relation errors or missing pretraining facts), their actual robustness is unverified. Consequently, researchers cannot determine whether observed gains reflect genuine mitigation of multimodal hallucinations or hyperparameter overfitting to specific evaluation setups.

Why it matters

Provides a clear empirical map of where existing vision-language mitigation methods generalize and where they fail across model families and error types. Enables the field to benchmark future interventions against distinct hallucination failure modes rather than narrow entity-mention subsets.

Ways to approach it

Prior-work checks are free with an account. Results someone already ran are shown to everyone.

  1. 1

    Conduct a cross-architecture evaluation of standard decoding-time mitigation algorithms across multiple open-weight vision-language models (e.g., LLaVA, InstructBLIP, Qwen-VL) under varied hyperparameter configurations, measuring performance on both nominal (CHAIR) and non-nominal hallucination benchmarks.

  2. 2

    Construct a multi-source diagnostic evaluation separating hallucinations driven by visual attention misalignments from those rooted in pretraining data gaps or missing visual evidence, measuring the exact performance boundaries of decoding-time versus retrieval-augmented mitigations.

Have a different approach?

Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.

Free · 3 checks per day

Why it might fail

Rapid architectural convergence to proprietary multimodal APIs could make internal attention- and decoding-level interventions obsolete, or unified large-scale benchmarks may be concurrently released by major industrial labs.

Evidence

Each paper's own statement of the limitation, verbatim.

Nearest existing work

Related open problems

Vision-Language Understanding & Captioning

Effect to explainOpen

Robust Vision-Language Pre-training Under Cascading Pseudo-Label Noise

Modern fine-grained vision-language and region-captioning models rely predominantly on automated annotation pipelines combining specialist detectors, segmenters, OCR, and large multimodal models without human verification. These multi-stage machine annotations suffer from compounded errors, such as mislocalized bounding boxes and hallucinated visual attributes, which directly propagate into downstream vision-language representations and degrade generation quality. Current learning-with-noisy-labels methods assume human inter-annotator disagreement or synthetic classification noise, failing to model the structured, cascading failure modes inherent to vision-language pseudo-labeling pipelines.

Possible candidate · 3/5 runs4 papers report this50% from 2025+

Vision-Language Understanding & Captioning

Barrier to removeOpen

Cross-Cultural Knowledge Grounding for Vision-Language Understanding

Current knowledge-grounded vision-language benchmarks rely almost exclusively on English Wikipedia as their ground-truth reference base. Because Wikipedia has documented geographic and cultural footprint skews, models evaluated on imagery from low-footprint cultures suffer from missing reference facts rather than failures in visual-semantic reasoning. Consequently, benchmarks penalize models on underrepresented cultures and researchers cannot diagnose whether errors stem from visual perception, language generation, or knowledge retrieval omissions.

Strong candidate · 4/5 runs3 papers report this100% from 2025+
Generated automatically, not curated by hand. Automated prior-work checks catch about a third of existing work, so treat this problem as a lead to investigate.