Medical Imaging
Open-Vocabulary Diagnostic Knowledge Retrieval for Medical Imaging Without Predefined Ontologies
Generated automatically from the limitations stated in 3 papers (CVPR, ICCV, EMNLP), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Current diagnostic and retrieval models in medical imaging depend on fixed disease categories and rigid hierarchical ontologies. When patients present with rare, novel, or out-of-vocabulary pathologies—or when clinical environments lack standardized label taxonomies—these systems fail to retrieve relevant medical knowledge or characterize visual findings. Consequently, automated diagnostic support is structurally blocked from assisting on atypical and long-tail conditions where clinical reference support is most needed.
Why it matters
Enables automated medical decision-support systems to retrieve relevant clinical evidence and characterize anomalies for rare, emerging, or uncataloged pathologies without requiring manual ontology curation or model retraining.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Unstructured literature retrieval baseline: Replace graph-based ontology queries with dense vector retrieval directly over uncurated medical literature (e.g., PubMed Central open access articles and clinical reference texts), measuring zero-shot top-$k$ diagnostic retrieval accuracy on rare pathology subsets withheld from training.
- 2
Dynamic terminology extraction from free-text reports: Train a token-level medical concept extractor that maps visual anomaly regions directly to free-form clinical descriptions without fixed category bins, measuring concept-level overlap (F1 score) against expert radiologist annotations on out-of-vocabulary findings.
- 3
Multi-modal open-vocabulary grounding: Evaluate vision-language models trained on uncurated image-report pairs for zero-shot localization and retrieval of rare disease patterns, measuring recall@$k$ on rare and long-tail imaging cases compared to closed-vocabulary baselines.
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
Unstructured open-vocabulary retrieval may retrieve clinically non-specific or noisy text that fails expert verification, making unconstrained retrieval unsafe in high-stakes differential diagnosis compared to curated ontologies.
Evidence
Each paper's own statement of the limitation, verbatim.
- Instance-level Expert Knowledge and Aggregate Discriminative Attention for Radiology Report GenerationCVPR 2024
Relies on a pre-defined set of disease categories for the knowledge retrieval and diagnostic components, which may limit applicability to rare or novel pathologies.
- Visual-Textual Attentive Semantic Consistency for Medical Report GenerationICCV 2021
Relies on a pre-defined set of 18 disease and 32 description pattern labels, potentially failing to describe rare or out-of-vocabulary pathologies
- Ontological attention ensembles for capturing semantic concepts in ICD code prediction from clinical textEMNLP 2019
Requires a strictly defined hierarchical ontology, which may not be available for all clinical labeling tasks
Nearest existing work
- Improving Neural Models for Radiology Report Retrieval with Lexicon-based Automated AnnotationNAACL 2022
- AHIVE: Anatomy-aware Hierarchical Vision Encoding for Interactive Radiology Report RetrievalCVPR 2024
- MedKLIP: Medical Knowledge Enhanced Language-Image Pre-Training for X-ray DiagnosisICCV 2023
- Fact-Aware Multimodal Retrieval Augmentation for Accurate Medical Radiology Report GenerationNAACL 2025
- Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community RetrievalICLR 2025
- MedRAX: Medical Reasoning Agent for Chest X-rayICML 2025
- EviRAG: Evidence-Guided Retrieval-Augmented Generation for Medical Vision-Language ModelsSIGIR 2026
- Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-trainingICCV 2025
- GLoRIA: A Multimodal Global-Local Representation Learning Framework for Label-Efficient Medical Image RecognitionICCV 2021
- MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQACVPR 2026
- Pushing the Limits of Radiology with Joint Modeling of Visual and Textual InformationACL 2018
- FACT: Fuzzy Alignment with Comorbidity Topology for Reliable Multi-Label Medical Image DiagnosisICML 2026
- UniMedVL: Unifying Medical Multimodal Understanding and Generation through Observation-Knowledge-AnalysisICML 2026
- ProbMED: A Probabilistic Framework for Medical Multimodal BindingICCV 2025
- MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for MedicineICLR 2025