Open Problems

Medical Imaging

Open-Vocabulary Diagnostic Knowledge Retrieval for Medical Imaging Without Predefined Ontologies

Barrier to removeOpen
Possible candidate · 2/5 runs3 papers report this0% from 2025+

Generated automatically from the limitations stated in 3 papers (CVPR, ICCV, EMNLP), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.

The problem

Current diagnostic and retrieval models in medical imaging depend on fixed disease categories and rigid hierarchical ontologies. When patients present with rare, novel, or out-of-vocabulary pathologies—or when clinical environments lack standardized label taxonomies—these systems fail to retrieve relevant medical knowledge or characterize visual findings. Consequently, automated diagnostic support is structurally blocked from assisting on atypical and long-tail conditions where clinical reference support is most needed.

Why it matters

Enables automated medical decision-support systems to retrieve relevant clinical evidence and characterize anomalies for rare, emerging, or uncataloged pathologies without requiring manual ontology curation or model retraining.

Ways to approach it

Prior-work checks are free with an account. Results someone already ran are shown to everyone.

  1. 1

    Unstructured literature retrieval baseline: Replace graph-based ontology queries with dense vector retrieval directly over uncurated medical literature (e.g., PubMed Central open access articles and clinical reference texts), measuring zero-shot top-$k$ diagnostic retrieval accuracy on rare pathology subsets withheld from training.

  2. 2

    Dynamic terminology extraction from free-text reports: Train a token-level medical concept extractor that maps visual anomaly regions directly to free-form clinical descriptions without fixed category bins, measuring concept-level overlap (F1 score) against expert radiologist annotations on out-of-vocabulary findings.

  3. 3

    Multi-modal open-vocabulary grounding: Evaluate vision-language models trained on uncurated image-report pairs for zero-shot localization and retrieval of rare disease patterns, measuring recall@$k$ on rare and long-tail imaging cases compared to closed-vocabulary baselines.

Have a different approach?

Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.

Free · 3 checks per day

Why it might fail

Unstructured open-vocabulary retrieval may retrieve clinically non-specific or noisy text that fails expert verification, making unconstrained retrieval unsafe in high-stakes differential diagnosis compared to curated ontologies.

Evidence

Each paper's own statement of the limitation, verbatim.

Nearest existing work

Generated automatically, not curated by hand. Automated prior-work checks catch about a third of existing work, so treat this problem as a lead to investigate.