Open Problems

Semantic & Instance Segmentation

Overcoming the Mask Proposal Recall Ceiling in Open-Vocabulary Segmentation

Barrier to removeOpen
Strong candidate · 4/5 runs9 papers report this22% from 2025+

Generated automatically from the limitations stated in 9 papers (CVPR, ICCV, NeurIPS), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.

The problem

Current open-vocabulary and weakly supervised segmentation pipelines decouple class-agnostic mask proposal generation from segment classification, freezing an upstream proposal model like SAM, FreeSOLO, or Mask2Former. Because the downstream classifier operates solely on candidate masks, segmentation accuracy on novel classes and sub-object regions is strictly bounded by the proposal generator's recall, creating a ceiling that downstream improvements cannot fix. Failures to isolate novel or fine-grained instances at the proposal stage permanently propagate through the pipeline without any mechanism for recovery.

Why it matters

Enables segmentation frameworks to generalize beyond the fixed objecthood definitions and training class budgets of pre-trained proposal generators, segmenting fine-grained sub-parts and novel object categories previously missed.

Ways to approach it

Prior-work checks are free with an account. Results someone already ran are shown to everyone.

  1. 1

    Iterative Text-Guided Proposal Refinement: Construct an iterative mask adaptation module that takes text embeddings of candidate classes to modulate and deform frozen proposal masks, measuring mask average recall (AR) and panoptic quality (PQ) on novel classes in ADE20k/LVIS.

  2. 2

    Multi-Granularity Candidate Ensembling: Design a lightweight merging and cross-scale pruning mechanism combining low-level unsupervised boundaries (e.g., self-supervised ViT patch features) with frozen proposal masks, evaluating recall gains across sub-object and rare entity benchmarks.

  3. 3

    End-to-End Proposal-Classifier Co-adaptation: Train a lightweight feedback adapter that updates proposal boundaries based on downstream classification uncertainty, measuring the delta between fixed-proposal oracle upper bounds and adapted-proposal mIoU.

Have a different approach?

Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.

Free · 3 checks per day

Why it might fail

If foundation vision models (such as iterative promptable foundation segmenters) scale to near-100% zero-shot mask recall across all sub-object granularities natively, eliminating the proposal bottleneck via brute pretraining scale.

Evidence

Each paper's own statement of the limitation, verbatim.

Show all 9 papers

Nearest existing work

Related open problems

Semantic & Instance Segmentation

Barrier to removePartly addressed

Benchmarking and Bridging Idealized Preconditions in Promptable Segmentation

Current promptable segmentation models systematically rely on idealized preconditions during evaluation and training, including ground-truth-derived prompts (points or tight boxes extracted directly from GT masks), dense instance annotations, and perfect spatial feature alignment. In real-world interactive deployment, prompts originate from noisy human clicks, coarse bounding boxes, or imperfect auxiliary detectors, leading to uncharacterized performance degradation. Furthermore, models optimized exclusively on point prompts fail to maintain performance across alternative prompt modalities or without high-resolution decoders. Consequently, reported benchmark gains reflect idealized prompting setups rather than actual operational capability.

Possible candidate · 3/5 runs5 papers report this60% from 2025+

Semantic & Instance Segmentation

Barrier to removePartly addressed

Benchmarking and Developing Segmentation Without Dense Initial Annotation Preconditions

Current semantic and instance segmentation methods rely on strict supervision preconditions, such as dense pixel-wise ground truth or accurate initial bounding-box-derived masks. When applied to domains where only weak annotations (such as image-level tags or sparse points) are available from scratch, these methods cannot be initialized or trained. Consequently, segmentation models cannot be deployed in annotation-scarce domains without first undertaking costly manual pixel-level or box-level labeling campaigns.

Possible candidate · 2/5 runs3 papers report this0% from 2025+

Semantic & Instance Segmentation

Barrier to removePartly addressed

Hierarchical Layout and Vector Segmentation Under Incomplete Multi-Level Supervision

Current instance and semantic segmentation methods for structured documents and vector graphics rely on full, multi-tier hierarchical annotations (simultaneously requiring word, line, and paragraph labels) and fail when vector elements are deeply nested, tangled, or overlapping. Because multi-level hierarchical annotations are significantly more expensive than standard single-level bounding boxes, deploying segmentation pipelines to diverse real-world documents remains cost-prohibitive. Furthermore, the inability to parse overlapping and nested primitives prevents reliable digitization of technical figures, CAD drawings, and complex vector assets.

Possible candidate · 2/5 runs3 papers report this33% from 2025+
Generated automatically, not curated by hand. Automated prior-work checks catch about a third of existing work, so treat this problem as a lead to investigate.