Open Problems

Semantic & Instance Segmentation

Hierarchical Layout and Vector Segmentation Under Incomplete Multi-Level Supervision

Barrier to removePartly addressed
Possible candidate · 2/5 runs3 papers report this33% from 2025+

Generated automatically from the limitations stated in 3 papers (CVPR), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.

The problem

Current instance and semantic segmentation methods for structured documents and vector graphics rely on full, multi-tier hierarchical annotations (simultaneously requiring word, line, and paragraph labels) and fail when vector elements are deeply nested, tangled, or overlapping. Because multi-level hierarchical annotations are significantly more expensive than standard single-level bounding boxes, deploying segmentation pipelines to diverse real-world documents remains cost-prohibitive. Furthermore, the inability to parse overlapping and nested primitives prevents reliable digitization of technical figures, CAD drawings, and complex vector assets.

Why it matters

Enables automated extraction and structured hierarchy recovery for complex vector graphics and document layouts using only cheap, single-level bounding box annotations.

Ways to approach it

Prior-work checks are free with an account. Results someone already ran are shown to everyone.

  1. 1

    Weakly-supervised multi-level grouping: Formulate a bottom-up relational grouping objective that infers nested line and paragraph structures from single-level bounding box supervision, measuring hierarchical layout mAP and boundary F1 against fully supervised baselines.

  2. 2

    Topology-aware vector overlap parsing: Build a graph-based segmentation model that operates over vector primitive intersections and containment relations, benchmarked on datasets of tangled vector illustrations and multi-layer diagrams for instance segmentation accuracy under varying overlap depth.

Have a different approach?

Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.

Free · 3 checks per day

Why it might fail

If simple rule-based spatial clustering and heuristic post-processing over OCR tokens already recover the hierarchy accurately enough for downstream use, the added complexity of learning nested vector relations will provide insufficient empirical gain.

Evidence

Each paper's own statement of the limitation, verbatim.

Nearest existing work

Related open problems

Semantic & Instance Segmentation

Barrier to removeOpen

Overcoming the Mask Proposal Recall Ceiling in Open-Vocabulary Segmentation

Current open-vocabulary and weakly supervised segmentation pipelines decouple class-agnostic mask proposal generation from segment classification, freezing an upstream proposal model like SAM, FreeSOLO, or Mask2Former. Because the downstream classifier operates solely on candidate masks, segmentation accuracy on novel classes and sub-object regions is strictly bounded by the proposal generator's recall, creating a ceiling that downstream improvements cannot fix. Failures to isolate novel or fine-grained instances at the proposal stage permanently propagate through the pipeline without any mechanism for recovery.

Strong candidate · 4/5 runs9 papers report this22% from 2025+

Semantic & Instance Segmentation

Barrier to removePartly addressed

Benchmarking and Bridging Idealized Preconditions in Promptable Segmentation

Current promptable segmentation models systematically rely on idealized preconditions during evaluation and training, including ground-truth-derived prompts (points or tight boxes extracted directly from GT masks), dense instance annotations, and perfect spatial feature alignment. In real-world interactive deployment, prompts originate from noisy human clicks, coarse bounding boxes, or imperfect auxiliary detectors, leading to uncharacterized performance degradation. Furthermore, models optimized exclusively on point prompts fail to maintain performance across alternative prompt modalities or without high-resolution decoders. Consequently, reported benchmark gains reflect idealized prompting setups rather than actual operational capability.

Possible candidate · 3/5 runs5 papers report this60% from 2025+

Semantic & Instance Segmentation

Barrier to removePartly addressed

Benchmarking and Developing Segmentation Without Dense Initial Annotation Preconditions

Current semantic and instance segmentation methods rely on strict supervision preconditions, such as dense pixel-wise ground truth or accurate initial bounding-box-derived masks. When applied to domains where only weak annotations (such as image-level tags or sparse points) are available from scratch, these methods cannot be initialized or trained. Consequently, segmentation models cannot be deployed in annotation-scarce domains without first undertaking costly manual pixel-level or box-level labeling campaigns.

Possible candidate · 2/5 runs3 papers report this0% from 2025+
Generated automatically, not curated by hand. Automated prior-work checks catch about a third of existing work, so treat this problem as a lead to investigate.