3D Shape Modeling & Surface Reconstruction
Cross-Category and Vocabulary Generalization Benchmark for Text-to-3D Shape Models
Generated automatically from the limitations stated in 3 papers (CVPR), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Current text-driven 3D shape generation and modeling methods are predominantly trained and benchmarked strictly on Text2Shape, which is limited to chairs and tables (~11.5k shapes and a ~3.6k-word vocabulary). Consequently, it is unknown whether existing architectural designs, text encodings, and shape priors maintain generative fidelity when exposed to out-of-domain object categories or open-vocabulary text prompts. Without evaluating existing methods across broader category and vocabulary distributions, the field cannot determine if current performance gains reflect generalizable shape modeling or dataset-specific memorization.
Why it matters
Establishes the empirical boundaries of current text-conditioned 3D shape generation methods and provides a standardized protocol for evaluating cross-category robustness.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Systematic multi-category benchmark: Evaluate existing open-source Text2Shape models zero-shot and with fine-tuning on diverse ShapeNet classes (e.g., cars, airplanes, vessels) and measure text-shape retrieval accuracy, Chamfer Distance, and F-score across unseen classes.
- 2
Open-vocabulary prompt stress-testing: Construct an evaluation suite of synthetic and real-world natural language descriptions spanning compositional and out-of-vocabulary terms to quantify failure modes in semantic binding and geometry generation.
- 3
Patch- and component-level transfer evaluation: Assess whether localized geometric priors transfer across dissimilar categories when conditioned on shared compositional descriptions (e.g., legs, flat surfaces, thin supports).
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
The problem could be dissolved if recent large-scale open-vocabulary 3D generative frameworks (e.g., trained on Objaverse) render chair/table-constrained Text2Shape baselines entirely obsolete for research benchmarking.
Evidence
Each paper's own statement of the limitation, verbatim.
- HyperSDFusion: Bridging Hierarchical Structures in Language and Geometry for Enhanced 3D Text2Shape GenerationCVPR 2024
Evaluated only on Text2Shape, which covers only chairs and tables from ShapeNet; generality to other shape categories is untested
- Parts2Words: Learning Joint Embedding of Point Clouds and Texts by Bidirectional Matching Between Parts and WordsCVPR 2023
Evaluated only on Text2Shape (chairs and tables, ~11.5k training shapes, vocabulary of 3587 words), so generalization to open vocabulary or other object categories is untested
- Diffusion-SDF: Text-To-Shape via Voxelized DiffusionCVPR 2023
Trained and evaluated only on Text2Shape, which covers just chairs and tables from ShapeNet, so generalization to other categories is unvalidated
Nearest existing work
- OpenShape: Scaling Up 3D Shape Representation Towards Open-World UnderstandingNeurIPS 2023
- ShapeCrafter: A Recursive Text-Conditioned 3D Shape Generation ModelNeurIPS 2022
- Texturify: Generating Textures on 3D Shape SurfacesECCV 2022
- Towards Implicit Text-Guided 3D Shape GenerationCVPR 2022
- ZeroShape: Regression-based Zero-shot Shape ReconstructionCVPR 2024
- CLIP-Forge: Towards Zero-Shot Text-To-Shape GenerationCVPR 2022
- Parts2Words: Learning Joint Embedding of Point Clouds and Texts by Bidirectional Matching Between Parts and WordsCVPR 2023
- Planes vs. Chairs: Category-Guided 3D Shape Learning without Any 3D CuesECCV 2022
- Geometry-Guided Modeling of Foundation Features Enables Generalizable Object Shape Deformation LearningICML 2026
- Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent RepresentationNeurIPS 2023
- ShapeWords: Guiding Text-to-Image Synthesis with 3D Shape-Aware PromptsCVPR 2025
- TAPS3D: Text-Guided 3D Textured Shape Generation From Pseudo SupervisionCVPR 2023
- SurfGen: Adversarial 3D Shape Synthesis With Explicit Surface DiscriminatorsICCV 2021
- ShapeScaffolder: Structure-Aware 3D Shape Generation from TextICCV 2023
- Domain-Adaptive Single-View 3D ReconstructionICCV 2019