Digital Humans, Avatars & Virtual Try-On
Overcoming Parametric 3DMM Expressiveness and Mouth Interior Bottlenecks in Monocular Facial Avatars
Generated automatically from the limitations stated in 11 papers (CVPR, NeurIPS), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Neural head avatar pipelines universally rely on parametric mesh models like FLAME as geometric proxies, inheriting their topological limitations. Because FLAME lacks internal oral anatomy (teeth and tongue) and fine-grained geometric detail, avatars suffer from severe projection artifacts such as teeth textures baked directly onto lip surfaces and distorted speech articulation. Furthermore, low-dimensional linear expression spaces cap the capture of dynamic micro-expressions, dynamic wrinkles, and non-parametric regions like hair.
Why it matters
Enables realistic conversational avatars capable of natural speech articulation without oral rendering artifacts or expression damping.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
Hybrid Parametric-Volumetric Mouth Interior: Augment a FLAME-guided neural radiance or Gaussian avatar with a dedicated articulated oral cavity representation (combining rigid dental structures with a deformable tongue model), measuring lip-mouth rendering fidelity (LPIPS/PSNR) and oral depth/penetration error on speaking video datasets.
- 2
Residual Per-Vertex Dynamic Displacement Learning: Train a lightweight neural deformation field conditioned on audio and high-resolution video frames that predicts non-linear geometric residuals on top of 3DMM coefficients, measuring surface reconstruction error against multi-view/3D scan benchmarks (e.g., capture of micro-wrinkles and non-parametric facial dynamics).
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
End-to-end generative 3D diffusion and video models might bypass explicit 3DMM mesh conditioning altogether if they learn temporal 3D consistency directly from large video corpora. Additionally, monocular oral capture remains inherently under-constrained due to severe self-occlusion during speech.
Evidence
Each paper's own statement of the limitation, verbatim.
- PhysHead: Simulation-Ready Gaussian Head AvatarsCVPR 2026
FLAME geometry approximations cause texture proxy artifacts in regions like the mouth interior (e.g., teeth baked onto lips).
- OMG-Avatar: One-shot Multi-LOD Gaussian Head AvatarCVPR 2026
Expressiveness is bounded by the FLAME parametric model, failing to capture non-parametric dynamics such as tongue motions and complex hair deformation.
- FlexAvatar: Learning Complete 3D Head Avatars with Partial SupervisionCVPR 2026
Relying on FLAME expression codes limits expressive facial motions and fine geometry, such as tongue movements.
- Feed-Forward One-Shot Animatable Textured Mesh Avatar ReconstructionCVPR 2026
Animation capabilities are bounded by FLAME's parametric expressiveness, preventing the generation of fine dynamic wrinkles and tongue movements.
- FG-Portrait: 3D Flow Guided Editable Portrait AnimationCVPR 2026
FLAME's limited mesh resolution struggles to represent fine-grained expressions, so subtle facial movements may not be captured
- MVPortrait: Text-Guided Motion and Emotion Control for Multi-view Vivid Portrait AnimationCVPR 2025
FLAME cannot capture micro-expressions, so fine-grained facial details beyond the parametric space are lost
- HRAvatar: High-Quality and Relightable Gaussian Head AvatarCVPR 2025
Remains constrained by FLAME priors when training data is insufficient, limiting control over hair and accessories
- GAF: Gaussian Avatar Reconstruction from Monocular Videos via Multi-view DiffusionCVPR 2025
Quality is capped by the expressiveness of the FLAME parametric model, which lacks detailed hair geometry and animation
Show all 11 papers
- CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion ModelsCVPR 2025
The FLAME-based 3DMM does not model tongue or hair motion, causing artifacts in those regions
- BecomingLit: Relightable Gaussian Avatars with Hybrid Neural ShadingNeurIPS 2025
Inherits FLAME's limits: low expressiveness, sensitivity to tracking failures (especially gaze), no mouth interior modeling
- Instant Volumetric Head AvatarsCVPR 2023
Mouth interior quality is poor because FLAME has no teeth geometry; extreme or extrapolated expressions cause mouth-region artifacts
Nearest existing work
- Teeth Reconstruction and Performance Capture Using a Phone CameraICCV 2025
- GeoAvatar: Adaptive Geometrical Gaussian Splatting for 3D Head AvatarICCV 2025
- VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single ImageNeurIPS 2025
- MeGA: Hybrid Mesh-Gaussian Head Avatar for High-Fidelity Rendering and Head EditingCVPR 2025
- Neural Head Avatars From Monocular RGB VideosCVPR 2022
- GaussianAvatars: Photorealistic Head Avatars with Rigged 3D GaussiansCVPR 2024
- DiffusionAvatars: Deferred Diffusion for High-fidelity 3D Head AvatarsCVPR 2024
- Instant Volumetric Head AvatarsCVPR 2023
- Expressive Talking Human from Single-Image with Imperfect PriorsICCV 2025
- GANHead: Towards Generative Animatable Neural Head AvatarsCVPR 2023
- FlashAvatar: High-fidelity Head Avatar with Efficient Gaussian EmbeddingCVPR 2024
- GaussianSpeech: Audio-Driven Personalized 3D Gaussian AvatarsICCV 2025
- I M Avatar: Implicit Morphable Head Avatars From VideosCVPR 2022
- GeoDiff4D: Geometry-Aware Diffusion for 4D Head Avatar ReconstructionCVPR 2026
- Learning Personalized High Quality Volumetric Head Avatars From Monocular RGB VideosCVPR 2023
Related open problems
Digital Humans, Avatars & Virtual Try-On
Joint Pose Refinement and Error-Tolerant Conditioning for Downstream Human Avatar Synthesis
Current avatar creation and virtual try-on pipelines treat SMPL/SMPL-X estimations, camera calibrations, and segmentation masks as fixed, ground-truth preconditions. When off-the-shelf estimators inevitably produce misaligned or out-of-distribution body parameters on in-the-wild images, these errors cascade into severe geometric distortion, stitching artifacts, and generation failure. Because existing downstream pipelines do not jointly optimize human body parameters or account for parameter uncertainty, high-fidelity human generation remains restricted to controlled studio captures with pristine initializations.
Digital Humans, Avatars & Virtual Try-On
Reducing Downstream Dependency of Garment/Avatar Generation on Upstream 2D and 3D Human-Parsing Predictions
Every pipeline in this space chains off-the-shelf estimators — pose keypoint detectors, DensePose, SMPL/SMPL-X fitting, segmentation/parsing models, SAM — and treats their outputs as ground truth. When these upstream predictions are noisy, fail on exaggerated poses, unusual garments, or garment-only imagery, the downstream model inherits or amplifies the errors (garment leaking, misinitialized garment geometry, failed reposing). No one measures or controls how error propagates through this preprocessing chain, so each new pipeline is brittle in its own undocumented way.
Digital Humans, Avatars & Virtual Try-On
Benchmarking and Improving 3D Human Body Model Robustness Across Underrepresented Demographics and Non-Standard Morphologies
Existing 3D statistical body models and avatar pipelines are predominantly trained on narrow demographic subsets (such as adult European CAESAR scans) and evaluated solely on standard normative body templates. Consequently, downstream avatars and virtual try-on methods systematically produce distorted geometry, unnatural surface artifacts, and inaccurate anthropometric measurements on children, adolescents, diverse global populations, individuals with disabilities, and non-standard body proportions. Because existing benchmarks do not measure generalization across diverse body shapes, failure modes on these populations remain unquantified and unaddressed.