Hand-Object Interaction & Affordance
Robust Hand-Object Interaction Learning Beyond Cascaded Parametric Pose Estimators
Generated automatically from the limitations stated in 42 papers (CVPR, ICCV, ECCV), listed under Evidence. It is not a paper, and it does not come from papers submitted to CSPaper.
The problem
Downstream hand-object interaction (HOI) and affordance models are universally bottlenecked by a strict reliance on upstream off-the-shelf 3D hand pose estimators (e.g., FrankMocap, HaMeR, MANO parameterizations) and bounding-box detectors. When these upstream tools encounter severe object occlusions, fast manipulation dynamics, or egocentric perspective distortion, their estimation errors cascade directly into downstream HOI training, causing catastrophic mis-projections and synthetic contact artifacts. Because downstream methods require clean parametric 3D meshes as preconditions, massive repositories of in-the-wild manipulation videos cannot be utilized without fragile filtering heuristics that discard over half the data.
Why it matters
Affordance and manipulation models can train directly on millions of uncurated in-the-wild interaction videos without discarding frames due to upstream tracking failures. It also enables robust deployment on real-world egocentric video where severe hand-object occlusion causes off-the-shelf parametric hand estimators to fail completely.
Ways to approach it
Prior-work checks are free with an account. Results someone already ran are shown to everyone.
- 1
End-to-end joint contact and affordance learning directly from uncropped 2D RGB/video sequences using implicit or voxelized contact fields, bypassing intermediate parametric MANO fitting; measure interaction synthesis and affordance prediction error (AUC / Chamfer distance) on DexYCB and HO-3D when upstream pose estimates are degraded or withheld.
- 2
An error-aware HOI training framework that models upstream pose/mask uncertainty distributions (e.g., from HaMeR/SAM2) and marginalizes over pose ambiguity during contact optimization; measure downstream task degradation against injected upstream keypoint/pose noise compared to rigid cascading baselines.
- 3
Self-supervised cross-view and temporal photometric consistency objectives on unannotated egocentric manipulation video to jointly refine hand-object contact boundaries without ground-truth 3D meshes; measure 3D contact accuracy and frame yield rate on uncurated datasets like EPIC-KITCHENS and Ego4D.
Have a different approach?
Describe how you would tackle this problem and we'll look for papers that already do it. Free; your text stays private.
Why it might fail
If generalist vision foundation models or next-generation hand trackers solve in-the-wild 3D hand-object tracking with near-zero error under complete occlusion, the cascading error problem would dissolve from the upstream side.
Evidence
Each paper's own statement of the limitation, verbatim.
- ViHOI: Human-Object Interaction Synthesis with Visual PriorsCVPR 2026
Cannot generate detailed finger motions or dexterous manipulations due to the lack of fine-grained hand annotations in the underlying training datasets.
- UniHand: A Unified Model for Diverse Controlled 4D Hand Motion ModelingICLR 2026
Requires an initial hand pose token to bind the hand perceptron attention to the correct hand instance when multiple hands are present.
- SHOW3D: Capturing Scenes of 3D Hands and Objects in the WildCVPR 2026
Marker-less automated filtering yields only 1.9M high-confidence frames for hand pose estimation out of 4.3M total captured frames (~44% yield rate).
- Interaction-aware Representation Modeling With Co-Occurrence Consistency for Egocentric Hand-Object ParsingICLR 2026
Performance is sensitive to the hand pixel count threshold hyperparameter (tau) in the CoCo loss, risking false positive hand detections if too low or missing partially visible hands if too high
- Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing GlovesCVPR 2026
Requires optimizing a separate subject-specific 3D Gaussian hand avatar for each target hand identity.
- GenHOI: Towards Object-Consistent Hand-Object Interaction with Temporally Balanced and Spatially Selective Object InjectionCVPR 2026
Requires binary mask videos defining hand-object interaction regions.
- Enhancing Hands in 3D Whole-Body Pose Estimation with Conditional Hands ModulatorCVPR 2026
Performance depends on the quality of upstream hand localization and 2D bounding box crops.
- EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric VideoICLR 2026
Hand pose annotations rely on on-device visual tracking algorithms and can fail or lose confidence during severe object occlusions (such as folding laundry) or high-speed hand movements.
Show all 42 papers
- Clay-to-Stone: Phase-wise 3D Gaussian Splatting for Monocular Articulated Hand-Object Manipulation ModelingCVPR 2026
Requires known or pre-computed MANO hand pose parameters and 6-DoF object poses for canonical space mapping and initialization.
- AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp SynthesisCVPR 2026
Cannot train directly on datasets lacking MANO hand parameterizations or containing noisy hand annotations (e.g., AffordPose).
- VTON-HandFit: Virtual Try-on for Arbitrary Hand Pose Guided by Hand Priors EmbeddingCVPR 2025
Requires HaMeR-based 3D hand reconstruction, pose estimation, and DensePose as inputs, so errors in these upstream parsers propagate; the paper notes existing parsing models fail on tiny fragmented finger areas
- Improving Handshape Representations for Sign Language Processing: A Graph Neural Network ApproachEMNLP 2025
No frame-level handshape annotations, so static-frame selection relies on a crude minimum-displacement heuristic that may miss the canonical frame
- Hearing Hands: Generating Sounds from Physical Interactions in 3D ScenesCVPR 2025
3D hand detection (HaMeR) errors propagate into the dataset, causing inaccurate hand motions and desynchronized predictions when hands are less visible or motion is ambiguous
- HaWoR: World-Space Hand Motion Reconstruction from Egocentric VideosCVPR 2025
Depends on an external hand detector/tracker; method evaluated only on DexYCB and HOT3D benchmarks
- HOIGPT: Learning Long-Sequence Hand-Object Interaction with Language ModelsCVPR 2025
Object geometry must be available as a point cloud at encoding/decoding time, and hand poses depend on MANO parameterization, so arbitrary objects/hand representations are out of scope
- HOI-Dyn: Learning Interaction Dynamics for Human-Object Motion DiffusionNeurIPS 2025
Assumes rigid objects and SMPL-X humans with coarse hand representations, causing minor inaccuracies in hand-object interactions, particularly rotations
- Guiding Human-Object Interactions with Rich Geometry and RelationsCVPR 2025
The dataset excludes hand/finger motions, so generated hand details are low-fidelity
- FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image GenerationCVPR 2025
Conditioning requires 2D keypoints; the paper notes hand masks and keypoint images can be ambiguous under heavy finger occlusion, and NVS depends on an off-the-shelf 3D pose estimator with assumed camera intrinsics
- EgoDTM: Towards 3D-Aware Egocentric Video-Language PretrainingNeurIPS 2025
Hand segmentation quality is substantially below supervised upper bound (61.01 vs 90.87 hand-AP on VISOR) due to noisy HOI detection prompts and SAM2 struggles in cluttered egocentric backgrounds
- Weakly-Supervised 3D Hand Reconstruction with Knowledge Prior and Uncertainty GuidanceECCV 2024
Restricted to static hand knowledge; hand dynamics from video are not addressed
- Reconstructing Hands in 3D with TransformersCVPR 2024
Requires an external hand detector and handedness classifier at test time, and handles only one hand per crop (two-hand scenes processed independently)
- Prompting Future Driven Diffusion Model for Hand Motion PredictionECCV 2024
Hand-object interactions not modeled despite datasets containing object pose information
- On the Utility of 3D Hand Poses for Action RecognitionECCV 2024
Depends on hand pose availability; poses estimated from visual inputs suffer from out-of-view cases and noise
- InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object InteractionNeurIPS 2024
Hand poses are set to an average pose because the BEHAVE dataset lacks hand annotations, degrading interaction realism
- D-SCo: Dual-Stream Conditional Diffusion for Monocular Hand-Held Object ReconstructionECCV 2024
Requires predicted hand pose (MANO parameters) and camera view from off-the-shelf estimators; uses ground-truth hand poses in ObMan evaluation, so errors in hand pose estimation are not fully assessed
- Controlling the World by Sleight of HandECCV 2024
Performance degrades with noisy or low-detail hand masks; more detailed segmentation masks produce finer-grained outputs
- 3D Reconstruction of Objects in Hands without Real World 3D SupervisionECCV 2024
Inaccurate FrankMocap hand pose predictions cause mis-projection so sampled points do not cover the entire object, excluding parts of the object from supervision
- HandR2N2: Iterative 3D Hand Pose Estimation Using a Residual Recurrent Neural NetworkICCV 2023
Performance degrades when hand parts are severely occluded because the resampler cannot sample informative points near occluded joints
- HaMuCo: Hand Pose Estimation via Multiview Collaborative Self-Supervised LearningICCV 2023
Focused on hands without heavy occlusions; hand-object interaction scenarios and unsynchronized multi-view inputs are excluded
- Dynamic Hyperbolic Attention Network for Fine Hand-object ReconstructionICCV 2023
Requires datasets with 3D hand and object mesh ground truth for supervision; performance on real data relies on filtered subsets (e.g., FHB excludes frames with hand-object distance >10 mm)
- Diffusion-Guided Reconstruction of Everyday Hand-Object Interaction ClipsICCV 2023
Relies on off-the-shelf hand pose prediction (FrankMocap) and masks; while robust to doubled hand-pose error, severe initial failures require a separate trajectory-optimization correction step
- Decoupled Iterative Refinement Framework for Interacting Hands Reconstruction from a Single RGB ImageICCV 2023
Does not fully exploit estimated 3D mesh information, which could help fine-grained inter-hand understanding
- ContactGen: Generative Contact Modeling for Grasp GenerationICCV 2023
Requires a per-object optimization (200 + 1000 Adam iterations) to recover the hand pose from each sampled ContactGen
- CHORD: Category-level Hand-held Object Reconstruction via Shape DeformationICCV 2023
Relies on preceding hand pose estimation and category-level object pose estimation; GT-pose results (CD 2.30 vs 7.69 for mug) show large sensitivity to pose input quality
- Affordances From Human Videos as a Versatile Representation for RoboticsCVPR 2023
Labels come from off-the-shelf hand-object detection, pose, and egomotion estimators, so label noise from those tools propagates into training
- Joint Hand Motion and Interaction Hotspots Prediction From Egocentric VideosCVPR 2022
Pseudo-ground-truth data collection depends entirely on the accuracy and error propagation of off-the-shelf active hand-object detectors and skin/fingertip segmentation.
- Generative Adversarial Network for Future Hand Segmentation from Egocentric VideoECCV 2022
Ground-truth hand masks do not distinguish left from right hands, so the model cannot make separate left/right hand predictions
- Semi-Supervised 3D Hand-Object Poses Estimation With Interactions in TimeCVPR 2021
Removing either spatial or temporal filtering constraints degrades hand pose accuracy below the supervised baseline, so the pseudo-label pipeline is fragile to filter design
- Hand Image Understanding via Deep Multi-Task LearningICCV 2021
Performance depends on the newly collected HIU-Data (33k manually annotated images); without it, evaluation of mask/2D pose claims against other methods is limited since no prior dataset has high-quality hand masks
- CrossInfoNet: Multi-Task Information Sharing Based Hand Pose EstimationCVPR 2019
Relies on a pre-cropped, normalized hand region (96x96) and assumes a fixed palm/finger joint partition, which may not transfer to hands in unconstrained scenes or other skeletal topologies
- Point-to-Point Regression PointNet for 3D Hand Pose EstimationECCV 2018
Requires hand segmentation (via a separate hourglass network on NYU/ICVL) and PCA pose prior fitting
- Learning to Predict Gaze in Egocentric VideoICCV 2013
Performance is sensitive to the accuracy of hand segmentation and configuration classification, which remains a non-trivial task.
Nearest existing work
- AffordPose: A Large-Scale Dataset of Hand-Object Interactions with Affordance-Driven Hand PoseICCV 2023
- Semi-Supervised 3D Hand-Object Poses Estimation With Interactions in TimeCVPR 2021
- HOLD: Category-agnostic 3D Reconstruction of Interacting Hands and Objects from VideoCVPR 2024
- Towards A Richer 2D Understanding of Hands at ScaleNeurIPS 2023
- Reconstructing Hand-Object Interactions in the WildICCV 2021
- Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with ObjectsECCV 2024
- Dense Hand-Object(HO) GraspNet with Full Grasping Taxonomy and DynamicsECCV 2024
- MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion GenerationNeurIPS 2025
- Human Hands As Probes for Interactive Object UnderstandingCVPR 2022
- Robust Hand Pose Estimation During the Interaction With an Unknown ObjectICCV 2017
- Leveraging Photometric Consistency Over Time for Sparsely Supervised Hand-Object ReconstructionCVPR 2020
- 3D Interacting Hand Pose Estimation by Hand De-Occlusion and RemovalECCV 2022
- HOnnotate: A Method for 3D Annotation of Hand and Object PosesCVPR 2020
- Transformer-Based Unified Recognition of Two Hands Manipulating ObjectsCVPR 2023
- TOUCH: Text-guided Controllable Generation of Free-Form Hand-Object InteractionsICLR 2026
Related open problems
Hand-Object Interaction & Affordance
Robust 3D Hand-Object Pose and Interaction Estimation Under Upstream 2D Segmentation and Localization Noise
Current hand-object interaction and 3D hand pose estimation models rely heavily on upstream 2D bounding box crops, hand segmentation masks, or pre-extracted interaction regions. In real-world egocentric scenes, upstream 2D segmenters and detectors degrade severely due to object occlusions, motion blur, and cluttered backgrounds, causing downstream 3D estimation to fail or hallucinate. Because models are trained and benchmarked assuming near-ideal 2D crops or masks, existing pipelines cannot operate reliably in autonomous, in-the-wild video streams.
Hand-Object Interaction & Affordance
Benchmarking Hand-Object Interaction Models Across Diverse Hand Morphologies and Embodiments
Existing hand representation and hand-object interaction models are predominantly evaluated on small demographic cohorts or fixed kinematic setups (e.g., standard anthropomorphic models like MANO or fixed robotic hands like ShadowHand). Because existing pipelines implicitly bind kinematics, palm-to-finger ratios, or subject identities into their representations, their performance on unseen human hand proportions and non-standard robotic end-effectors remains untested. As a result, it is unknown whether current interaction and affordance estimators degrade gracefully or fail catastrophically when applied outside their narrow training morphologies.
Hand-Object Interaction & Affordance
Fine-Grained 3D Hand Affordance Grounding and Generation for Small-Part Articulated Interactions
Current affordance learning pipelines rely on self-training loops, video extraction, or upstream generative HOI models that fail when contact targets are small functional components such as bottle caps, zippers, or rotary knobs. Because generated hand poses easily drift to adjacent object surfaces, models exhibit catastrophic failure on precision affordances (e.g., 0% accuracy on pull or sub-50% on twist), propagating noisy pseudo-labels across unobserved geometries. Downstream manipulation systems are consequently blocked from executing precision tasks like uncapping, pulling sliders, or fine bimanual manipulation.
Hand-Object Interaction & Affordance
Hand-Object Affordance Estimation Beyond Clean Pre-Scanned Scenes and Predefined 3D Grids
Current methods for 3D hand-object interaction and affordance estimation depend strictly on clean geometric preconditions: high-fidelity 3D scans, pre-detected object instances, static overhead views, or tightly predefined 3D bounding grids. In practical manipulation and augmented reality scenarios, environments are unsegmented, viewpoints are dynamic, and point clouds are noisy and incomplete. Relying on manually tuned grid volumes and pre-extracted meshes prevents existing models from operating directly on raw, uncurated sensor streams.