Multi-Modal 6D Object-in-Hand Pose Estimation for Robotic Manipulation
Estimating the full 6D pose (3D position and 3D orientation) of an object grasped in a robotic hand by combining three complementary sensing modalities: external vision, whole-hand tactile readings, and proprioceptive joint state. Fusing the modalities resolves the occlusions and ambiguities that defeat vision alone, giving the accurate in-hand pose that dexterous manipulation requires; models are trained on large-scale simulated data and transferred to real hardware.
2501.00510
VinT-6D is the first large-scale multi-modal dataset for estimating the 6D pose (position and orientation) of an object held in a robotic hand, combining synchronized vision, whole-hand tactile sensi…