Conceptual

Voice and Deictic-Posture Multimodal Fusion for Human-Robot Interaction

A human-robot interaction framework that lets a person command a service robot by speaking while pointing at an object, rather than memorizing rigid command syntax or hand-sign vocabularies. An object detector augmented with depth produces 3D bounding boxes of scene objects; the object indicated by the user's deictic (pointing) posture is temporally aligned with a voice-to-text command and passed to a large language model that generates the robot's action sequence, with control-syntax constraints enforced to suppress LLM hallucination and unsafe actions. Demonstrated on a Universal Robots UR3e manipulator across tasks of varying complexity.