LLM-Based Pose-Conditioned Sign Language Production and Translation
A unified approach to bidirectional sign-language systems built on a paired text-to-sign dataset with skeletal keypoints. A large language model with retrieval augmentation maps natural language to gestures and drives pose-conditioned video synthesis that coordinates hand gestures and facial expressions, while a companion self-supervised translation model converts signing back to text; the pose keypoints enable direct evaluation of production accuracy rather than back-translation alone.
2501.00765
Bidirectional sign-language systems need both Sign Language Production (turning text into signing) and Sign Language Translation (turning signing back into text), but progress is held back by the abs…