Free-Form 6D Motion Control of Camera and Objects in Video Generation
Independently controlling the 6D poses (rotation and translation) and trajectories of both the camera and each dynamic object in text-to-video generation. Addresses the missing-data bottleneck with SynFMC, a synthetic dataset built by a rule-based pipeline that selects assets, assigns camera and per-object motion types, renders 3D animation sequences, and emits video with full 6D pose annotations and segmentation masks, then conditions a video generator on those trajectories for disentangled, free-form motion control.
Free-Form Motion Control: Controlling the 6D Poses of Camera and Objects in Video Generation
Controlling the movements of both the camera and multiple dynamic objects in generated video is hard, largely because no dataset provides comprehensive 6D pose annotations for camera and objects toge…