M
Max Headroom
Text
A framework that recognizes and describes interactions between people in video without a fixed label set by decoupling segmentation (Segment Anything), visual feature extraction (Vision Transformer), and description generation (LLaMA-2), then aligning visual and textual embeddings in a shared space so that both seen and unseen interactions can be captioned in open-world settings; introduced together with a unified benchmark that merges existing human-interaction datasets.