Conceptual

CAREBENCH: Fine-Grained Video Captioning and Retrieval Benchmark

A benchmark of 1,000 human-annotated videos with captions manually split into spatial and temporal parts, paired with two metrics (ReBias for retrieval, CapST for captioning) that expose spatiotemporal biases in video-language models, plus CARE, a unified MLLM baseline trained by two-stage supervised fine-tuning that performs both fine-grained captioning and retrieval.