Face-Human-Bench: Evaluating Face and Human Understanding in Multimodal Assistants
A benchmark and hierarchical (three-level) ability taxonomy for measuring how well multimodal large language models understand faces and humans, spanning tasks such as face recognition, expression and age estimation, human attribute/action recognition, spatial and social relation understanding, person re-identification, and face-attack detection. Built from public datasets via a semi-automatic pipeline (1800-problem dev and test sets, English and Chinese) and used to evaluate 25 MLLMs, studying inter-ability correlation, target relative position, Chain-of-Thought prompting effects, and where specialist models remain necessary.
Face-Human-Bench: A Comprehensive Benchmark of Face and Human Understanding for Multi-modal
Introduces Face-Human-Bench, a benchmark for evaluating the face- and human-understanding abilities of multimodal large language models (MLLMs). Proposes a three-level hierarchical ability taxonomy s…