Conceptual

Face-Human-Bench: Evaluating Face and Human Understanding in Multimodal Assistants

A benchmark and hierarchical (three-level) ability taxonomy for measuring how well multimodal large language models understand faces and humans, spanning tasks such as face recognition, expression and age estimation, human attribute/action recognition, spatial and social relation understanding, person re-identification, and face-attack detection. Built from public datasets via a semi-automatic pipeline (1800-problem dev and test sets, English and Chinese) and used to evaluate 25 MLLMs, studying inter-ability correlation, target relative position, Chain-of-Thought prompting effects, and where specialist models remain necessary.