Conceptual

Part-Level Cross-Modal Correspondence for Text-Based Person Search

How to match free-form text descriptions to person images by aligning fine-grained body-part information across the image and text modalities. Students learn how a coarse-to-fine attention-based encoder-decoder aligns the two modalities without explicit alignment labels, and how a commonality-based margin ranking loss learns discriminative part details using only person-ID supervision.