HieA2G: Hierarchical Alignment and Adaptive Counting for Generalized Referring Expression Comprehension
A vision-language grounding network for Generalized Referring Expression Comprehension (GREC), where a free-form text expression may refer to zero, one, or several objects rather than exactly one. HieA2G combines two ideas: a Hierarchical Multi-modal Semantic Alignment module that aligns text and image at three granularities -- word-object, phrase-object, and text-image -- for robust cross-modal understanding, and an Adaptive Grounding Counter that dynamically predicts how many targets the expression denotes (including none), trained with an auxiliary contrastive loss that clusters multi-modal features by their target count. This removes the fixed-single-output assumption of classic referring-expression methods and reaches state-of-the-art results across GREC, REC, phrase grounding, and referring/generalized referring segmentation.