Conceptual

Scale-wise Bidirectional Vision-Language Alignment for Referring Remote Sensing Segmentation

How to segment the specific aerial-image region named by a text phrase when objects span very different scales: aligning visual and linguistic features in both directions (not just language guiding vision) with learnable query tokens, selecting macro global-context and micro local-detail features dynamically, and using a text-conditioned aggregator to exchange information across scales between encoder and decoder. Students learn why the neglected vision-to-language flow and scale diversity limit prior methods.