Scale-wise Bidirectional Vision-Language Alignment for Referring Remote Sensing Segmentation
How to segment the specific aerial-image region named by a text phrase when objects span very different scales: aligning visual and linguistic features in both directions (not just language guiding vision) with learnable query tokens, selecting macro global-context and micro local-detail features dynamically, and using a text-conditioned aggregator to exchange information across scales between encoder and decoder. Students learn why the neglected vision-to-language flow and scale diversity limit prior methods.
JOURNAL OF LATEX CLASS FILES, JANUARY 2025 1 Scale-wise Bidirectional Alignment Network for
A computer-vision paper proposing SBANet (Scale-wise Bidirectional Alignment Network) for referring remote sensing image segmentation (RRSIS) - extracting the exact pixel region of an aerial image de…