Conceptual

Alignment and Misalignment in Large Vision-Language Models

How large vision-language models map visual and textual features into a shared representation, and how that alignment fails. Covers the representational and behavioral aspects of alignment and a taxonomy of misalignment at the object, attribute, and relational levels, with its origins traced to the data, model, and inference stages, plus mitigation strategies grouped into parameter-frozen and parameter-tuning approaches, all framed through model explainability.