Conceptual

CLIP-Based Unanswerable Question Detection in Visual Question Answering

A lightweight technique that equips a frozen vision-language model to detect when a visual-question-answering query cannot be answered from the image, such as a question about an object that is not present, and to abstain rather than hallucinate an answer. It extracts CLIP-based question-image alignment signals and trains only a few added layers, preserving the base model's original performance on answerable tasks.