Conceptual

Probing and Reducing Visual Language Priors in Vision-Language Models

Introduces ViLP, a benchmark that pairs each question with deliberately out-of-distribution synthesized images and a text-prior 'distractor' answer, isolating when a Vision-Language Model answers from learned language priors rather than the pixels. Proposes Image-DPO, a self-improving method that builds good/bad image pairs via controlled corruptions for DPO-style training to increase a model's reliance on visual input, shown to optimize an upper bound of the RLHF objective.