Conceptual

Benchmarking Vision-Language Models on Visual Illusion Understanding

How to measure whether a vision-language model perceives visual illusions the way people do: assembling a dataset that mixes classic synthetic illusions with real-scene illusions, color-blindness plates, and trap images; probing it with true-or-false, multiple-choice, and open-ended description tasks against human baselines; and testing prompting strategies that make the model name an illusion's cause before describing content. Students learn why real-world and anti-overfitting images matter and how step-by-step reasoning changes model accuracy.