Conceptual

Benchmarking and Improving Cultural Understanding in Vision-Language Models

How to measure and reduce the Western-centric cultural bias of vision-language models: building a large-scale multimodal benchmark of culturally specific symbols, gestures, and artifacts spanning many countries, using it to expose systematic regional performance gaps across many models, and fine-tuning on culturally rich image-text data to raise cultural understanding while preserving general capabilities and generalizing across cultures, continents, and datasets.