Conceptual

Zero-Shot Image Safety Judgment Using Constitutional Multimodal Language Models

A training-free pipeline that uses a pre-trained multimodal large language model to decide whether an image violates a set of safety rules, without collecting human-labeled examples. Students learn how subjective safety rules are objectified into checkable criteria, how a CLIP image-text similarity filter prunes irrelevant rules for efficiency, how each rule is decomposed into logically complete precondition chains that make long rules tractable, and how debiased token probabilities correct language-prior and spatial biases to yield a confident verdict, escalating to cascaded chain-of-thought reasoning only when needed.