Conceptual

Symbolic World Models from Vision-Language-Model Predicates

A method for solving long-horizon robotic decision-making from only low-level skills and a handful of short-horizon image demonstrations. A pretrained vision-language model proposes a large set of candidate visual predicates (object properties and relations) and evaluates them directly from camera images; an optimization-based model-learning algorithm then selects a compact subset of these predicates and learns an abstract symbolic world model (operators over predicates). At test time the VLM builds a symbolic description of a novel scene and a search-based planner sequences low-level skills to reach a novel goal, giving zero-shot generalization without task-specific retraining.