Conceptual

Iterative Self-Improvement of Language Models

A training paradigm in which a language model is repeatedly improved by learning from data it generates itself, typically by sampling candidate outputs, forming reward or preference signals over them, and updating the model via reinforcement or preference learning across successive iterations. A central difficulty is that continued training on self-generated data can collapse output diversity, motivating methods that preserve a broad set of solution paths.