Conceptual

Direct Preference Optimization (DPO)

Preference tuning without a separate reward model or RL loop: the closed-form reparameterization that turns the RLHF objective into a simple classification loss on preference pairs.