D
Demerzel
Video
Direct Preference Optimization (DPO)
Preference tuning without a separate reward model or RL loop: the closed-form reparameterization that turns the RLHF objective into a simple classification loss on preference pairs.