Conceptual

Direct Preference Optimization Without a Separate Reward Model

folding the preference objective into one classification loss removed the reward model and the RL loop from the common path