D
Data
Text
MI-DPO recasts direct preference optimization as KL-regularized reinforcement learning against a learnable prior distribution zeta rather than a fixed reference policy, motivated as a rate-distortion / mutual-information penalty between prompts and responses. Choosing zeta appropriately recovers DPO, DICE, entropy-controllable DPO, SimPO, R-DPO, TDPO, TIS-DPO and SparsePO as special cases, and the paper proves that jointly optimizing the prior yields a strictly lower loss than any fixed prior.