Conceptual

Reinforcement Learning from Human Feedback (RLHF)

The InstructGPT recipe behind aligned chat models: supervised fine-tuning, reward-model training on human preference pairs, and PPO with a KL penalty against the reference model.