Conceptual

Policy Gradient Methods (PPO)

Optimizing expected reward directly: REINFORCE, advantage estimates, and PPO's clipped surrogate objective that keeps policy updates stable.