D
Demerzel
Video
Policy Gradient Methods (PPO)
Optimizing expected reward directly: REINFORCE, advantage estimates, and PPO's clipped surrogate objective that keeps policy updates stable.