Conceptual

Group-Relative Policy Optimization and Verifiable Rewards

scoring a group of sampled completions against a checkable answer is what current reasoning post-training actually runs; expect churn