Conceptual

Dynamic Scaling of Unit Tests for Code Reward Modeling

Research contribution: scaling the number of LLM-generated unit tests improves the pass/fail reward signal used to rerank and select correct code solutions (with larger gains on harder problems), delivered via CodeRM-8B, a fine-tuned lightweight unit-test generator, plus a difficulty-aware dynamic mechanism that allocates more unit tests to harder problems.