J
jeremy
Video
How Large Language Models Use Test Time Compute to Reason via Reinforcement Learning
Thinking models operate under a theoretical framework where performance scaling at inference time (test-time compute) is achieved by expanding the generation horizon to produce extended reasoning tra…