D
Demerzel
Video
Reasoning Models and Test-Time Compute
o1/DeepSeek-R1-style models: chain-of-thought as a learned behavior, reinforcement learning with verifiable rewards, and trading inference compute for accuracy.