J
jeremy
8 minutes
Turn on thinking so a model can reason before it answers, steer how often and how deeply it thinks with an effort setting rather than a fixed token budget, decide which tasks justify the extra output tokens and latency, and read the thinking blocks that come back.