Claude with Extended Thinking: When Deep Reasoning Is Worth the Extra Cost

By Carlos Montiel | Enterprise AI Specialist
Leer en español →
Published: 2026-07-28 | By: Carlos Montiel | Reading time: ~4 minutes

Turning on extended thinking without a second thought can multiply your Claude bill by 6x on the same task, without necessarily improving the answer. The key is knowing which tasks it actually pays off on.

How it's billed, in concrete numbers

Thinking tokens are billed at the output rate. A typical extended-thinking request might use between 5,000 and 20,000 thinking tokens plus 1,000-4,000 response tokens, costing roughly $0.09 to $0.36 per request, compared to $0.015-$0.06 in standard mode — up to six times more expensive for the same question. With Opus, a heavy-thinking request (50,000 thinking tokens) costs around $3.75 in reasoning tokens alone.

When to turn it on — and when not to

Turn on extended thinking for genuinely complex reasoning tasks: competition-level math (AIME), coding problems with multiple simultaneous constraints, extensive analysis with contradictory evidence. Avoid it for direct generation tasks (summarization, classification, simple Q&A) where the overhead adds cost and latency without improving result quality — it's exactly the same "don't use a hammer for everything" principle that applies to choosing between models.

The budget control that landed in March 2026

The `budget_tokens` field launched in March 2026: you set a ceiling on how much Claude can think before it has to answer. The model allocates that budget however it wants — it simply can't exceed it. It's the first time Anthropic gave developers direct control over thinking spend, instead of leaving it entirely to the model's discretion.

How much budget to allocate in practice

The practical recommendation: budget between 5,000 and 10,000 thinking tokens for most use cases, and 20,000+ only for genuinely hard problems. It's a decision worth making per task type, not globally for the whole application — a classification endpoint doesn't need the same thinking budget as a complex legal-analysis endpoint.

Build your criteria with evidence, not intuition

The most important recommendation is to build your own evidence-based heuristics for when the cost is worth it — running the same set of representative tasks with and without extended thinking, and measuring whether the quality difference justifies the cost difference for your specific use case, instead of enabling extended thinking "just in case" across the whole application.

Carlos Montiel
Enterprise AI Solutions Architect
Specialist in LLMs, Agents, and Orchestration
guatemalia.com/en/#contact · info@guatemalia.com

Need to implement AI at your company?

Carlos Montiel is an enterprise AI solutions architect. He implements LLMs, Agents, RAG, and orchestrators for companies across Guatemala and Latin America. Reach out for a consultation.

Contact Carlos Montiel

info@guatemalia.com