You deployed an o3-powered agent last week. The baseline pricing looked reasonable: $2 input, $8 output per million tokens. Your math said this feature would cost $50/month. Your bill was $312.
Welcome to reasoning token economics. OpenAI's newest models (o3, o4-mini, o3-mini) charge for something they don't explicitly advertise: thinking tokens. These aren't regular output tokens. They're the internal reasoning the model does before it writes an answer. And they cost exactly as much as output tokens — $8 per million for o3, $4.40 for o4-mini.
Here's what nobody tells you: the default reasoning budget is enormous. Your o3 agent thinks for 30-60 seconds before responding, consuming tens of thousands of tokens per request. Those tokens bill at full output rates. You do the math on 100 requests/day, and suddenly your $50/month feature costs $300+.
This isn't a bug. It's how reasoning models work. But it's also a builder problem that catches founders by surprise. Understanding reasoning token economics is the difference between a sustainable feature and a money pit.
What Changed: Thinking Tokens Now Have a Price Tag
For the last 18 months, reasoning models (o1, o1-preview) charged per token but didn't expose how many tokens reasoning consumed. Your invoice just said "o1: $X." You didn't know if you'd used 1 million tokens or 10 million — OpenAI bundled it together.
On Sept 28, 2026, OpenAI changed the game. They released o3, o4-mini, and o3-mini with a critical difference: thinking tokens are now visible and charged separately. Every API response now returns two token counts: input tokens and output tokens. The "output" number includes both regular output tokens and all the thinking tokens the model generated while reasoning.
Why this matters: Reasoning tokens cost the same as output tokens. For o3, that's $8 per million. If your model spends 50,000 reasoning tokens before writing a 100-token answer, you're paying for 50,100 tokens at the $8/M output rate — not the $2/M input rate.
The hidden multiplier: Default reasoning effort ("medium" mode) consumes 20,000-80,000 reasoning tokens per request depending on task complexity. High-effort reasoning for complex problems (research synthesis, multi-step coding, architecture decisions) can hit 200,000+ reasoning tokens per request.
Math: 200,000 reasoning tokens at $8/M = $1.60 per request. That's 80x the $2 input cost of your original prompt. One request costs more than 100 regular ChatGPT queries.
Real Cost Examples: What Reasoning Actually Costs
Scenario 1: Customer support chatbot using o3 for complex issues
Your chatbot normally answers FAQs in 50 tokens using GPT-4. For edge cases, you route to o3 with medium reasoning. Real request breakdown (from OpenAI API responses):
- Input tokens: 3,200 (customer history + question)
- Reasoning tokens: 45,000 (thinking through edge case resolution)
- Output tokens: 820 (final answer + recommendations)
- Total billed: 45,820 tokens at $8/M output rate = $0.37 per request Your GPT-4 baseline was $0.03. You're now 12x more expensive per request.
Scale: 200 edge case requests/month costs $74. Budget was $6.
Scenario 2: Code generation agent with o4-mini (cheaper reasoning model)
You built an agent that generates boilerplate code. Requests are:
- Input tokens: 2,100 (code context + requirements)
- Reasoning tokens: 18,000 (breaking down architecture, planning implementation)
- Output tokens: 2,400 (generated code)
- Total billed: 20,400 tokens at $4.40/M output rate = $0.09 per request o4-mini is cheaper than o3, so this looks reasonable. But you're still paying $0.09 vs $0.02 for a Claude Opus equivalent (reasoning built into training, not charged separately).
Scale: 5,000 code generation requests/month = $450. You budgeted $100.
Scenario 3: Research synthesis agent with o3 high-effort reasoning
Your agent reads 20 research papers and synthesizes findings. High-effort reasoning is mandatory for quality output:
- Input tokens: 180,000 (20 papers, each 9,000 tokens)
- Reasoning tokens: 420,000 (deep synthesis, cross-referencing, critical analysis)
- Output tokens: 6,200 (synthesis report)
- Total billed: 426,200 tokens at $8/M output rate = $3.41 per request One report costs $3.41. You generate 50 reports/month = $170. Acceptable. But if you iterate on prompts during development (20 test runs), development alone costs $68 before reaching production.
The trap: High-effort reasoning sounds necessary for quality, so you set it to "auto" and let the model decide. Result: Every request reasons hard, and costs scale with request volume, not value delivered.
How Reasoning Budgets Work (And How to Control Them)
OpenAI offers three reasoning effort levels for o3:
Low effort (default for chat): 1,000-5,000 reasoning tokens. Fast (3-7 seconds), cheap. Best for straightforward questions, summarization, basic coding.
Medium effort: 20,000-80,000 reasoning tokens. Balanced (15-30 seconds), moderate cost. Best for moderately complex problems, edge cases, multi-step logic.
High effort: 100,000-200,000+ reasoning tokens. Slow (30-60+ seconds), expensive. Best for novel research, architecture decisions, complex multi-step reasoning.
The problem: Most founders don't adjust these settings. They use the OpenAI Playground default (medium), which is fine for occasional use but catastrophic at scale.
For API requests, you specify reasoning effort per request:
"model": "o3",
"thinking": {
"type": "enabled",
"budget_tokens": 50000 # cap reasoning at 50k tokens
}
Without explicit caps, the model can use up to the default budget. Without effort setting, it defaults to medium. You're paying full price on autopilot.
For o4-mini, the cost structure is similar but cheaper ($4.40/M vs $8/M). For o3-mini, you get different performance tiers (default matches o1-mini reasoning ability). The principle is identical: reasoning tokens bill at output rates, and the bill surprises founders who didn't budget for it.
When Reasoning Costs Spiral (And When They Don't)
Reasoning costs spiral when:
Your agent makes many requests and reasoning effort is high by default. Customer support with 1,000 daily requests at medium effort: you're spending $2,000+/month just on reasoning tokens. That's unsustainable for most founders.
You're using o3 for tasks that don't need reasoning. Summarizing text, extracting structured data, or answering FAQs don't benefit from 50,000 reasoning tokens. Using o3 here is paying for capability you don't use.
You're in development and iterating on prompts. Each test run with high-effort reasoning costs $0.50-$3. Twenty iterations cost $10-60. Multiply across a team of 3 building agents, and development costs exceed production costs.
You're using reasoning for deterministic tasks (formatting, templating, simple routing). These don't need o3's reasoning at all. You're overpaying 10x vs GPT-4 or Claude.
Reasoning costs make sense when:
The task genuinely requires multi-step reasoning. Research synthesis, architecture decisions, complex debugging, novel problem-solving. These are rare enough that high reasoning cost is justified.
You've capped reasoning budgets per request and profiled real usage. "Medium effort" for this task actually costs $0.08 per request, but value delivered is $0.50+. ROI is positive.
You're using o4-mini, not o3, and reasoning effort is low. Reasoning tokens at $4.40/M are more defensible than $8/M. If you're getting 24% faster results than o1-mini at this price, the trade-off works.
You're not iterating on this feature anymore. Development is sunk cost; production reasoning tokens are the only ongoing expense. If usage is low (10 requests/day), even high reasoning cost is acceptable ($2-3/month).
Decision Framework: Is o3 Reasoning Worth It?
Ask these three questions:
1. Does this task actually need deep reasoning? If you can describe the task in a sentence and the answer is deterministic (yes/no, classification, simple transformation), reasoning is wasted. Use GPT-4 or Claude. If the task requires synthesis, novel problem-solving, or multi-step logic that the model needs to think through, reasoning is necessary.
2. Can you afford the cost at scale? Calculate: (expected requests/month) × (expected reasoning tokens per request / 1M) × ($4.40 or $8 depending on model). If that's >10% of your feature's monthly budget, you're overpaying. Consider GPT-4 (no reasoning cost, 80% cheaper) or Claude (reasoning baked in, no separate charge). Use o3 only if reasoning unlocks value that justifies the cost.
3. Have you profiled actual usage and reasoning token consumption? Deploy a test version, log reasoning tokens for 100-500 real requests, calculate average and p95 reasoning tokens per request. Compare to your budget. If average reasoning is 10,000 tokens and you budgeted for 5,000, you're already 2x over. Adjust before scaling.
Recommended approach for founders:
Start with GPT-4 or Claude for most workloads. Identify the 10-20% of requests that truly need reasoning. Route only those to o3, with explicit reasoning budget caps (20,000 tokens max). Profile monthly costs. If reasoning costs are <10% of feature budget and performance improvement justifies it, keep o3. If reasoning costs are >20% of budget and alternatives exist, switch to Claude (reasoning built-in) or GPT-4 (70% cheaper, sufficient for most tasks).
The Bigger Pattern
OpenAI's pricing transparency (showing reasoning tokens separately) is actually good for founders — it forces you to be honest about costs. But it's also a trap if you don't understand what you're buying.
Reasoning token economics reveal a fundamental trade-off: thinking time costs money. The harder your model thinks, the higher your bill. This is mathematically unavoidable for any reasoning-capable system. The question isn't "how do I avoid reasoning token costs?" It's "which tasks actually need reasoning, and which am I overpaying for?"
For most founder-built agents (customer support, document processing, simple routing, code generation), reasoning is unnecessary overkill. Claude's built-in reasoning through training means you pay one price per token, reasoning included. GPT-4 is cheaper and sufficient. o3 should be reserved for the rare tasks where deep reasoning unlocks value — research synthesis, novel architecture decisions, complex debugging — and even then, only if you've profiled costs and they're defensible.
The founders winning with o3 aren't using it for everything. They're using it for 5-10% of requests where reasoning provides outsized value, and they've capped reasoning budgets to keep costs predictable. Everyone else is paying $300/month to run customer support queries that could be handled for $30 with simpler models.
Cost Control Checklist
Before deploying o3 to production:
- Profile reasoning tokens for 100 real requests. Calculate average and p95.
- Set explicit reasoning budget caps based on profiling:
"budget_tokens": <p95 + 20%> - Compare total monthly cost to budget. If >15% of feature budget, evaluate alternatives.
- Set up billing alerts. If monthly reasoning spend exceeds projection by >20%, trigger a review.
- Log reasoning tokens per request and track trends. Degradation in reasoning efficiency (same task, more tokens) is a warning sign.
- Test low-effort reasoning for your task. Does "low" produce acceptable results at 80% cost savings? Use it.
- For development and testing, use a separate API key with monthly spend limits. Prevent runaway test costs.
- Measure value delivered per reasoning token spent. Is the outcome better enough to justify the cost? If not, downgrade to GPT-4.
Next Steps
- Audit your o3 requests. If you've deployed o3 already, pull API logs and calculate average reasoning tokens per request and total monthly spend. Compare to budget.
- Profile alternatives. Test GPT-4 and Claude for the same tasks. Compare reasoning tokens (for o3), cost per request, and output quality. Build a comparison matrix.
- Set budgets. If you decide to keep o3, set explicit reasoning budget caps per request type. Document your reasoning effort decisions.
- Alert on overruns. OpenAI's billing dashboard allows alerts; set one at 110% of monthly projection. Catch cost overruns early.
For a detailed breakdown of reducing API costs, see our guide on reducing Claude API costs.