Estimated token cost per request
- Input / 1M
- $2
- Output / 1M
- $10
- Model limits
- 1M total / 128K output
Daily coding and agent workflows
Anthropic model details ↗Frontier models
A current, practical comparison of the leading cloud model families—what each is good at, what API use costs, and where to get access.
Same workload, clear costs
Standard paid text API rates · reviewed September 7, 2026
Estimated token cost per request
Daily coding and agent workflows
Anthropic model details ↗Estimated token cost per request
Fast multimodal work and grounded research
50% promotional rate through December 31; $1.50 / $7.50 from January 1, 2027.
Google model details ↗Estimated token cost per request
High-volume extraction, routing, and simpler coding tasks
OpenAI model details ↗This is a cost calculation, not a quality ranking or a guarantee that your request fits a model’s limits. Long-context rate changes apply to the whole request. Check the limits above; input and output share a total window where specified. Caching, reasoning/output usage, tools, retries, media, taxes, and regional rates can change your bill.
Fast recommendations
Editorial candidates based on documented capabilities, not a measured leaderboard. Read our method.
Pay the premium only when deeper reasoning changes the outcome.
Strong follow-through without defaulting to the most expensive tier.
A strong fit for documents, images, video, search, and high-volume work.
Useful when native web/X search and quick agentic work are central.
Classify, extract, rewrite, and route at a fraction of frontier cost.
Provider comparison
Prices below are standard API text rates per 1M input / output tokens unless noted. Consumer subscriptions use separate plan pricing and usage limits.
| Exact model | Input / 1M | Output / 1M | Conditions and limits |
|---|---|---|---|
GPT-6 Astra ↗gpt-6-astraReviewed September 7, 2026 | $10 | $50 | 1.05M total; 922K input / 128K output At >272,000 input tokens: $20 input / $75 output per 1M. |
GPT-5.6 Sol ↗gpt-5.6-solReviewed September 7, 2026 | $4 | $20 | 1.05M total; 922K input / 128K output At >272,000 input tokens: $8 input / $30 output per 1M. Promotional rate confirmed at least through November 21; later rate is not established. |
GPT-5.6 Terra ↗gpt-5.6-terraReviewed September 7, 2026 | $2 | $12 | 1.05M total; 922K input / 128K output At >272,000 input tokens: $4 input / $18 output per 1M. |
GPT-5.6 Luna ↗gpt-5.6-lunaReviewed September 7, 2026 | $0.2 | $1.2 | 1.05M total; 922K input / 128K output At >272,000 input tokens: $0.4 input / $1.8 output per 1M. |
Claude Fable 5.1 ↗claude-fable-5-1Reviewed September 7, 2026 | $10 | $50 | 1M total / 128K output |
Claude Opus 5 ↗claude-opus-5Reviewed September 7, 2026 | $5 | $25 | 1M total / 128K output |
Claude Sonnet 5 ↗claude-sonnet-5Reviewed September 7, 2026 | $2 | $10 | 1M total / 128K output |
Claude Haiku 4.5 ↗claude-haiku-4-5-20251001Reviewed September 7, 2026 | $1 | $5 | 200K total / 64K output |
Gemini 3.8 Flash ↗gemini-3.8-flashReviewed September 7, 2026 | $0.75 | $3.75 | 1,048,576 input / 65,536 output 50% promotional rate through December 31; $1.50 / $7.50 from January 1, 2027. |
Gemini 3.5 Flash-Lite ↗gemini-3.5-flash-liteReviewed September 7, 2026 | $0.3 | $2.5 | 1,048,576 input / 65,536 output |
Gemini 3.1 Pro Preview ↗gemini-3.1-pro-previewReviewed September 7, 2026 | $2 | $12 | 1,048,576 input / 65,536 output At >200,000 input tokens: $4 input / $18 output per 1M. |
Grok 4.6 ↗grok-4.6Reviewed September 7, 2026 | $2 | $6 | 500K total At ≥200,000 input tokens: $4 input / $12 output per 1M. |
Grok Build 0.1 ↗grok-build-0.1Reviewed September 7, 2026 | $1 | $2 | 256K total At ≥200,000 input tokens: $2 input / $4 output per 1M. |
Grok 4.3 ↗grok-4.3Reviewed September 7, 2026 | $1.25 | $2.5 | 1M total At ≥200,000 input tokens: $2.5 input / $5 output per 1M. |
Sonnet 5’s scheduled September price increase was cancelled; its $2 / $10 base rate is permanent as documented on the review date. Grok 4.6 does not support Batch. Other tiers and tool charges are separate.
Cost in plain English
This rough example sends a 60K-token document and produces a 5K-token answer. It excludes cached input, tools, long-context multipliers, and retries.
The calculator above starts with this same 60K-input / 5K-output example and recalculates when you change the workload. Its prices come from the same dated records as the table.
Set maximum steps, tool-call budgets, timeouts, and explicit stop conditions before you let an agent work unattended.
System instructions, long reference material, and stable prefixes are exactly where provider caching can change the economics.
Use a smaller model for extraction and triage. Escalate only the ambiguous cases to the flagship.
How to get access
Use the provider’s consumer app first. Upload a real document, test a difficult prompt, and learn the interaction model before you write code.
Create a developer account, add billing, generate a project-scoped key, set a spend limit, and start with the provider’s official quickstart.
A gateway can make provider switching, logs, budgets, and fallback routing easier. It also becomes another security and reliability dependency—treat it accordingly.
Make the choice durable
Collect ten to thirty representative inputs. Define what good looks like. Run the same set whenever a model, prompt, tool, or retrieval system changes.