Pricing reviewed September 7, 2026Review dates and sources appear on each record.

Frontier models

Choose the model. Not the mythology.

A current, practical comparison of the leading cloud model families—what each is good at, what API use costs, and where to get access.

Same workload, clear costs

Compare three models.

Standard paid text API rates · reviewed September 7, 2026

$0.17

Estimated token cost per request

Input / 1M
$2
Output / 1M
$10
Model limits
1M total / 128K output

Daily coding and agent workflows

Anthropic model details ↗
$0.0638

Estimated token cost per request

Input / 1M
$0.75
Output / 1M
$3.75
Model limits
1,048,576 input / 65,536 output

Fast multimodal work and grounded research

50% promotional rate through December 31; $1.50 / $7.50 from January 1, 2027.

Google model details ↗
$0.018

Estimated token cost per request

Input / 1M
$0.2
Output / 1M
$1.2
Model limits
1.05M total; 922K input / 128K output

High-volume extraction, routing, and simpler coding tasks

OpenAI model details ↗

This is a cost calculation, not a quality ranking or a guarantee that your request fits a model’s limits. Long-context rate changes apply to the whole request. Check the limits above; input and output share a total window where specified. Caching, reasoning/output usage, tools, retries, media, taxes, and regional rates can change your bill.

Fast recommendations

Good starting points by job.

Editorial candidates based on documented capabilities, not a measured leaderboard. Read our method.

Hardest professional work

GPT-6 Astra, Claude Opus 5, or Claude Fable 5.1

Pay the premium only when deeper reasoning changes the outcome.

Daily agent and coding work

Claude Sonnet 5 or GPT-5.6 Terra

Strong follow-through without defaulting to the most expensive tier.

Fast multimodal + grounded research

Gemini 3.8 Flash

A strong fit for documents, images, video, search, and high-volume work.

Current web and X context

Grok 4.6

Useful when native web/X search and quick agentic work are central.

High-volume routine tasks

GPT-5.6 Luna or Gemini 3.5 Flash-Lite

Classify, extract, rewrite, and route at a fraction of frontier cost.

Provider comparison

The current frontier, side by side.

Prices below are standard API text rates per 1M input / output tokens unless noted. Consumer subscriptions use separate plan pricing and usage limits.

Exact modelInput / 1MOutput / 1MConditions and limits
GPT-6 Astragpt-6-astraReviewed September 7, 2026$10$501.05M total; 922K input / 128K output

At >272,000 input tokens: $20 input / $75 output per 1M.

GPT-5.6 Solgpt-5.6-solReviewed September 7, 2026$4$201.05M total; 922K input / 128K output

At >272,000 input tokens: $8 input / $30 output per 1M.

Promotional rate confirmed at least through November 21; later rate is not established.

GPT-5.6 Terragpt-5.6-terraReviewed September 7, 2026$2$121.05M total; 922K input / 128K output

At >272,000 input tokens: $4 input / $18 output per 1M.

GPT-5.6 Lunagpt-5.6-lunaReviewed September 7, 2026$0.2$1.21.05M total; 922K input / 128K output

At >272,000 input tokens: $0.4 input / $1.8 output per 1M.

Claude Fable 5.1claude-fable-5-1Reviewed September 7, 2026$10$501M total / 128K output
Claude Opus 5claude-opus-5Reviewed September 7, 2026$5$251M total / 128K output
Claude Sonnet 5claude-sonnet-5Reviewed September 7, 2026$2$101M total / 128K output
Claude Haiku 4.5claude-haiku-4-5-20251001Reviewed September 7, 2026$1$5200K total / 64K output
Gemini 3.8 Flashgemini-3.8-flashReviewed September 7, 2026$0.75$3.751,048,576 input / 65,536 output

50% promotional rate through December 31; $1.50 / $7.50 from January 1, 2027.

Gemini 3.5 Flash-Litegemini-3.5-flash-liteReviewed September 7, 2026$0.3$2.51,048,576 input / 65,536 output
Gemini 3.1 Pro Previewgemini-3.1-pro-previewReviewed September 7, 2026$2$121,048,576 input / 65,536 output

At >200,000 input tokens: $4 input / $18 output per 1M.

Grok 4.6grok-4.6Reviewed September 7, 2026$2$6500K total

At 200,000 input tokens: $4 input / $12 output per 1M.

Grok Build 0.1grok-build-0.1Reviewed September 7, 2026$1$2256K total

At 200,000 input tokens: $2 input / $4 output per 1M.

Grok 4.3grok-4.3Reviewed September 7, 2026$1.25$2.51M total

At 200,000 input tokens: $2.5 input / $5 output per 1M.

Sonnet 5’s scheduled September price increase was cancelled; its $2 / $10 base rate is permanent as documented on the review date. Grok 4.6 does not support Batch. Other tiers and tool charges are separate.

Cost in plain English

Tokens are cheap. Unbounded workflows are not.

This rough example sends a 60K-token document and produces a 5K-token answer. It excludes cached input, tools, long-context multipliers, and retries.

The calculator above starts with this same 60K-input / 5K-output example and recalculates when you change the workload. Its prices come from the same dated records as the table.

Control the loop

Set maximum steps, tool-call budgets, timeouts, and explicit stop conditions before you let an agent work unattended.

Cache what repeats

System instructions, long reference material, and stable prefixes are exactly where provider caching can change the economics.

Route by difficulty

Use a smaller model for extraction and triage. Escalate only the ambiguous cases to the flagship.

How to get access

Use the app. Build with the API.

Try models as a person

Use the provider’s consumer app first. Upload a real document, test a difficult prompt, and learn the interaction model before you write code.

Use a model gateway

A gateway can make provider switching, logs, budgets, and fallback routing easier. It also becomes another security and reliability dependency—treat it accordingly.

  • Keep provider model IDs configurable
  • Never expose API keys in a browser app
  • Log cost, latency, and failure reason
  • Pin versions for repeatable production behavior

Make the choice durable

Build a tiny evaluation before a big integration.

Collect ten to thirty representative inputs. Define what good looks like. Run the same set whenever a model, prompt, tool, or retrieval system changes.

CorrectnessDid it reach the right result and use the supplied evidence?
CompletenessDid it cover the requirements without inventing new ones?
FormatCan downstream software reliably use the response?
CostWhat was the total cost including retries and tool calls?
LatencyDoes it feel fast enough for the actual user?
Choose a project to test

Beyond the shortlist

The frontier table is a decision guide. The catalog is the knowledgebase.

Search current and prior families, open weights, image, video, audio, retrieval, specialist models, and the major platforms where each is available.

Search every tracked family