Claude Opus 4.6
anthropic/claude-opus-4-6
The last Opus to accept sampling parameters.
Claude Opus 4.6 is the oldest Opus still in the current line. It is the last one that accepts temperature, top_p and top_k, and the last that still honours a manual thinking budget alongside adaptive thinking — useful if you need a hard ceiling on thinking spend while you migrate. 1M context, 128K output.
- ModalitiesText + Imageout: Text
- Price$5 / $25per 1M in / out
- Context1M128K max output
- ReleasedNot publishedcutoff May 2025
Providers
Every routable endpoint for this model, with the rates the gateway bills against. One endpoint today, so there is no routing decision to make — when there is more than one, the router picks between them on price, health and latency.
| Provider | Input / M | Output / M | Cache read / M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|
| Anthropic · Messages APIclaude-opus-4-6 · global | $5 | $25 | $0.50 | 690ms | 70 tok/s | 99.92% |
Effective Pricing Computed
What input tokens actually cost once prompt caching is working. This is arithmetic over Anthropic's published prices, not a measurement — cached input bills at $0.50 per million instead of $5, so the blended rate falls as more of your prefix is reused: list × (1 − hit) + cache read × hit.
Chart requires JavaScript. The figures are in the table below.
- Blended input rate
- List price, no caching
View as table
| Cache hit ratio | 0% | 25% | 50% | 75% | 90% |
|---|---|---|---|---|---|
| Blended input rate | $5 | $3.88 | $2.75 | $1.63 | $0.95 |
| List price, no caching | $5 | $5 | $5 | $5 | $5 |
| Cache hit ratio | Blended input / M | Saving | 1M in + 1M out |
|---|---|---|---|
| 0% | $5 | — | $30 |
| 25% | $3.88 | 22% | $28.88 |
| 50% | $2.75 | 45% | $27.75 |
| 75% | $1.63 | 68% | $26.63 |
| 90% | $0.95 | 81% | $25.95 |
Performance Sample data
Throughput, round-trip latency and time to first token. Nothing here has been measured — there are no probes running and no traffic to learn from yet, so these are placeholders showing what the hydration layer will fill in.
- 70tokens / secondoutput, median
- 690mstime to first tokenmedian
- 690msp50 latencyend to end
- 2140msp99 latencyend to end
Throughput across Anthropic
Uptime Sample data
Share of requests that completed without an upstream error, over a rolling 30-day window. Sample data — the gateway has not made enough requests to have an opinion.
- 99.92%success raterolling 30d
- 0.08%error raterolling 30d
Uptime across Anthropic
Apps Sample data
How many distinct applications sent traffic to this model each week. We deliberately do not list app names: the per-app breakdown arrives with label passthrough on the event stream, and inventing customer names to fill the gap would be worse than an empty section.
Activity Sample data
Token volume and request traffic over the trailing four weeks, rolled up from the event stream. Sample data until the stream has something real in it.
Chart requires JavaScript. The figures are in the table below.
View as table
| Metric | 26 Apr | 27 Apr | 28 Apr | 29 Apr | 30 Apr | 1 May | 2 May | 3 May | 4 May | 5 May | 6 May | 7 May | 8 May | 9 May | 10 May | 11 May | 12 May | 13 May | 14 May | 15 May | 16 May | 17 May | 18 May | 19 May | 20 May | 21 May | 22 May | 23 May | 24 May | 25 May | 26 May | 27 May | 28 May | 29 May | 30 May | 31 May | 1 Jun | 2 Jun | 3 Jun | 4 Jun | 5 Jun | 6 Jun | 7 Jun | 8 Jun | 9 Jun | 10 Jun | 11 Jun | 12 Jun | 13 Jun | 14 Jun | 15 Jun | 16 Jun | 17 Jun | 18 Jun | 19 Jun | 20 Jun | 21 Jun | 22 Jun | 23 Jun | 24 Jun | 25 Jun | 26 Jun | 27 Jun | 28 Jun | 29 Jun | 30 Jun | 1 Jul | 2 Jul | 3 Jul | 4 Jul | 5 Jul | 6 Jul | 7 Jul | 8 Jul | 9 Jul | 10 Jul | 11 Jul | 12 Jul | 13 Jul | 14 Jul | 15 Jul | 16 Jul | 17 Jul | 18 Jul | 19 Jul | 20 Jul | 21 Jul | 22 Jul | 23 Jul | 24 Jul |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Tokens | 8.2M | 11M | 11M | 12M | 14M | 14M | 9.8M | 8.8M | 11M | 12M | 13M | 15M | 15M | 10M | 9.3M | 12M | 12M | 14M | 16M | 16M | 11M | 9.8M | 13M | 13M | 15M | 17M | 17M | 12M | 10M | 13M | 14M | 16M | 17M | 18M | 12M | 11M | 14M | 14M | 16M | 18M | 19M | 13M | 11M | 15M | 15M | 17M | 19M | 20M | 13M | 12M | 15M | 16M | 18M | 20M | 21M | 14M | 12M | 16M | 17M | 19M | 21M | 21M | 14M | 13M | 17M | 17M | 20M | 22M | 22M | 15M | 13M | 17M | 18M | 21M | 23M | 23M | 15M | 14M | 18M | 19M | 21M | 24M | 24M | 16M | 14M | 19M | 20M | 22M | 25M | 9.4M |
| Week | Tokens | Requests | Apps |
|---|---|---|---|
| 20 Jul 2026 | 154.0M | 160.0K | 95 |
| 13 Jul 2026 | 168.0M | 174.0K | 101 |
| 6 Jul 2026 | 181.0M | 188.0K | 106 |
| 29 Jun 2026 | 195.0M | 202.0K | 110 |
This model ranks #8 across the catalog this week — see the full rankings.
Call it
Standard OpenAI client, one changed line. Streaming is the default path — non-streaming is just collecting the stream.
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.aigate.dev/v1",
apiKey: process.env.AIGATE_API_KEY,
});
const stream = await client.chat.completions.create({
model: "anthropic/claude-opus-4-6",
messages: [{ role: "user", content: "Hello" }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}Capabilities
- Streaming
- Tool use
- Vision
- PDF input
- Citations
- Adaptive thinking
- Extended thinking
- Effort control
- Prompt caching
- Compaction
- Batch API
- 1M context
Accepted parameters
- max_tokens
- stream
- stop_sequences
- system
- tools
- tool_choice
- thinking
- effort
- temperature
- top_p
- top_k
Specification
- Model ID
- anthropic/claude-opus-4-6
- Upstream ID
- claude-opus-4-6
- Context window
- 1,000,000 tokens
- Max output
- 128,000 tokens
- API dialect
- anthropic — translated for you
- Knowledge cutoff
- May 2025
- Min cacheable prefix
- 4,096 tokens
- Also available on
- claude-api, bedrock, google-cloud
FAQ
What does Claude Opus 4.6 cost?
$5 per million input tokens and $25 per million output tokens — Anthropic's published list price. AI Gate adds a flat 5% routing fee on top, billed from prepaid credits. Cached input is far cheaper at $0.50 per million tokens.
What is the context window for Claude Opus 4.6?
1,000,000 tokens of context, with up to 128,000 tokens in a single response. Its knowledge is most reliable through May 2025.
How do I call Claude Opus 4.6?
Point any OpenAI-compatible client at https://api.aigate.dev/v1 and pass "anthropic/claude-opus-4-6" as the model string. There is no SDK to install — we translate the request into Anthropic's anthropic dialect on the way out and the response stream back on the way in.
How is prompt caching billed?
Reading from cache costs $0.50 per million tokens, one tenth of the list input price. Writing to a five-minute cache costs $6.25 (1.25× list), and a one-hour cache costs $10 (2× list). Prefixes shorter than 4,096 tokens will not cache at all.
Which API dialect does Claude Opus 4.6 speak?
The anthropic dialect. You never write it: the gateway accepts OpenAI-shaped requests and translates them at the edge, streaming byte by byte rather than buffering. Switching to a model on a different dialect means changing the model string and nothing else.
Is Claude Opus 4.6 being retired?
No retirement date has been announced. Anthropic publishes deprecations ahead of time, and this page tracks that field, so a date will appear here if one is set.
Closest alternatives
By input price. Switch between any of them by changing the model string.