Claude Opus 5
anthropic/claude-opus-5
For complex agentic coding and enterprise work.
Claude Opus 5 is Anthropic's flagship model for complex agentic coding and enterprise work — a step change over Opus 4.8 on deep reasoning, long-horizon autonomy and test-time compute scaling, at the same price. Thinking is on by default and the full effort ladder runs from low through max, so you tune depth per request rather than per model. It holds a 1M-token context window and writes up to 128K tokens in a single response.
- ModalitiesText + Imageout: Text
- Price$5 / $25per 1M in / out
- Context1M128K max output
- ReleasedNot publishedcutoff May 2026
Providers
Every routable endpoint for this model, with the rates the gateway bills against. One endpoint today, so there is no routing decision to make — when there is more than one, the router picks between them on price, health and latency.
| Provider | Input / M | Output / M | Cache read / M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|
| Anthropic · Messages APIclaude-opus-5 · global | $5 | $25 | $0.50 | 780ms | 62 tok/s | 99.94% |
Effective Pricing Computed
What input tokens actually cost once prompt caching is working. This is arithmetic over Anthropic's published prices, not a measurement — cached input bills at $0.50 per million instead of $5, so the blended rate falls as more of your prefix is reused: list × (1 − hit) + cache read × hit.
Chart requires JavaScript. The figures are in the table below.
- Blended input rate
- List price, no caching
View as table
| Cache hit ratio | 0% | 25% | 50% | 75% | 90% |
|---|---|---|---|---|---|
| Blended input rate | $5 | $3.88 | $2.75 | $1.63 | $0.95 |
| List price, no caching | $5 | $5 | $5 | $5 | $5 |
| Cache hit ratio | Blended input / M | Saving | 1M in + 1M out |
|---|---|---|---|
| 0% | $5 | — | $30 |
| 25% | $3.88 | 22% | $28.88 |
| 50% | $2.75 | 45% | $27.75 |
| 75% | $1.63 | 68% | $26.63 |
| 90% | $0.95 | 81% | $25.95 |
Performance Sample data
Throughput, round-trip latency and time to first token. Nothing here has been measured — there are no probes running and no traffic to learn from yet, so these are placeholders showing what the hydration layer will fill in.
- 62tokens / secondoutput, median
- 780mstime to first tokenmedian
- 780msp50 latencyend to end
- 2400msp99 latencyend to end
Throughput across Anthropic
Uptime Sample data
Share of requests that completed without an upstream error, over a rolling 30-day window. Sample data — the gateway has not made enough requests to have an opinion.
- 99.94%success raterolling 30d
- 0.06%error raterolling 30d
Uptime across Anthropic
Apps Sample data
How many distinct applications sent traffic to this model each week. We deliberately do not list app names: the per-app breakdown arrives with label passthrough on the event stream, and inventing customer names to fill the gap would be worse than an empty section.
Activity Sample data
Token volume and request traffic over the trailing four weeks, rolled up from the event stream. Sample data until the stream has something real in it.
Chart requires JavaScript. The figures are in the table below.
View as table
| Metric | 26 Apr | 27 Apr | 28 Apr | 29 Apr | 30 Apr | 1 May | 2 May | 3 May | 4 May | 5 May | 6 May | 7 May | 8 May | 9 May | 10 May | 11 May | 12 May | 13 May | 14 May | 15 May | 16 May | 17 May | 18 May | 19 May | 20 May | 21 May | 22 May | 23 May | 24 May | 25 May | 26 May | 27 May | 28 May | 29 May | 30 May | 31 May | 1 Jun | 2 Jun | 3 Jun | 4 Jun | 5 Jun | 6 Jun | 7 Jun | 8 Jun | 9 Jun | 10 Jun | 11 Jun | 12 Jun | 13 Jun | 14 Jun | 15 Jun | 16 Jun | 17 Jun | 18 Jun | 19 Jun | 20 Jun | 21 Jun | 22 Jun | 23 Jun | 24 Jun | 25 Jun | 26 Jun | 27 Jun | 28 Jun | 29 Jun | 30 Jun | 1 Jul | 2 Jul | 3 Jul | 4 Jul | 5 Jul | 6 Jul | 7 Jul | 8 Jul | 9 Jul | 10 Jul | 11 Jul | 12 Jul | 13 Jul | 14 Jul | 15 Jul | 16 Jul | 17 Jul | 18 Jul | 19 Jul | 20 Jul | 21 Jul | 22 Jul | 23 Jul | 24 Jul |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Tokens | 197M | 248M | 223M | 220M | 243M | 277M | 214M | 209M | 263M | 237M | 234M | 259M | 294M | 227M | 221M | 278M | 251M | 248M | 275M | 312M | 240M | 234M | 293M | 264M | 262M | 290M | 330M | 253M | 246M | 308M | 278M | 276M | 306M | 347M | 267M | 258M | 323M | 292M | 290M | 322M | 365M | 280M | 270M | 338M | 305M | 305M | 338M | 383M | 293M | 282M | 352M | 319M | 319M | 355M | 401M | 306M | 294M | 367M | 332M | 333M | 371M | 419M | 319M | 306M | 382M | 346M | 347M | 387M | 437M | 332M | 318M | 396M | 359M | 362M | 403M | 455M | 345M | 330M | 411M | 373M | 376M | 420M | 473M | 358M | 342M | 425M | 386M | 390M | 436M | 187M |
| Week | Tokens | Requests | Apps |
|---|---|---|---|
| 20 Jul 2026 | 3.15B | 2.4M | 980 |
| 13 Jul 2026 | 2.72B | 2.1M | 910 |
| 6 Jul 2026 | 2.38B | 1.8M | 845 |
| 29 Jun 2026 | 2.05B | 1.6M | 774 |
This model ranks #2 across the catalog this week — see the full rankings.
Call it
Standard OpenAI client, one changed line. Streaming is the default path — non-streaming is just collecting the stream.
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.aigate.dev/v1",
apiKey: process.env.AIGATE_API_KEY,
});
const stream = await client.chat.completions.create({
model: "anthropic/claude-opus-5",
messages: [{ role: "user", content: "Hello" }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}Capabilities
- Streaming
- Tool use
- Vision
- High-res vision
- PDF input
- Citations
- Structured outputs
- Adaptive thinking
- Effort control
- Task budgets
- Fast mode
- Prompt caching
- Compaction
- Batch API
- 1M context
Accepted parameters
- max_tokens
- stream
- stop_sequences
- system
- tools
- tool_choice
- thinking
- effort
- output_config.format
- speed
Specification
- Model ID
- anthropic/claude-opus-5
- Upstream ID
- claude-opus-5
- Context window
- 1,000,000 tokens
- Max output
- 128,000 tokens
- API dialect
- anthropic — translated for you
- Knowledge cutoff
- May 2026
- Min cacheable prefix
- 512 tokens
- Also available on
- claude-api, bedrock, google-cloud
FAQ
What does Claude Opus 5 cost?
$5 per million input tokens and $25 per million output tokens — Anthropic's published list price. AI Gate adds a flat 5% routing fee on top, billed from prepaid credits. Cached input is far cheaper at $0.50 per million tokens.
What is the context window for Claude Opus 5?
1,000,000 tokens of context, with up to 128,000 tokens in a single response. Its knowledge is most reliable through May 2026.
How do I call Claude Opus 5?
Point any OpenAI-compatible client at https://api.aigate.dev/v1 and pass "anthropic/claude-opus-5" as the model string. There is no SDK to install — we translate the request into Anthropic's anthropic dialect on the way out and the response stream back on the way in.
How is prompt caching billed?
Reading from cache costs $0.50 per million tokens, one tenth of the list input price. Writing to a five-minute cache costs $6.25 (1.25× list), and a one-hour cache costs $10 (2× list). Prefixes shorter than 512 tokens will not cache at all.
Which API dialect does Claude Opus 5 speak?
The anthropic dialect. You never write it: the gateway accepts OpenAI-shaped requests and translates them at the edge, streaming byte by byte rather than buffering. Switching to a model on a different dialect means changing the model string and nothing else.
Is Claude Opus 5 being retired?
No retirement date has been announced. Anthropic publishes deprecations ahead of time, and this page tracks that field, so a date will appear here if one is set.
Closest alternatives
By input price. Switch between any of them by changing the model string.