Claude Sonnet 4.5
anthropic/claude-sonnet-4-5
Legacy Sonnet. 200K context, manual thinking budgets.
Claude Sonnet 4.5 is a legacy model that remains callable, with a 200K context window, 64K output cap and manual extended thinking. It has no effort parameter — asking for one is an error — so depth is controlled purely by the thinking budget you set.
- ModalitiesText + Imageout: Text
- Price$3 / $15per 1M in / out
- Context200K64K max output
- Released29 Sept 2025cutoff Jan 2025
Providers
Every routable endpoint for this model, with the rates the gateway bills against. One endpoint today, so there is no routing decision to make — when there is more than one, the router picks between them on price, health and latency.
| Provider | Input / M | Output / M | Cache read / M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|
| Anthropic · Messages APIclaude-sonnet-4-5 · global | $3 | $15 | $0.30 | 390ms | 112 tok/s | 99.93% |
Effective Pricing Computed
What input tokens actually cost once prompt caching is working. This is arithmetic over Anthropic's published prices, not a measurement — cached input bills at $0.30 per million instead of $3, so the blended rate falls as more of your prefix is reused: list × (1 − hit) + cache read × hit.
Chart requires JavaScript. The figures are in the table below.
- Blended input rate
- List price, no caching
View as table
| Cache hit ratio | 0% | 25% | 50% | 75% | 90% |
|---|---|---|---|---|---|
| Blended input rate | $3 | $2.33 | $1.65 | $0.97 | $0.57 |
| List price, no caching | $3 | $3 | $3 | $3 | $3 |
| Cache hit ratio | Blended input / M | Saving | 1M in + 1M out |
|---|---|---|---|
| 0% | $3 | — | $18 |
| 25% | $2.33 | 22% | $17.32 |
| 50% | $1.65 | 45% | $16.65 |
| 75% | $0.97 | 68% | $15.97 |
| 90% | $0.57 | 81% | $15.57 |
Performance Sample data
Throughput, round-trip latency and time to first token. Nothing here has been measured — there are no probes running and no traffic to learn from yet, so these are placeholders showing what the hydration layer will fill in.
- 112tokens / secondoutput, median
- 390mstime to first tokenmedian
- 390msp50 latencyend to end
- 1240msp99 latencyend to end
Throughput across Anthropic
Uptime Sample data
Share of requests that completed without an upstream error, over a rolling 30-day window. Sample data — the gateway has not made enough requests to have an opinion.
- 99.93%success raterolling 30d
- 0.07%error raterolling 30d
Uptime across Anthropic
Apps Sample data
How many distinct applications sent traffic to this model each week. We deliberately do not list app names: the per-app breakdown arrives with label passthrough on the event stream, and inventing customer names to fill the gap would be worse than an empty section.
Activity Sample data
Token volume and request traffic over the trailing four weeks, rolled up from the event stream. Sample data until the stream has something real in it.
Chart requires JavaScript. The figures are in the table below.
View as table
| Metric | 26 Apr | 27 Apr | 28 Apr | 29 Apr | 30 Apr | 1 May | 2 May | 3 May | 4 May | 5 May | 6 May | 7 May | 8 May | 9 May | 10 May | 11 May | 12 May | 13 May | 14 May | 15 May | 16 May | 17 May | 18 May | 19 May | 20 May | 21 May | 22 May | 23 May | 24 May | 25 May | 26 May | 27 May | 28 May | 29 May | 30 May | 31 May | 1 Jun | 2 Jun | 3 Jun | 4 Jun | 5 Jun | 6 Jun | 7 Jun | 8 Jun | 9 Jun | 10 Jun | 11 Jun | 12 Jun | 13 Jun | 14 Jun | 15 Jun | 16 Jun | 17 Jun | 18 Jun | 19 Jun | 20 Jun | 21 Jun | 22 Jun | 23 Jun | 24 Jun | 25 Jun | 26 Jun | 27 Jun | 28 Jun | 29 Jun | 30 Jun | 1 Jul | 2 Jul | 3 Jul | 4 Jul | 5 Jul | 6 Jul | 7 Jul | 8 Jul | 9 Jul | 10 Jul | 11 Jul | 12 Jul | 13 Jul | 14 Jul | 15 Jul | 16 Jul | 17 Jul | 18 Jul | 19 Jul | 20 Jul | 21 Jul | 22 Jul | 23 Jul | 24 Jul |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Tokens | 6.0M | 8.3M | 9.2M | 11M | 11M | 11M | 7.1M | 6.4M | 8.9M | 9.8M | 11M | 12M | 12M | 7.5M | 6.8M | 9.4M | 10M | 12M | 13M | 12M | 7.9M | 7.2M | 9.9M | 11M | 13M | 13M | 13M | 8.3M | 7.5M | 10M | 12M | 13M | 14M | 14M | 8.7M | 7.9M | 11M | 12M | 14M | 15M | 14M | 9.2M | 8.3M | 12M | 13M | 15M | 15M | 15M | 9.6M | 8.7M | 12M | 14M | 15M | 16M | 15M | 10.0M | 9.0M | 13M | 14M | 16M | 17M | 16M | 10M | 9.4M | 13M | 15M | 17M | 18M | 17M | 11M | 9.8M | 14M | 15M | 17M | 18M | 17M | 11M | 10M | 14M | 16M | 18M | 19M | 18M | 12M | 11M | 15M | 17M | 19M | 20M | 7.1M |
| Week | Tokens | Requests | Apps |
|---|---|---|---|
| 20 Jul 2026 | 121.0M | 190.0K | 88 |
| 13 Jul 2026 | 134.0M | 205.0K | 93 |
| 6 Jul 2026 | 149.0M | 224.0K | 99 |
| 29 Jun 2026 | 161.0M | 242.0K | 104 |
This model ranks #9 across the catalog this week — see the full rankings.
Call it
Standard OpenAI client, one changed line. Streaming is the default path — non-streaming is just collecting the stream.
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.aigate.dev/v1",
apiKey: process.env.AIGATE_API_KEY,
});
const stream = await client.chat.completions.create({
model: "anthropic/claude-sonnet-4-5",
messages: [{ role: "user", content: "Hello" }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}Capabilities
- Streaming
- Tool use
- Vision
- PDF input
- Citations
- Extended thinking
- Prompt caching
- Batch API
Accepted parameters
- max_tokens
- stream
- stop_sequences
- system
- tools
- tool_choice
- thinking
- temperature
- top_p
- top_k
Specification
- Model ID
- anthropic/claude-sonnet-4-5
- Upstream ID
- claude-sonnet-4-5
- Context window
- 200,000 tokens
- Max output
- 64,000 tokens
- API dialect
- anthropic — translated for you
- Knowledge cutoff
- Jan 2025
- Min cacheable prefix
- 1,024 tokens
- Also available on
- claude-api, bedrock, google-cloud
FAQ
What does Claude Sonnet 4.5 cost?
$3 per million input tokens and $15 per million output tokens — Anthropic's published list price. AI Gate adds a flat 5% routing fee on top, billed from prepaid credits. Cached input is far cheaper at $0.30 per million tokens.
What is the context window for Claude Sonnet 4.5?
200,000 tokens of context, with up to 64,000 tokens in a single response. Its knowledge is most reliable through Jan 2025.
How do I call Claude Sonnet 4.5?
Point any OpenAI-compatible client at https://api.aigate.dev/v1 and pass "anthropic/claude-sonnet-4-5" as the model string. There is no SDK to install — we translate the request into Anthropic's anthropic dialect on the way out and the response stream back on the way in.
How is prompt caching billed?
Reading from cache costs $0.30 per million tokens, one tenth of the list input price. Writing to a five-minute cache costs $3.75 (1.25× list), and a one-hour cache costs $6 (2× list). Prefixes shorter than 1,024 tokens will not cache at all.
Which API dialect does Claude Sonnet 4.5 speak?
The anthropic dialect. You never write it: the gateway accepts OpenAI-shaped requests and translates them at the edge, streaming byte by byte rather than buffering. Switching to a model on a different dialect means changing the model string and nothing else.
Is Claude Sonnet 4.5 being retired?
No retirement date has been announced. Anthropic publishes deprecations ahead of time, and this page tracks that field, so a date will appear here if one is set.
Closest alternatives
By input price. Switch between any of them by changing the model string.