Claude Sonnet 4.5

anthropic/claude-sonnet-4-5

Legacy Sonnet. 200K context, manual thinking budgets.

Claude Sonnet 4.5 is a legacy model that remains callable, with a 200K context window, 64K output cap and manual extended thinking. It has no effort parameter — asking for one is an error — so depth is controlled purely by the thinking budget you set.

  • ModalitiesText + Imageout: Text
  • Price$3 / $15per 1M in / out
  • Context200K64K max output
  • Released29 Sept 2025cutoff Jan 2025

Providers

Every routable endpoint for this model, with the rates the gateway bills against. One endpoint today, so there is no routing decision to make — when there is more than one, the router picks between them on price, health and latency.

ProviderInput / MOutput / MCache read / MLatencyThroughputUptime
Anthropic · Messages APIclaude-sonnet-4-5 · global$3$15$0.30390ms112 tok/s99.93%
Latency, throughput and uptime are sample data. Prices are Anthropic's published figures.

Effective Pricing Computed

What input tokens actually cost once prompt caching is working. This is arithmetic over Anthropic's published prices, not a measurement — cached input bills at $0.30 per million instead of $3, so the blended rate falls as more of your prefix is reused: list × (1 − hit) + cache read × hit.

Chart requires JavaScript. The figures are in the table below.

  • Blended input rate
  • List price, no caching
Computed from list prices — not a measurement. Output tokens are unaffected by caching and always bill at list price.
View as table
Blended input price for Claude Sonnet 4.5 by cache hit ratio
Cache hit ratio0%25%50%75%90%
Blended input rate$3$2.33$1.65$0.97$0.57
List price, no caching$3$3$3$3$3
Cache hit ratioBlended input / MSaving1M in + 1M out
0%$3$18
25%$2.3322%$17.32
50%$1.6545%$16.65
75%$0.9768%$15.97
90%$0.5781%$15.57
Writing to the cache costs more than a plain read the first time: $3.75 per million for a five-minute prefix and $6 for a one-hour one. That is a one-off per prefix, so it is excluded from the steady-state rates above. Prefixes under 1,024 tokens never cache.

Performance Sample data

Throughput, round-trip latency and time to first token. Nothing here has been measured — there are no probes running and no traffic to learn from yet, so these are placeholders showing what the hydration layer will fill in.

  • 112tokens / secondoutput, median
  • 390mstime to first tokenmedian
  • 390msp50 latencyend to end
  • 1240msp99 latencyend to end

Throughput across Anthropic

Sample data — illustrative ordering, not a benchmark.

Uptime Sample data

Share of requests that completed without an upstream error, over a rolling 30-day window. Sample data — the gateway has not made enough requests to have an opinion.

  • 99.93%success raterolling 30d
  • 0.07%error raterolling 30d

Uptime across Anthropic

Sample data. Bars start at 99.45% so the differences are visible — they are all within a fraction of a percent.

Apps Sample data

How many distinct applications sent traffic to this model each week. We deliberately do not list app names: the per-app breakdown arrives with label passthrough on the event stream, and inventing customer names to fill the gap would be worse than an empty section.

  • 29 Jun 2026104 apps
  • 6 Jul 202699 apps
  • 13 Jul 202693 apps
  • 20 Jul 202688 apps
Sample data.

Activity Sample data

Token volume and request traffic over the trailing four weeks, rolled up from the event stream. Sample data until the stream has something real in it.

Chart requires JavaScript. The figures are in the table below.

Sample data. The final bar is hatched because that day is still in progress.
View as table
Daily tokens processed by Claude Sonnet 4.5
Metric26 Apr27 Apr28 Apr29 Apr30 Apr1 May2 May3 May4 May5 May6 May7 May8 May9 May10 May11 May12 May13 May14 May15 May16 May17 May18 May19 May20 May21 May22 May23 May24 May25 May26 May27 May28 May29 May30 May31 May1 Jun2 Jun3 Jun4 Jun5 Jun6 Jun7 Jun8 Jun9 Jun10 Jun11 Jun12 Jun13 Jun14 Jun15 Jun16 Jun17 Jun18 Jun19 Jun20 Jun21 Jun22 Jun23 Jun24 Jun25 Jun26 Jun27 Jun28 Jun29 Jun30 Jun1 Jul2 Jul3 Jul4 Jul5 Jul6 Jul7 Jul8 Jul9 Jul10 Jul11 Jul12 Jul13 Jul14 Jul15 Jul16 Jul17 Jul18 Jul19 Jul20 Jul21 Jul22 Jul23 Jul24 Jul
Tokens6.0M8.3M9.2M11M11M11M7.1M6.4M8.9M9.8M11M12M12M7.5M6.8M9.4M10M12M13M12M7.9M7.2M9.9M11M13M13M13M8.3M7.5M10M12M13M14M14M8.7M7.9M11M12M14M15M14M9.2M8.3M12M13M15M15M15M9.6M8.7M12M14M15M16M15M10.0M9.0M13M14M16M17M16M10M9.4M13M15M17M18M17M11M9.8M14M15M17M18M17M11M10M14M16M18M19M18M12M11M15M17M19M20M7.1M
WeekTokensRequestsApps
20 Jul 2026121.0M190.0K88
13 Jul 2026134.0M205.0K93
6 Jul 2026149.0M224.0K99
29 Jun 2026161.0M242.0K104

Call it

Standard OpenAI client, one changed line. Streaming is the default path — non-streaming is just collecting the stream.

example.ts
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.aigate.dev/v1",
  apiKey: process.env.AIGATE_API_KEY,
});

const stream = await client.chat.completions.create({
  model: "anthropic/claude-sonnet-4-5",
  messages: [{ role: "user", content: "Hello" }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}

Capabilities

  • Streaming
  • Tool use
  • Vision
  • PDF input
  • Citations
  • Extended thinking
  • Prompt caching
  • Batch API

Accepted parameters

  • max_tokens
  • stream
  • stop_sequences
  • system
  • tools
  • tool_choice
  • thinking
  • temperature
  • top_p
  • top_k

Specification

Model ID
anthropic/claude-sonnet-4-5
Upstream ID
claude-sonnet-4-5
Context window
200,000 tokens
Max output
64,000 tokens
API dialect
anthropic — translated for you
Knowledge cutoff
Jan 2025
Min cacheable prefix
1,024 tokens
Also available on
claude-api, bedrock, google-cloud

FAQ

What does Claude Sonnet 4.5 cost?

$3 per million input tokens and $15 per million output tokens — Anthropic's published list price. AI Gate adds a flat 5% routing fee on top, billed from prepaid credits. Cached input is far cheaper at $0.30 per million tokens.

What is the context window for Claude Sonnet 4.5?

200,000 tokens of context, with up to 64,000 tokens in a single response. Its knowledge is most reliable through Jan 2025.

How do I call Claude Sonnet 4.5?

Point any OpenAI-compatible client at https://api.aigate.dev/v1 and pass "anthropic/claude-sonnet-4-5" as the model string. There is no SDK to install — we translate the request into Anthropic's anthropic dialect on the way out and the response stream back on the way in.

How is prompt caching billed?

Reading from cache costs $0.30 per million tokens, one tenth of the list input price. Writing to a five-minute cache costs $3.75 (1.25× list), and a one-hour cache costs $6 (2× list). Prefixes shorter than 1,024 tokens will not cache at all.

Which API dialect does Claude Sonnet 4.5 speak?

The anthropic dialect. You never write it: the gateway accepts OpenAI-shaped requests and translates them at the edge, streaming byte by byte rather than buffering. Switching to a model on a different dialect means changing the model string and nothing else.

Is Claude Sonnet 4.5 being retired?

No retirement date has been announced. Anthropic publishes deprecations ahead of time, and this page tracks that field, so a date will appear here if one is set.

Closest alternatives

By input price. Switch between any of them by changing the model string.