Claude Sonnet 4.6

anthropic/claude-sonnet-4-6

The previous-generation Sonnet, still on 1M context.

Claude Sonnet 4.6 keeps a 1M-token context window and a 128K output cap at $3 in / $15 out, with adaptive thinking and the older tokenizer. Because the Sonnet 5 tokenizer counts roughly 30% more tokens for the same text, 4.6 is still the cheaper option for some long-context workloads even at identical sticker prices.

  • ModalitiesText + Imageout: Text
  • Price$3 / $15per 1M in / out
  • Context1M128K max output
  • ReleasedNot publishedcutoff Aug 2025

Providers

Every routable endpoint for this model, with the rates the gateway bills against. One endpoint today, so there is no routing decision to make — when there is more than one, the router picks between them on price, health and latency.

ProviderInput / MOutput / MCache read / MLatencyThroughputUptime
Anthropic · Messages APIclaude-sonnet-4-6 · global$3$15$0.30410ms108 tok/s99.95%
Latency, throughput and uptime are sample data. Prices are Anthropic's published figures.

Effective Pricing Computed

What input tokens actually cost once prompt caching is working. This is arithmetic over Anthropic's published prices, not a measurement — cached input bills at $0.30 per million instead of $3, so the blended rate falls as more of your prefix is reused: list × (1 − hit) + cache read × hit.

Chart requires JavaScript. The figures are in the table below.

  • Blended input rate
  • List price, no caching
Computed from list prices — not a measurement. Output tokens are unaffected by caching and always bill at list price.
View as table
Blended input price for Claude Sonnet 4.6 by cache hit ratio
Cache hit ratio0%25%50%75%90%
Blended input rate$3$2.33$1.65$0.97$0.57
List price, no caching$3$3$3$3$3
Cache hit ratioBlended input / MSaving1M in + 1M out
0%$3$18
25%$2.3322%$17.32
50%$1.6545%$16.65
75%$0.9768%$15.97
90%$0.5781%$15.57
Writing to the cache costs more than a plain read the first time: $3.75 per million for a five-minute prefix and $6 for a one-hour one. That is a one-off per prefix, so it is excluded from the steady-state rates above. Prefixes under 1,024 tokens never cache.

Performance Sample data

Throughput, round-trip latency and time to first token. Nothing here has been measured — there are no probes running and no traffic to learn from yet, so these are placeholders showing what the hydration layer will fill in.

  • 108tokens / secondoutput, median
  • 410mstime to first tokenmedian
  • 410msp50 latencyend to end
  • 1290msp99 latencyend to end

Throughput across Anthropic

Sample data — illustrative ordering, not a benchmark.

Uptime Sample data

Share of requests that completed without an upstream error, over a rolling 30-day window. Sample data — the gateway has not made enough requests to have an opinion.

  • 99.95%success raterolling 30d
  • 0.05%error raterolling 30d

Uptime across Anthropic

Sample data. Bars start at 99.45% so the differences are visible — they are all within a fraction of a percent.

Apps Sample data

How many distinct applications sent traffic to this model each week. We deliberately do not list app names: the per-app breakdown arrives with label passthrough on the event stream, and inventing customer names to fill the gap would be worse than an empty section.

  • 29 Jun 2026461 apps
  • 6 Jul 2026446 apps
  • 13 Jul 2026428 apps
  • 20 Jul 2026410 apps
Sample data.

Activity Sample data

Token volume and request traffic over the trailing four weeks, rolled up from the event stream. Sample data until the stream has something real in it.

Chart requires JavaScript. The figures are in the table below.

Sample data. The final bar is hatched because that day is still in progress.
View as table
Daily tokens processed by Claude Sonnet 4.6
Metric26 Apr27 Apr28 Apr29 Apr30 Apr1 May2 May3 May4 May5 May6 May7 May8 May9 May10 May11 May12 May13 May14 May15 May16 May17 May18 May19 May20 May21 May22 May23 May24 May25 May26 May27 May28 May29 May30 May31 May1 Jun2 Jun3 Jun4 Jun5 Jun6 Jun7 Jun8 Jun9 Jun10 Jun11 Jun12 Jun13 Jun14 Jun15 Jun16 Jun17 Jun18 Jun19 Jun20 Jun21 Jun22 Jun23 Jun24 Jun25 Jun26 Jun27 Jun28 Jun29 Jun30 Jun1 Jul2 Jul3 Jul4 Jul5 Jul6 Jul7 Jul8 Jul9 Jul10 Jul11 Jul12 Jul13 Jul14 Jul15 Jul16 Jul17 Jul18 Jul19 Jul20 Jul21 Jul22 Jul23 Jul24 Jul
Tokens63M99M107M106M96M86M61M67M106M114M112M102M91M65M71M112M121M119M107M97M68M75M119M128M125M113M102M72M80M126M135M131M119M107M76M84M132M141M138M124M112M80M88M139M148M144M130M117M84M92M146M155M150M135M122M87M97M152M162M157M141M127M91M101M159M169M163M146M132M95M105M166M176M169M152M137M99M110M172M182M176M157M142M103M114M179M189M182M163M56M
WeekTokensRequestsApps
20 Jul 20261.18B1.9M410
13 Jul 20261.31B2.1M428
6 Jul 20261.44B2.3M446
29 Jun 20261.52B2.4M461

Call it

Standard OpenAI client, one changed line. Streaming is the default path — non-streaming is just collecting the stream.

example.ts
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.aigate.dev/v1",
  apiKey: process.env.AIGATE_API_KEY,
});

const stream = await client.chat.completions.create({
  model: "anthropic/claude-sonnet-4-6",
  messages: [{ role: "user", content: "Hello" }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}

Capabilities

  • Streaming
  • Tool use
  • Vision
  • PDF input
  • Citations
  • Adaptive thinking
  • Extended thinking
  • Effort control
  • Prompt caching
  • Compaction
  • Batch API
  • 1M context

Accepted parameters

  • max_tokens
  • stream
  • stop_sequences
  • system
  • tools
  • tool_choice
  • thinking
  • effort
  • temperature
  • top_p
  • top_k

Specification

Model ID
anthropic/claude-sonnet-4-6
Upstream ID
claude-sonnet-4-6
Context window
1,000,000 tokens
Max output
128,000 tokens
API dialect
anthropic — translated for you
Knowledge cutoff
Aug 2025
Min cacheable prefix
1,024 tokens
Also available on
claude-api, bedrock, google-cloud

FAQ

What does Claude Sonnet 4.6 cost?

$3 per million input tokens and $15 per million output tokens — Anthropic's published list price. AI Gate adds a flat 5% routing fee on top, billed from prepaid credits. Cached input is far cheaper at $0.30 per million tokens.

What is the context window for Claude Sonnet 4.6?

1,000,000 tokens of context, with up to 128,000 tokens in a single response. Its knowledge is most reliable through Aug 2025.

How do I call Claude Sonnet 4.6?

Point any OpenAI-compatible client at https://api.aigate.dev/v1 and pass "anthropic/claude-sonnet-4-6" as the model string. There is no SDK to install — we translate the request into Anthropic's anthropic dialect on the way out and the response stream back on the way in.

How is prompt caching billed?

Reading from cache costs $0.30 per million tokens, one tenth of the list input price. Writing to a five-minute cache costs $3.75 (1.25× list), and a one-hour cache costs $6 (2× list). Prefixes shorter than 1,024 tokens will not cache at all.

Which API dialect does Claude Sonnet 4.6 speak?

The anthropic dialect. You never write it: the gateway accepts OpenAI-shaped requests and translates them at the edge, streaming byte by byte rather than buffering. Switching to a model on a different dialect means changing the model string and nothing else.

Is Claude Sonnet 4.6 being retired?

No retirement date has been announced. Anthropic publishes deprecations ahead of time, and this page tracks that field, so a date will appear here if one is set.

Closest alternatives

By input price. Switch between any of them by changing the model string.