The catalog
11 models, one baseURL
Prices, context windows and model IDs are the provider's published figures, maintained by hand. Token counts and latency are sample data until the probes and the event stream are live — we would rather show you the shape of the page than invent a measurement.
- 11models
- 1provider
- $1cheapest / 1M in
- 1Mlargest context
- 1dialect
Filters
Filters
Category membership and token counts come from sample usage data.
Claude Opus 5 is Anthropic's flagship model for complex agentic coding and enterprise work — a step change over Opus 4.8 on deep reasoning, long-horizon autonomy and test-time compute scaling, at the same price. Thinking is on by default and the full effort ladder runs from low through max, so you tune depth per request rather than per model. It holds a 1M-token context window and writes up to 128K tokens in a single response.
Claude Sonnet 5 is the workhorse: near-Opus quality on coding and agentic work at a third of the price, with adaptive thinking on by default and the full low-to-max effort range. It is the first Sonnet-tier model with high-resolution vision and xhigh effort. Introductory pricing of $2 in / $10 out per million tokens applies through 31 August 2026.
Claude Haiku 4.5 is the cheapest and quickest model in the catalog — a dollar per million input tokens, with a 200K context window and a 64K output cap. It uses the older extended-thinking parameter rather than adaptive thinking, which makes it predictable to budget for. Built for the calls you make thousands of times an hour.
Claude Fable 5 is Anthropic's most capable widely released model, aimed at the most demanding reasoning and long-horizon agentic work. Thinking is always on — there is no way to disable it — and the raw chain of thought is never returned, only a summary. Single requests on hard tasks can run for many minutes, so plan for streaming and asynchronous check-ins rather than blocking calls.
Claude Opus 4.8 is the previous Opus flagship — highly autonomous, strong on long-horizon agentic work, knowledge work and memory, with a noticeably warmer and less hedged writing voice than 4.7. Same request surface as Opus 5 minus the two Opus 5 breaking changes, which makes it the usual fallback target when a newer model declines a request.
Claude Opus 4.7 introduced the tokenizer, the xhigh effort level and the 2576-pixel high-resolution vision path that the newer models inherit. It follows instructions more literally than 4.6 and reaches for tools less often by default. Still callable and still strong on long-horizon agentic work.
Claude Opus 4.6 is the oldest Opus still in the current line. It is the last one that accepts temperature, top_p and top_k, and the last that still honours a manual thinking budget alongside adaptive thinking — useful if you need a hard ceiling on thinking spend while you migrate. 1M context, 128K output.
Claude Sonnet 4.6 keeps a 1M-token context window and a 128K output cap at $3 in / $15 out, with adaptive thinking and the older tokenizer. Because the Sonnet 5 tokenizer counts roughly 30% more tokens for the same text, 4.6 is still the cheaper option for some long-context workloads even at identical sticker prices.
Claude Opus 4.5 is a legacy model that remains callable. It predates adaptive thinking, so it uses an explicit thinking budget, and it caps at a 200K context window with 64K of output. It was the first model to ship the effort parameter, limited to low, medium and high.
Claude Sonnet 4.5 is a legacy model that remains callable, with a 200K context window, 64K output cap and manual extended thinking. It has no effort parameter — asking for one is an error — so depth is controlled purely by the thinking budget you set.
Claude Opus 4.1 is deprecated and will stop answering on 5 August 2026. It is also the most expensive model in the catalog at $15 in / $75 out per million tokens — three times the price of Opus 5 for a 200K context window and a 32K output cap. Migrate to Claude Opus 5 before the retirement date.
| Model | Provider | Context | Max out | In / 1M | Out / 1M | TTFT | Uptime |
|---|---|---|---|---|---|---|---|
| Claude Opus 5anthropic/claude-opus-51M contextAdaptive thinkingHigh-res vision | Anthropic | 1M | 128K | $5 | $25 | 780ms | 99.94% |
| Claude Sonnet 5anthropic/claude-sonnet-51M contextAdaptive thinkingHigh-res vision | Anthropic | 1M | 128K | $3 | $15 | 430ms | 99.96% |
| Claude Haiku 4.5anthropic/claude-haiku-4-5Extended thinkingVisionStructured outputs | Anthropic | 200K | 64K | $1 | $5 | 210ms | 99.97% |
| Claude Fable 5anthropic/claude-fable-51M contextAdaptive thinkingHigh-res vision | Anthropic | 1M | 128K | $10 | $50 | 1450ms | 99.9% |
| Claude Opus 4.8anthropic/claude-opus-4-81M contextAdaptive thinkingHigh-res vision | Anthropic | 1M | 128K | $5 | $25 | 720ms | 99.95% |
| Claude Opus 4.7anthropic/claude-opus-4-71M contextAdaptive thinkingHigh-res vision | Anthropic | 1M | 128K | $5 | $25 | 700ms | 99.93% |
| Claude Opus 4.6anthropic/claude-opus-4-61M contextAdaptive thinkingExtended thinking | Anthropic | 1M | 128K | $5 | $25 | 690ms | 99.92% |
| Claude Sonnet 4.6anthropic/claude-sonnet-4-61M contextAdaptive thinkingExtended thinking | Anthropic | 1M | 128K | $3 | $15 | 410ms | 99.95% |
| Claude Opus 4.5anthropic/claude-opus-4-5Extended thinkingVisionStructured outputs | Anthropic | 200K | 64K | $5 | $25 | 660ms | 99.9% |
| Claude Sonnet 4.5anthropic/claude-sonnet-4-5Extended thinkingVision | Anthropic | 200K | 64K | $3 | $15 | 390ms | 99.93% |
| Claude Opus 4.1Retiringanthropic/claude-opus-4-1Extended thinkingVisionStructured outputs | Anthropic | 200K | 32K | $15 | $75 | 640ms | 99.85% |
Nothing matches those filters..
Sample dataToken counts total 50.27B across the catalog over four weeks, and all of it is illustrative — see the rankings for the full picture and the caveat.
Providers
Each provider is one dialect spec in the adapter engine. Adding the next one is a config file and a row, not an integration.