Claude Fable 5
anthropic/claude-fable-5
Next-generation intelligence for long-running agents.
Claude Fable 5 is Anthropic's most capable widely released model, aimed at the most demanding reasoning and long-horizon agentic work. Thinking is always on — there is no way to disable it — and the raw chain of thought is never returned, only a summary. Single requests on hard tasks can run for many minutes, so plan for streaming and asynchronous check-ins rather than blocking calls.
- ModalitiesText + Imageout: Text
- Price$10 / $50per 1M in / out
- Context1M128K max output
- Released9 Jun 2026cutoff Jan 2026
Providers
Every routable endpoint for this model, with the rates the gateway bills against. One endpoint today, so there is no routing decision to make — when there is more than one, the router picks between them on price, health and latency.
| Provider | Input / M | Output / M | Cache read / M | Latency | Throughput | Uptime |
|---|---|---|---|---|---|---|
| Anthropic · Messages APIclaude-fable-5 · global | $10 | $50 | $1 | 1450ms | 44 tok/s | 99.9% |
Effective Pricing Computed
What input tokens actually cost once prompt caching is working. This is arithmetic over Anthropic's published prices, not a measurement — cached input bills at $1 per million instead of $10, so the blended rate falls as more of your prefix is reused: list × (1 − hit) + cache read × hit.
Chart requires JavaScript. The figures are in the table below.
- Blended input rate
- List price, no caching
View as table
| Cache hit ratio | 0% | 25% | 50% | 75% | 90% |
|---|---|---|---|---|---|
| Blended input rate | $10 | $7.75 | $5.50 | $3.25 | $1.90 |
| List price, no caching | $10 | $10 | $10 | $10 | $10 |
| Cache hit ratio | Blended input / M | Saving | 1M in + 1M out |
|---|---|---|---|
| 0% | $10 | — | $60 |
| 25% | $7.75 | 22% | $57.75 |
| 50% | $5.50 | 45% | $55.50 |
| 75% | $3.25 | 68% | $53.25 |
| 90% | $1.90 | 81% | $51.90 |
Performance Sample data
Throughput, round-trip latency and time to first token. Nothing here has been measured — there are no probes running and no traffic to learn from yet, so these are placeholders showing what the hydration layer will fill in.
- 44tokens / secondoutput, median
- 1450mstime to first tokenmedian
- 1450msp50 latencyend to end
- 5200msp99 latencyend to end
Throughput across Anthropic
Uptime Sample data
Share of requests that completed without an upstream error, over a rolling 30-day window. Sample data — the gateway has not made enough requests to have an opinion.
- 99.9%success raterolling 30d
- 0.1%error raterolling 30d
Uptime across Anthropic
Apps Sample data
How many distinct applications sent traffic to this model each week. We deliberately do not list app names: the per-app breakdown arrives with label passthrough on the event stream, and inventing customer names to fill the gap would be worse than an empty section.
Activity Sample data
Token volume and request traffic over the trailing four weeks, rolled up from the event stream. Sample data until the stream has something real in it.
Chart requires JavaScript. The figures are in the table below.
View as table
| Metric | 26 Apr | 27 Apr | 28 Apr | 29 Apr | 30 Apr | 1 May | 2 May | 3 May | 4 May | 5 May | 6 May | 7 May | 8 May | 9 May | 10 May | 11 May | 12 May | 13 May | 14 May | 15 May | 16 May | 17 May | 18 May | 19 May | 20 May | 21 May | 22 May | 23 May | 24 May | 25 May | 26 May | 27 May | 28 May | 29 May | 30 May | 31 May | 1 Jun | 2 Jun | 3 Jun | 4 Jun | 5 Jun | 6 Jun | 7 Jun | 8 Jun | 9 Jun | 10 Jun | 11 Jun | 12 Jun | 13 Jun | 14 Jun | 15 Jun | 16 Jun | 17 Jun | 18 Jun | 19 Jun | 20 Jun | 21 Jun | 22 Jun | 23 Jun | 24 Jun | 25 Jun | 26 Jun | 27 Jun | 28 Jun | 29 Jun | 30 Jun | 1 Jul | 2 Jul | 3 Jul | 4 Jul | 5 Jul | 6 Jul | 7 Jul | 8 Jul | 9 Jul | 10 Jul | 11 Jul | 12 Jul | 13 Jul | 14 Jul | 15 Jul | 16 Jul | 17 Jul | 18 Jul | 19 Jul | 20 Jul | 21 Jul | 22 Jul | 23 Jul | 24 Jul |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Tokens | 28M | 35M | 33M | 36M | 41M | 45M | 33M | 30M | 37M | 36M | 38M | 44M | 48M | 35M | 32M | 39M | 38M | 40M | 46M | 51M | 37M | 34M | 42M | 40M | 43M | 49M | 53M | 39M | 35M | 44M | 42M | 45M | 51M | 56M | 40M | 37M | 46M | 44M | 48M | 54M | 59M | 42M | 39M | 48M | 46M | 50M | 57M | 62M | 44M | 40M | 50M | 48M | 52M | 60M | 65M | 46M | 42M | 52M | 50M | 55M | 62M | 68M | 48M | 44M | 54M | 53M | 57M | 65M | 70M | 50M | 46M | 57M | 55M | 60M | 68M | 73M | 52M | 47M | 59M | 57M | 62M | 70M | 76M | 54M | 49M | 61M | 59M | 64M | 73M | 30M |
| Week | Tokens | Requests | Apps |
|---|---|---|---|
| 20 Jul 2026 | 486.0M | 210.0K | 145 |
| 13 Jul 2026 | 402.0M | 176.0K | 128 |
| 6 Jul 2026 | 331.0M | 142.0K | 112 |
| 29 Jun 2026 | 262.0M | 112.0K | 94 |
This model ranks #6 across the catalog this week — see the full rankings.
Call it
Standard OpenAI client, one changed line. Streaming is the default path — non-streaming is just collecting the stream.
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.aigate.dev/v1",
apiKey: process.env.AIGATE_API_KEY,
});
const stream = await client.chat.completions.create({
model: "anthropic/claude-fable-5",
messages: [{ role: "user", content: "Hello" }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}Capabilities
- Streaming
- Tool use
- Vision
- High-res vision
- PDF input
- Citations
- Structured outputs
- Adaptive thinking
- Effort control
- Task budgets
- Prompt caching
- Compaction
- Batch API
- 1M context
Accepted parameters
- max_tokens
- stream
- stop_sequences
- system
- tools
- tool_choice
- thinking
- effort
- output_config.format
Specification
- Model ID
- anthropic/claude-fable-5
- Upstream ID
- claude-fable-5
- Context window
- 1,000,000 tokens
- Max output
- 128,000 tokens
- API dialect
- anthropic — translated for you
- Knowledge cutoff
- Jan 2026
- Min cacheable prefix
- 512 tokens
- Also available on
- claude-api, bedrock, google-cloud
FAQ
What does Claude Fable 5 cost?
$10 per million input tokens and $50 per million output tokens — Anthropic's published list price. AI Gate adds a flat 5% routing fee on top, billed from prepaid credits. Cached input is far cheaper at $1 per million tokens.
What is the context window for Claude Fable 5?
1,000,000 tokens of context, with up to 128,000 tokens in a single response. Its knowledge is most reliable through Jan 2026.
How do I call Claude Fable 5?
Point any OpenAI-compatible client at https://api.aigate.dev/v1 and pass "anthropic/claude-fable-5" as the model string. There is no SDK to install — we translate the request into Anthropic's anthropic dialect on the way out and the response stream back on the way in.
How is prompt caching billed?
Reading from cache costs $1 per million tokens, one tenth of the list input price. Writing to a five-minute cache costs $12.50 (1.25× list), and a one-hour cache costs $20 (2× list). Prefixes shorter than 512 tokens will not cache at all.
Which API dialect does Claude Fable 5 speak?
The anthropic dialect. You never write it: the gateway accepts OpenAI-shaped requests and translates them at the edge, streaming byte by byte rather than buffering. Switching to a model on a different dialect means changing the model string and nothing else.
Is Claude Fable 5 being retired?
No retirement date has been announced. Anthropic publishes deprecations ahead of time, and this page tracks that field, so a date will appear here if one is set.
Closest alternatives
By input price. Switch between any of them by changing the model string.