Key Takeaways
- An upgrade at the same price: $10 input / $50 output per 1M tokens, identical to Fable 5, with capabilities a clear step up
- Cache reads cost a quarter: cache reads drop from $1.00 to $0.25 / 1M tokens (0.025x base input, versus 0.1x on other Claude models), which lands directly on long-running agent sessions. APIYI has already matched the cut
- Benchmark jumps, not nudges: Terminal-Bench-Science goes from 24.7% to 52.6%, AutomationBench from 17.1% to 31.4% — both more than doubling
- ⚠️ Three breaking changes: forced tool use (
tool_choiceofany/tool) now returns 400, thinking blocks are bound to the model that produced them, and editing earlier turns invalidates thinking blocks — read before migrating from Fable 5 - Live on APIYI: both
claude-fable-5-1andclaude-fable-5-1-thinking, on the OpenAI-compatible and native Anthropic endpoints, in thedefault/svip/ClaudeCodegroups
Background
On September 1, 2026, Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 — the first iteration of the Mythos-class line since Fable 5 launched on June 9. Anthropic positions both as the most advanced models for coding and knowledge work. Unlike the usual generational price bump, Fable 5.1 holds input and output pricing exactly level with Fable 5. The only price change is a 75% cut to cache reads. Anthropic’s own framing: at low and medium effort, Fable 5.1 reaches similar or better results than Fable 5 at a much lower cost — roughly 25% lower overall for typical workloads, and up to 45% lower on agentic tasks. Fable 5.1 and Mythos 5.1 are the same model with different levels of safeguards. Mythos 5.1 is offered only to approved Project Glasswing customers; Fable 5.1 is the publicly available one. APIYI hasclaude-fable-5-1 and claude-fable-5-1-thinking live, with all four line items — input, output, cache reads, and cache writes — priced exactly in line with the provider, cache reads included at $0.25.
What’s New
Capability gains
Long-horizon agentic coding
Multi-file features, large refactors and migrations, debugging and code review across sessions that run for hours
Documents, spreadsheets, slides
From a first question to a finished document, a live-formula spreadsheet, or a deck built from a blank page
Research and search
Higher accuracy on multistep web research and deep-research tasks that follow up on what they find
Vision
Reads dense charts, filings and tables nested in PDFs, including crop-and-zoom on charts
Long context
Reasons over and connects details across the full 1M token context window
Computer use
Operates browsers and desktop apps more reliably and recovers from failed steps
Benchmark highlights
Official figures (Fable 5.1 vs Fable 5 vs Opus 5):
Terminal-Bench-Science and AutomationBench are the two that more than double — and they map onto exactly the workloads people run in production: long-horizon research agents and business process automation. Multilingual performance is on par with Fable 5.
Specifications
⚠️ Three Breaking Changes (read before migrating)
1. Forced tool use is not supported
tool_choice set to {"type": "any"} or {"type": "tool", "name": "..."} returns a 400 invalid_request_error:
{"type": "auto"} (the default) and {"type": "none"} are unchanged. The same validation applies to the token counting endpoint.
Thinking is always on for this model, and a forced tool call would skip it — the model would write its working-out into the tool arguments instead, which lowers argument quality.
What to do instead: keep tool_choice: {"type": "auto"} and use strict tool use (strict: true) or structured outputs for schema-valid JSON. To make the model call a tool rather than reply in text, state in the prompt when the tool applies (for example, “Use the get_weather tool to answer”) — Fable 5.1 follows explicit tool instructions reliably.
2. Thinking blocks are tied to the model that produced them
Every thinking block records which model produced it, and it is preserved in one direction only: Fable 5.1 reads earlier models’ thinking blocks, and no earlier model reads Fable 5.1’s. So a conversation that moves onto Fable 5.1 from Opus 5, Fable 5, or any earlier Claude model keeps its reasoning; a conversation that moves from Fable 5.1 back to those models loses it for the turns that run there. When a request carries a block the target model can’t read, the API drops it before the model sees it — dropped blocks don’t count towardinput_tokens and aren’t billed.
Worth a close look if your gateway or client does model routing or failure fallback that switches models mid-conversation. The
thinking-binding-controls-2026-08-01 beta header surfaces dropped blocks in a top-level input_transformations array; without it, the drop is silent.3. Editing earlier turns invalidates thinking blocks
Modifying anything before a Fable 5.1 thinking block — thesystem prompt, the tools array, or an earlier message — errors on the next request with The block is bound to a different conversation.
Patterns that invalidate every later thinking block:
- Editing, reordering, or removing an earlier turn while keeping later ones
- Injecting per-request text into an earlier turn (a reminder or status line) that you remove on the next request
- Rebuilding the top-level
systemprompt ortoolsarray between requests in the same conversation - An image or document URL that serves different bytes on a later request (the check covers the bytes, not the URL, so a rotating signed URL for the same file is fine)
cache_control markers, and changing effort between requests.
The check is enforced for accounts created on or after August 31, 2026. Earlier accounts have the mismatch recorded but not acted on, unless the request sets
thinking.block_binding.prefix_mismatch_behavior.Practical rule: treat the conversation as append-only. Add instructions with a turn-scoped system message (clear_at: "next_user_message", beta) and change tools with mid-conversation tool changes rather than editing system or tools. These patterns also keep the prompt cache warm.New Features (all beta)
Change effort mid-conversation
Raise it for a hard step, lower it for routine ones, without invalidating the prompt cache. Beta header:
mid-conversation-output-config-2026-07-01Turn-scoped system messages
clear_at: "next_user_message" applies for the current turn only; the message stays in messages and is sent back verbatim, so nothing earlier changes. Beta header: mid-conversation-system-clear-at-2026-08-21Progress updates between tool calls
Set
thinking.display to "updates" to receive progress updates as text while reasoning stays hidden. Beta header: thinking-display-updates-2026-08-18Behavior Differences from Fable 5 (no code change required to notice)
None of these error out, but they change how the model feels. Each has a prompting fix:Unchanged from Fable 5
- Adaptive thinking is always on:
thinking: {"type": "enabled"}withbudget_tokensand{"type": "disabled"}both return 400. Omitthinkingor send{"type": "adaptive"} thinking.displaydefaults to"omitted";"summarized"is available, and the raw chain of thought is never returned- Prefilling the assistant response returns 400
- Non-default
temperature,top_p, ortop_kvalues return 400 - The minimum cacheable prompt length is still 512 tokens
- Interleaved thinking is automatic, with no beta header
Refusals, Fallback, and Billing
Fable 5.1 carries safety classifiers covering the samestop_details categories as Fable 5:
- Refusals: a declined request returns HTTP 200 with
stop_reason: "refusal"and astop_detailsobject naming the policy area that fired - Fallback: retry a refused request on another model with server-side fallback. The permitted fallback targets for Fable 5.1 are Claude Opus 4.8 and Claude Opus 5
- Billing: you aren’t billed for a refusal that arrives before any output, and fallback credit refunds the prompt-cache cost of switching models
Data Retention and Compliance (important)
The retention happens at Anthropic and the cloud providers, not at APIYI: the 30-day retention above takes place provider-side (passed through to Anthropic for abuse detection). APIYI retains no data of its own and acts as a transparent proxy, forwarding requests only. Only Mythos-class models (Fable 5 / 5.1, Mythos 5 / 5.1) are subject to this requirement — other Claude models such as Opus 5, Opus 4.8, and Sonnet 5 are unaffected.
Putting It to Work
Where it fits
- Long-horizon agentic coding: large cross-session refactors, migrations, multi-file feature work
- Multistep research and deep retrieval: tasks that follow up on intermediate findings
- Document-heavy knowledge work: filings, papers, chart-dense PDFs, through to a finished deliverable
- Very long context analysis: connecting details across a 1M token window
Anthropic’s own guidance: start with Claude Opus 5 for most workloads. Reach for Fable 5.1 for demanding reasoning and long-horizon agentic work, or when your evals on Opus 5 at higher effort still fall short.
Code examples
Native Anthropic format
OpenAI-compatible format
Five-step migration checklist
- Change the model ID:
claude-fable-5→claude-fable-5-1 - Remove forced tool use: drop
tool_choiceof typeany/tool, and move schema enforcement to strict tool use or structured outputs - Keep history append-only: pass thinking blocks back unchanged, and don’t rebuild
system/toolsor edit earlier messages between requests - Re-tune effort: the default is
high; consider changing it mid-conversation instead of holding one level for the whole session - Re-run your evals: refusal handling, fallback, and token counts carry over, but cache reads cost less and the default behaviors above differ
Pricing and Availability
Pricing
Official pricing (USD per 1M tokens):
That table is also APIYI’s pricing — every line item matches the provider, with no markup.
Cache reads (hits and refreshes) cost 0.025x the base input price on these models, versus 0.1x on other Claude models — the saving lands hardest on long agentic sessions that re-read the same cached prefix. Batch processing is $5 input / $25 output per 1M tokens.
APIYI prices this model in line with the provider and adds no markup. On top of that, some groups carry a discount and stack with recharge bonus campaigns — that part is our own margin given back, separate from the model’s pricing.For what you were actually charged, treat the console’s cache billing detail as authoritative — the
usage cache fields echoed by the API are not a billing record.Groups and endpoints
The
ClaudeCode group is for clients speaking the native Anthropic protocol, such as Claude Code. Match the group to the protocol: use ClaudeCode for the native Anthropic protocol, and default / svip for the OpenAI-compatible protocol. That group also carries a discount, which stacks with recharge bonuses.Stacking with recharge promotions
Pair this with APIYI’s recharge bonus campaigns to bring the effective cost down further:docs.apiyi.com/en/faq/recharge-promotions.
Verdict
Claude Fable 5.1 is an unusual release: the price of input and output doesn’t move at all, cache reads fall to a quarter, and the agentic benchmarks that matter most in production more than double. For teams already running long-horizon agents, the migration pays for itself quickly. On APIYI every line item matches the provider, cache reads included at the new $0.25. Confirm three things before you migrate:- Whether your code sets
tool_choicetoanyortool— it will 400 outright - Whether your conversation history is append-only — injected-then-deleted reminders and rebuilt
system/toolsarrays both invalidate thinking blocks - Whether your team knows about the 30-day data retention requirement, and handles sensitive data under your own compliance policy
Sources: Anthropic’s launch announcement and platform documentation (
anthropic.com/claude-fable-and-mythos-5-1, platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1), released September 1, 2026. APIYI availability and pricing taken from the platform pricing endpoint; data retrieved September 2, 2026 (UTC+8). Final billing follows real-time platform data.