platform.claude.com/docs/en/build-with-claude/prompt-caching) and adapted to APIYI’s setup with copy-paste-ready examples.
In one sentence
Mark a long, reused prompt prefix (system instructions / a long document / few-shot examples) withcache_control. The server stores it; on the next request with the same prefix, it skips reprocessing — roughly 10× cheaper and faster. It expires after a period of inactivity.
Why bother — read the multipliers
Relative to the model’s base input token price (1×):
Break-even points:
- 5-min TTL: only 2 reuses of the same prefix to break even (1.25 + 0.1 = 1.35, cheaper than 2.0 for two uncached requests).
- 1-hour TTL: 3 reuses to break even (2 + 0.2 = 2.2, cheaper than 3.0).
TTL is a sliding window: every cache hit resets the expiration timer, so active conversations don’t expire from under you. Only true idleness beyond the TTL causes eviction.
Good fit
- The same long system prompt called many times (agents, chatbots)
- Multi-turn conversations (every prior turn becomes reusable prefix)
- Batch processing of one document (asking 50 questions about one contract)
- RAG, where stable retrieved chunks form the prefix
Bad fit
- Every prompt differs from the first character onward
- The whole thing is short and never crosses the per-model minimum (below)
- Chat clients that re-assemble uploaded attachments on every turn (see “Images and files cache too” below)
The three hard requirements
All three are mandatory.1. Explicit cache_control marker
content cannot be a plain string. It must be a content block array, with cache_control attached to the block you want cached:
2. Length must clear the per-model minimum
If the content is shorter than the model’s minimum, it won’t be cached even with the marker (no error, just silently skipped). Verified against Anthropic’s official docs:Measured on APIYI (2026-07-29). We probed the write threshold with a fixed prefix stepped up in size:
claude-opus-5 produced no cache write at 301 tokens but did at 614, bracketing the official 512; claude-sonnet-5 produced none at 612 but did at 1,250, bracketing the official 1,024. Both match the table above.3. Prefix must match byte-for-byte
Caching is prefix-based: from the start of the request up to thecache_control marker, the byte stream must be identical to the previous request. Any single character change — whitespace, JSON key ordering, a timestamp — counts as a new prefix and triggers a fresh write instead of a hit.
Practical rule: stable stuff up front, volatile stuff at the back.
Minimal runnable example
Send two requests using the same long document but different questions. The first writes, the second hits:read ≈ the first call’s write — the same prefix is being reused.
How to tell whether you hit — three usage fields
In every response,usage reports:
Total input tokens = sum of all three. As long as
cache_read_input_tokens > 0, you’re saving money.
Most common pitfalls
Images and files cache too — but attachment scenarios hit less reliably
First, a common misconception: it is not that Claude’s caching of images/files is “unreliable.” Anthropic documents explicitly that text, images and documents can all enter the Prompt Cache, on equal terms. The real reason is the hard requirement above: caching matches the complete prefix of the request, and the bytes of an image or file are part of that prefix. Any change in how the client assembles attachments between turns breaks it. Anthropic also calls out one case specifically: changing whether a prompt contains an image invalidates the corresponding cache.How the two scenarios differ
Coding clients like Claude Code, Cline and Cursor fall in the first column, which is why their hit rates are usually steady. GUI clients like Chatbox and Cherry Studio are equally steady for plain-text conversations — the hit rate only degrades when the conversation carries uploaded files or large image attachments.
Recovering the hit rate in attachment scenarios
- Put attachments first, in a fixed order — they are the largest and most stable content, so the prefix is where they belong
- Encode each attachment once and reuse that same base64; don’t re-read and re-encode the file every turn (re-encoding is not guaranteed to produce identical bytes)
- Don’t reorder content blocks between turns, especially not “image first this turn, image last the next”
- Put
cache_controlafter the attachments so they land inside the cached prefix - Avoid adding or removing attachments mid-conversation — adding or dropping one image invalidates the entire prefix after it
Advanced: multi-turn conversations
Placecache_control on the last content block of the most recent user message. Each new turn auto-extends the cached read range up to the end of the previous turn:
- At most 4
cache_controlbreakpoints per request. - Each breakpoint’s prefix lookup window is at most 20 content blocks back — anything older than that won’t be considered for a hit. In other words, in very long conversations, marking only the latest turn won’t cover the entire prior history.
On APIYI and caching
APIYI forwards cache fields end-to-end. The
cache_control you send is passed through to upstream Claude (AWS Claude or Claude Official) as-is, and the returned cache_creation_input_tokens / cache_read_input_tokens are passed straight back to you — no special adaptation needed in your code.- On the first request,
usage.cache_creation_input_tokens > 0(write succeeded). - Within seconds, send the same prefix again — you should see
usage.cache_read_input_tokens > 0(hit). - Your billing dashboard will itemize cache writes and cache reads separately, at the same official multipliers (1.25× / 2× / 0.1×).
Recap
1. Mark it
cache_control: {"type": "ephemeral"} on a content block — plain-string content is never cached.2. Long enough
Opus 5 ≥ 512; Sonnet 5 / Sonnet 4.6 ≥ 1,024; Opus 4.7 ≥ 2,048; Opus 4.6 / Haiku 4.5 ≥ 4,096 tokens, otherwise silently skipped.
3. Stable prefix
Stable up front, volatile in the back; one character of drift kills the hit.
4. Check usage
Only
cache_read_input_tokens > 0 proves you actually saved money; it should roughly equal the previous turn’s cache_creation_input_tokens.5. Keep attachments fixed
Images and files cache too, but re-assembling attachments each turn kills the prefix — put them first, keep the order, encode once.
Related links
- Parent page: Claude API Basics
- Client setup guides: Claude Code integration · Cherry Studio integration
- Get / manage tokens:
https://api.apiyi.com/token - Anthropic official docs:
platform.claude.com/docs/en/build-with-claude/prompt-caching