Skip to main content
If you’re using Claude Code, Cline, Cursor, or hand-rolling your own Claude API calls, Prompt Cache is the single biggest knob for lowering your bill — cached input tokens are billed at just 0.1×, a 90% discount. This page is based on Anthropic’s official documentation (platform.claude.com/docs/en/build-with-claude/prompt-caching) and adapted to APIYI’s setup with copy-paste-ready examples.

In one sentence

Mark a long, reused prompt prefix (system instructions / a long document / few-shot examples) with cache_control. The server stores it; on the next request with the same prefix, it skips reprocessing — roughly 10× cheaper and faster. It expires after a period of inactivity.

Why bother — read the multipliers

Relative to the model’s base input token price (): Break-even points:
  • 5-min TTL: only 2 reuses of the same prefix to break even (1.25 + 0.1 = 1.35, cheaper than 2.0 for two uncached requests).
  • 1-hour TTL: 3 reuses to break even (2 + 0.2 = 2.2, cheaper than 3.0).
TTL is a sliding window: every cache hit resets the expiration timer, so active conversations don’t expire from under you. Only true idleness beyond the TTL causes eviction.

Good fit

  • The same long system prompt called many times (agents, chatbots)
  • Multi-turn conversations (every prior turn becomes reusable prefix)
  • Batch processing of one document (asking 50 questions about one contract)
  • RAG, where stable retrieved chunks form the prefix

Bad fit

  • Every prompt differs from the first character onward
  • The whole thing is short and never crosses the per-model minimum (below)

The three hard requirements

All three are mandatory.

1. Explicit cache_control marker

content cannot be a plain string. It must be a content block array, with cache_control attached to the block you want cached:

2. Length must clear the per-model minimum

If the content is shorter than the model’s minimum, it won’t be cached even with the marker (no error, just silently skipped). Verified against Anthropic’s official docs:
This threshold does not decrease monotonically with version number — don’t guess it. The most counterintuitive pairs: Opus 5 needs only 512, while the older Opus 4.6 / 4.5 need 4,096 — an 8× difference. Haiku 4.5 is also 4,096, which is higher than the older Haiku 3.5 (2,048). So neither “newer models have lower thresholds” nor “smaller models have lower thresholds” holds. Check the table whenever you switch models.
English text averages roughly 0.75 words per token. In practice: Opus 5 caches from about 380+ words of stable content, Sonnet 5 / Sonnet 4.6 need around 770 words, and Opus 4.6 / Haiku 4.5 need roughly 3,000 words before caching does anything. Always refer to Anthropic’s official docs for the latest thresholds — they can change between model versions.
Measured on APIYI (2026-07-29). We probed the write threshold with a fixed prefix stepped up in size: claude-opus-5 produced no cache write at 301 tokens but did at 614, bracketing the official 512; claude-sonnet-5 produced none at 612 but did at 1,250, bracketing the official 1,024. Both match the table above.

3. Prefix must match byte-for-byte

Caching is prefix-based: from the start of the request up to the cache_control marker, the byte stream must be identical to the previous request. Any single character change — whitespace, JSON key ordering, a timestamp — counts as a new prefix and triggers a fresh write instead of a hit. Practical rule: stable stuff up front, volatile stuff at the back.

Minimal runnable example

Send two requests using the same long document but different questions. The first writes, the second hits:
Expected output:
The second call’s read ≈ the first call’s write — the same prefix is being reused.

How to tell whether you hit — three usage fields

In every response, usage reports: Total input tokens = sum of all three. As long as cache_read_input_tokens > 0, you’re saving money.

Most common pitfalls

Prompt Cache only works on the Anthropic native format (/v1/messages). When you call Claude through the OpenAI-compatible format (/v1/chat/completions), no cache fields will come back regardless of what you send. For Claude Code, Cline, Cursor and similar high-frequency clients, the native format is mandatory if you care about your bill.

Advanced: multi-turn conversations

Place cache_control on the last content block of the most recent user message. Each new turn auto-extends the cached read range up to the end of the previous turn:
Two hard limits to keep in mind:
  • At most 4 cache_control breakpoints per request.
  • Each breakpoint’s prefix lookup window is at most 20 content blocks back — anything older than that won’t be considered for a hit. In other words, in very long conversations, marking only the latest turn won’t cover the entire prior history.
A common pattern: place one breakpoint each on tool definitions, system prompt, long documents, and the latest conversation turn — using all 4 slots so that sections changing at different rates don’t invalidate each other’s cache.

On APIYI and caching

APIYI forwards cache fields end-to-end. The cache_control you send is passed through to upstream Claude (AWS Claude or Claude Official) as-is, and the returned cache_creation_input_tokens / cache_read_input_tokens are passed straight back to you — no special adaptation needed in your code.
How to self-verify:
  1. On the first request, usage.cache_creation_input_tokens > 0 (write succeeded).
  2. Within seconds, send the same prefix again — you should see usage.cache_read_input_tokens > 0 (hit).
  3. Your billing dashboard will itemize cache writes and cache reads separately, at the same official multipliers (1.25× / 2× / 0.1×).

Recap

1. Mark it

cache_control: {"type": "ephemeral"} on a content block — plain-string content is never cached.

2. Long enough

Opus 5 ≥ 512; Sonnet 5 / Sonnet 4.6 ≥ 1,024; Opus 4.7 ≥ 2,048; Opus 4.6 / Haiku 4.5 ≥ 4,096 tokens, otherwise silently skipped.

3. Stable prefix

Stable up front, volatile in the back; one character of drift kills the hit.

4. Check usage

Only cache_read_input_tokens > 0 proves you actually saved money.