Stacking TOON with Prompt Caching to Cut LLM Costs Further
Prompt caching cuts the price per token up to 90%; TOON cuts token count by about 43% on average. Combine both levers for compounding savings on Claude and GPT.
Yes, you can combine TOON with prompt caching — and the savings multiply. Caching discounts the price per token (Anthropic: 90% off reads on most models; OpenAI GPT-5.x: 90% off cached input, as of October 2026). TOON cuts the number of tokens: 42.6% overall and 58.7% on flat data in the current official benchmark. Stack them by encoding your static context in TOON and placing it in the cached prefix.
Why These Two Cost Levers Are Independent
Most developers treat prompt caching and format optimization as separate concerns. They are not competing strategies — they operate on different dimensions of the cost equation. Your total input cost is:
cost = token_count × price_per_tokenPrompt caching attacks price_per_token. According to the Anthropic prompt caching documentation, cache reads cost 10% of the standard input rate — a 90% discount. Cache writes cost 1.25x the base input rate for a 5-minute TTL or 2x for a 1-hour TTL. As of October 2026, Anthropic's pricing page lists two exceptions: cache reads cost 5% of the input rate on Claude Opus 5.5 and 2.5% on Claude Fable 5.1. On OpenAI, caching activates automatically once a stable prefix exceeds 1,024 tokens. OpenAI's pricing page lists cached input for GPT-5.x models at 10% of the input rate, for example $0.25 versus $2.50 per million tokens on GPT-5.4. Older models such as GPT-4o still bill cached input at 50%.
TOON attacks token_count. The current official toonformat.dev benchmark — 5,856 LLM calls across 244 questions, six formats, and four models — found that TOON uses 42.6% fewer tokens than JSON overall, while scoring 72.2% accuracy versus 71.4% for JSON. On flat datasets the reduction reaches 58.7%; on mixed-structure datasets it is 32.7%. Token counts in the benchmark use OpenAI's o200k_base tokenizer, so check your own provider's counts.
When you apply both levers to the same prefix block, the reduction compounds. A block that starts at 10,000 tokens becomes roughly 6,000 tokens after TOON encoding. Those 6,000 tokens are then cached and re-read at 10% of the normal rate on most current Claude and GPT-5.x models (5% on Claude Opus 5.5). Neither optimization cannibalizes the other.
How to Structure a Prompt for Maximum Cache + TOON Benefit
Prompt caching requires a stable prefix — content that stays identical across requests. This makes it a natural fit for static reference data: product catalogs, user permission tables, knowledge-base excerpts, configuration matrices. That kind of data is also exactly where TOON achieves its largest token savings.
The architecture is straightforward: put your static, TOON-encoded data in the cached portion of the prompt, and inject only the dynamic user question or task at the end.
// ── CACHED PREFIX (sent once, re-read cheaply on every call) ─────────────
// JSON version
[
{"product_id": "P001", "name": "Widget A", "price": 9.99, "in_stock": true},
{"product_id": "P002", "name": "Widget B", "price": 14.99, "in_stock": false},
{"product_id": "P003", "name": "Gadget X", "price": 49.99, "in_stock": true},
{"product_id": "P004", "name": "Gadget Y", "price": 39.99, "in_stock": true},
{"product_id": "P005", "name": "Part Z", "price": 4.99, "in_stock": false}
]
// TOON version (keys declared once; rows hold values only)
products[5]{product_id,name,price,in_stock}:
P001,Widget A,9.99,true
P002,Widget B,14.99,false
P003,Gadget X,49.99,true
P004,Gadget Y,39.99,true
P005,Part Z,4.99,false
// ── DYNAMIC SUFFIX (injected fresh on every call) ────────────────────────
// "Which products are in stock and under $15?"Scale this up to a 500-row product catalog and the TOON version might be 4,800 tokens instead of 9,600. On Claude Sonnet 5.5 ($2 per million input tokens as of October 2026), once cached, you re-read those 4,800 tokens at $0.20 per million instead of paying $2 per million for 9,600 tokens. Note that Claude 4.7 and later models use a newer tokenizer that Anthropic says produces approximately 30% more tokens for the same text, so measure counts on the model you actually use. The write cost is a one-time overhead that amortizes across every subsequent call that hits the cache.
For a broader look at how TOON performs across different use cases, see the API cost optimization guide and the case for replacing JSON in LLM prompts.
Caching Only vs TOON Only vs Both: Relative Cost Effect
The table below uses provider rates as of October 2026 to show the qualitative impact of each strategy. Dollar figures are illustrative — actual costs depend on model, tier, and cache hit rate. All figures assume a large static prefix re-read many times (a typical assistant or RAG pattern).
| Strategy | Token count effect | Price-per-token effect | Relative cost vs baseline |
|---|---|---|---|
| Baseline (JSON, no caching) | 100% (full count) | 100% (standard rate) | 100% |
| Prompt caching only (Anthropic reads) | 100% (unchanged) | 10% of standard (90% off) | ~10% |
| Prompt caching only (Claude Opus 5.5 reads) | 100% (unchanged) | 5% of standard (95% off) | ~5% |
| Prompt caching only (OpenAI GPT-5.x reads) | 100% (unchanged) | 10% of standard | ~10% |
| TOON only (flat table, no caching) | ~41% of JSON (58.7% reduction) | 100% (standard rate) | ~41% |
| TOON only (overall avg, no caching) | ~57% of JSON (42.6% reduction) | 100% (standard rate) | ~57% |
| TOON + Claude caching, 0.1x reads (flat table) | ~41% of JSON | 10% of standard on reads | ~4% |
| TOON + Claude Opus 5.5 caching (flat table) | ~41% of JSON | 5% of standard on reads | ~2% |
The "Both" rows at the bottom are multiplicative: you are not choosing between strategies, you are applying them to different dimensions of the same cost formula. Note that cache write costs are excluded from this table — they are a one-time overhead per TTL window and typically small relative to the cumulative read savings across many calls.
Does Caching Also Stack with the Batch API?
Yes. Both OpenAI and Anthropic offer a Batch API at about 50% off standard rates for asynchronous jobs. Anthropic's pricing page states that prompt caching multipliers stack with the Batch API discount; for OpenAI, check the batch rows on its pricing page for how cached input is billed. Apply TOON on top to cut the token count by roughly 40–59%, and the effective cost per equivalent data payload drops to a fraction of the baseline.
For bulk classification, extraction, or embedding jobs, this three-way combination — TOON encoding, Batch API, and cached prefix — represents the current practical ceiling for LLM cost reduction on data-heavy workloads. The sibling post Halving Costs Twice: The OpenAI Batch API Plus TOON covers the batch side in detail.
Practical Considerations Before You Implement
A few caveats to keep in mind:
- Cache invalidation on data change. If your static context changes, the cache is invalidated and you pay write cost again. Keep truly mutable data out of the cached prefix.
- TOON prompt tax on first call. If the model has not been primed with TOON format instructions, include them in the cached prefix itself — the instructions are static and benefit from caching too. A 2026 arXiv study (2603.03306) found that TOON's format-instruction overhead is a one-time cost that amortizes well on large repeated payloads.
- Minimum prefix size. OpenAI requires a stable prefix of at least 1,024 tokens before caching kicks in. A tiny TOON block might fall below this threshold; in that case, batch the format instructions together with the data to push over the limit.
- Model accuracy variance. In the current official TOON benchmark, accuracy varied far more by model than by format: Grok 4.5 scored 97.1% on TOON while gpt-5.4-nano scored 57.0%, and on each model TOON and JSON landed within about two points of each other (Claude Haiku 4.5: 65.6% vs 63.5%). Validate TOON accuracy on your specific model and query types before committing to a cached TOON prefix in production.
For a full breakdown of which workloads benefit most from TOON, see the guide to building cost-efficient LLM chatbots.
Frequently Asked Questions
Can I combine TOON with prompt caching?
Yes. Prompt caching discounts the price per token: Anthropic charges 10% of the normal input rate on cache reads for most Claude models (5% on Opus 5.5), and OpenAI bills cached input on GPT-5.x models at 10% of the input rate. TOON reduces the number of tokens. The two levers are independent and multiply: a smaller TOON-encoded prefix is cheaper to cache-write and cheaper on every cache-read.
How much does Anthropic charge for cached prompt reads?
As of October 2026, Anthropic charges 10% of the standard input rate for cache reads on most models, a 90% discount. Claude Opus 5.5 reads cost 5% and Claude Fable 5.1 reads cost 2.5%. Cache writes cost 1.25x the base input rate for a 5-minute TTL or 2x for a 1-hour TTL. Sources: Anthropic prompt caching documentation and the ngrok prompt-caching explainer.
What is the minimum prefix length for OpenAI prompt caching?
OpenAI caching activates automatically once a stable prefix exceeds 1,024 tokens. There is no special API call required; as of October 2026, GPT-5.x models bill cached input at 10% of the standard input rate (for example, $0.25 versus $2.50 per million tokens on GPT-5.4).
How much does TOON reduce token count?
According to the official toonformat.dev benchmarks — 5,856 LLM calls across 244 questions and four models — TOON uses 42.6% fewer tokens than JSON overall, and 58.7% fewer on flat datasets. The savings are largest on large, repetitive arrays of objects.
Do prompt caching discounts stack with batch API discounts?
Yes, on Anthropic, and the discounts multiply. OpenAI's Batch API is about 50% off standard rates, and Anthropic's pricing page states that prompt caching multipliers stack with the Batch API discount. Adding TOON on top reduces the token count, so the final cost per equivalent data payload drops even lower.
Recommended Reading
Using TOON with Claude: Token Savings and Prompt Caching
How to pass TOON context to Claude, why to benchmark accuracy on your model first, and how Anthropic's 90% cache-read discount stacks with TOON's token savings.
Halving Costs Twice: The OpenAI Batch API Plus TOON
The OpenAI Batch API takes 50% off input and output tokens; TOON removes about 43% of them first. Combine async batching and compact formatting for bulk jobs.
How to Calculate Your LLM Token Savings from TOON
A simple cost model for TOON token savings: tokens times price, minus the prompt tax, stacked with batch and caching discounts. With a worked example.