Using TOON with Claude: Token Savings and Prompt Caching on the Anthropic API
How to pass TOON context to Claude, why you should benchmark accuracy on your model first, and how Anthropic's 90% cache-read discount stacks with TOON's token savings.
To use TOON with Claude, encode your retrieved context as TOON in the input messages, place it in a cacheable prefix, and keep any structured output schema as JSON. TOON used 42.6% fewer tokens than JSON in the current official benchmark, and Anthropic cache reads cost 10% of the base input rate on most models, so the two discounts multiply. Test accuracy on your specific Claude model first. Prices below are as of October 2026.
How Accurate Is TOON with Claude?
The current official TOON benchmark (v4.1) ran 5,856 LLM calls: 244 questions, six formats, and four models. Overall, TOON scored 72.2% accuracy versus 71.4% for JSON while using 42.6% fewer tokens. The Claude model in the run is claude-haiku-4-5-20251001, which scored 65.6% on TOON versus 63.5% on JSON. TOON was slightly ahead on Claude, at a much lower token cost.
The benchmark uses Wilson 95% confidence intervals, and differences inside overlapping intervals are not statistically meaningful. Read the Haiku result as "TOON is at least as accurate as JSON here", not as a large win. Sonnet 5.5, Opus 5.5 and Fable 5.1 are not in the benchmark, so their TOON accuracy is unknown from this source. Before committing TOON to a production Claude pipeline, run a representative sample of your actual questions against your actual Claude model.
Accuracy also depends on the question type. Across all models, field retrieval ("what is field X for record Y") scored 97.8% on TOON versus 99.2% on JSON, while filtering scored 38.0% versus 41.1%. Structural validation was where TOON clearly led: 100% versus 50%, because the [N] length and field header let the model detect truncated or malformed data. An earlier run from early 2026 reported different per-model figures; those are superseded by the numbers above.
What a TOON Context Block Looks Like in a Claude Messages Call
Below is the same API call structured two ways — first with JSON context, then with TOON placed in a cacheable prefix. The TOON version is shorter and eligible for Anthropic's cache discount on every repeat call.
// --- JSON context (verbose; keys repeat per row) ---
{
"model": "claude-sonnet-5-5",
"system": "You are a data analyst. Answer based only on the provided records.",
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "Sales records:\n[ {\"region\":\"North\",\"rep\":\"Alice\",\"q1\":42000,\"q2\":38500}, {\"region\":\"South\",\"rep\":\"Bob\", \"q1\":31000,\"q2\":44000}, {\"region\":\"East\", \"rep\":\"Carol\",\"q1\":55000,\"q2\":51000}, {\"region\":\"West\", \"rep\":\"Dave\", \"q1\":29000,\"q2\":33000}]\nWhich rep had the highest Q2?"
}
]
}
]
}
// --- TOON context with cache_control on the prefix ---
{
"model": "claude-sonnet-5-5",
"system": [
{
"type": "text",
"text": "You are a data analyst. Answer based only on the provided records. Data is TOON format: header declares field names and count; rows are comma-separated values.",
"cache_control": { "type": "ephemeral" }
}
],
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "Sales records:\n\nsales[4]{region,rep,q1,q2}:\n North,Alice,42000,38500\n South,Bob,31000,44000\n East,Carol,55000,51000\n West,Dave,29000,33000",
"cache_control": { "type": "ephemeral" }
},
{
"type": "text",
"text": "Which rep had the highest Q2?"
}
]
}
]
}The cache_control breakpoint tells Anthropic's API to treat everything up to that point as a cacheable prefix. On the first call, the prefix is written to cache at 1.25x the base input rate (5-minute TTL) or 2x (1-hour TTL). On every subsequent call with the same prefix, it is re-read at 10% of the base rate on most models, 5% on Opus 5.5 and 2.5% on Fable 5.1. Because the TOON block is already about 40% smaller than the equivalent JSON, the cache-write cost is lower and each cache read is cheaper too. The question sits in a separate, uncached block so it can change per call. A four-row table is far below the minimum cacheable prompt length; in practice the cached block is a large dataset.
Claude Usage Scenarios: Which Format and Caching Approach to Use
The right strategy depends on the shape of your data, the size of your payload, and whether Claude is reading or writing structured content.
| Scenario | Recommended format | Caching note |
|---|---|---|
| Large uniform array (database rows, search results, logs) as context | TOON | Set cache_control on the context block; about 40% fewer tokens cuts write cost and re-read cost |
| Small context (<10 objects) or one-off payload | JSON | Prompt-tax overhead erases TOON savings; JSON requires no format instruction |
| Flat table data (uniform fields, no nesting) | TOON | 58.7% fewer tokens than JSON on flat datasets in the official benchmark; verify Claude model accuracy before production use |
| Deeply nested or non-uniform document | JSON or YAML | Savings shrink on irregular data, and compact JSON can beat TOON; measure before switching |
| Claude generating structured output (tool use, extraction) | JSON | arXiv 2603.03306: JSON wins on generation accuracy; use Claude structured outputs (GA) for schema-valid output |
| Tool definitions (when tools are supplied) | JSON Schema | Tool-use system prompt adds 286 tokens on Sonnet 5.5 and Opus 5.5, 496 on Haiku 4.5; keep tool order stable so it caches |
| Static context reused across many calls (RAG, system knowledge) | TOON + 1-hour cache | 2x write cost pays off after two reads; the 90% or larger discount on re-reads compounds across hundreds of calls |
How Anthropic Prompt Caching Works and Why Token Count Matters
Anthropic's prompt caching is opt-in: you add a cache_control object with {"type": "ephemeral"} to mark the end of the prefix you want cached. The pricing is asymmetric by TTL:
- Cache write (5-minute TTL): 1.25x the base input rate
- Cache write (1-hour TTL): 2x the base input rate
- Cache read: 10% of the base input rate (a 90% discount) on most models, 5% on Claude Opus 5.5 and 2.5% on Claude Fable 5.1
As of October 2026, Anthropic's pricing page lists Claude Sonnet 5.5 at $2 input / $10 output per million tokens, Opus 5.5 at $4 / $20, Haiku 4.5 at $1 / $5, and Fable 5.1 at $10 / $50. Cache reads cost $0.20 per million tokens on both Sonnet 5.5 (10% of $2) and Opus 5.5 (5% of $4).
According to Anthropic, caching pays off after one cache read for the 5-minute TTL (1.25x write) and after two reads for the 1-hour TTL (2x write).
TOON changes the economics on both sides. On the write side, a context block that is 40% smaller means a write cost that is 40% smaller. On the read side, every subsequent cache hit is charged on fewer tokens. For workloads that repeat the same context frequently — RAG pipelines, chatbots with a shared knowledge base, agents that always start with the same tool context — the compounding is significant.
Tool definitions also count toward the prefix. Anthropic adds a tool-use system prompt whenever tools are supplied: 286 tokens on Sonnet 5.5 and Opus 5.5, and 496 tokens on Haiku 4.5 with tool_choice set to auto (588 with any or tool). Keep tool definitions identical between calls so they stay inside the stable, cacheable prefix.
Does the newer Claude tokenizer change the math?
Yes, for absolute counts. Anthropic says Claude 4.7 and later models use a newer tokenizer that "produces approximately 30% more tokens for the same text" (source). If you move a pipeline from Sonnet 4.6 to Sonnet 5.5, the same JSON or TOON payload can bill around 30% more tokens, even though the per-token price is lower. TOON's relative savings come from removing repeated keys and punctuation, so they should carry over, but the benchmark measured tokens with OpenAI's o200k_base tokenizer. Count tokens with Anthropic's token-counting endpoint on your own data before you rely on a specific percentage.
For a broader treatment of caching strategy across providers, see our API cost optimization guide.
This Guide vs. the Claude Efficiency Post
The Claude and TOON efficiency guide covers token savings and cost math for the current Claude lineup. This guide is focused on the practical API integration: how to structure messages, where to place cache_control breakpoints, how to handle tool overhead, and when to fall back to JSON. Read the efficiency post first for the numbers; come back here for the implementation pattern.
The sibling post Using TOON with GPT-5 and the OpenAI API covers the same integration pattern for OpenAI's API, where caching is automatic rather than opt-in. In the current benchmark, gpt-5.4-nano scored 57.0% on TOON versus 57.4% on JSON, a statistical tie, while Claude Haiku 4.5 scored 65.6% versus 63.5%.
Practical Integration Steps for the Anthropic SDK
Step 1 — Convert your context data to TOON. Use the free json2toon.co converter for one-off testing, or the @toon-format/toon npm package for programmatic conversion. Target uniform arrays — database result sets, retrieved document chunks, structured log batches.
Step 2 — Add a brief format note to your system prompt. One sentence is enough: "Context data is encoded in TOON format. The header line declares field names and record count; rows are comma-separated values." Mark the system message with cache_control if the instructions are stable across calls.
Step 3 — Place TOON data early in the messages array, before any variable content. Everything up to the last cache_control breakpoint is cacheable. The user's actual question should come after, unmarked, so it can vary per call without busting the cache.
Step 4 — Keep output schemas as JSON. Use tool_use blocks or ask for JSON in the final user message. The arXiv 2603.03306 research is clear: for generation, JSON wins on reliability. Claude structured outputs are generally available: set output_config.format with a JSON schema, or strict: true on tools, to get schema-valid output.
Step 5 — Measure accuracy on your model before going to production. Claude Haiku 4.5's 65.6% in the benchmark is a starting point, not a ceiling or a floor for every task. Field retrieval on well-structured TOON data performs far better than aggregation or filtering. Run your own evaluation on a representative sample. See the TOON best practices guide for evaluation tips and edge-case handling.
Frequently Asked Questions
How do I use TOON with the Anthropic API?
Encode your context data as TOON and place it in a cacheable prefix — ideally inside the system message or an early user turn — then add a cache_control breakpoint. Keep Claude's structured output schema as JSON. TOON used 42.6% fewer tokens than JSON in the current official benchmark; Anthropic cache reads cost 10% of the input rate on most models (5% on Opus 5.5), so the two discounts multiply.
How accurate is TOON with Claude?
In the current official toonformat.dev benchmark (5,856 calls), Claude Haiku 4.5 scored 65.6% retrieval accuracy on TOON versus 63.5% on JSON, so TOON was slightly ahead while using far fewer tokens. Larger Claude models such as Sonnet 5.5 and Opus 5.5 were not tested. Always benchmark TOON accuracy on your specific Claude model and task type before committing to production.
How does Anthropic prompt caching work with TOON?
On most Claude models, cache reads cost 10% of the base input rate, a 90% discount; Opus 5.5 reads cost 5% and Fable 5.1 reads cost 2.5%. Cache writes cost 1.25x (5-minute TTL) or 2x (1-hour TTL) the base rate. A TOON context block is typically about 40% smaller than the equivalent JSON block, so both the write cost and the per-read cost are lower.
Should Claude output TOON or JSON?
Use JSON for output. A 2026 arXiv study (2603.03306) found plain JSON had the best one-shot and final accuracy when a model generates structured data. Claude structured outputs are generally available, so defining output and tool schemas in JSON with strict mode keeps generated output schema-valid.
Does tool use affect token counts when using TOON with Claude?
Yes. Anthropic adds a tool-use system prompt whenever tools are supplied: 286 tokens on Sonnet 5.5 and Opus 5.5, and 496 tokens on Haiku 4.5 with tool_choice auto (588 with any or tool). Keep tool definitions stable and place your TOON context block early in the prompt so both fall inside the cacheable prefix and the overhead is amortized across cache hits.
Recommended Reading
Stacking TOON with Prompt Caching to Cut LLM Costs Further
Prompt caching cuts the price per token up to 90%; TOON cuts the number of tokens up to 60%. Learn how to combine both levers for compounding savings on Claude and GPT.
TOON Benchmarks 2026: Token Savings and Accuracy Across GPT-5, Claude, Gemini & Grok
A data-driven look at TOON vs JSON across 5,856 LLM calls: 42.6% fewer tokens at 72.2% vs 71.4% retrieval accuracy, plus per-model and per-data-shape results.
How to Calculate Your LLM Token Savings from TOON
A simple cost model for estimating real savings from TOON: tokens times price, minus the prompt tax, stacked with batch and caching discounts. With a worked example.