Using TOON with Google Gemini
TOON with Google Gemini: Gemini 3.6 Flash read TOON as accurately as JSON in the official benchmark. Feed TOON context to the Gemini API and its large window.
TOON works well with Google Gemini as an input format. In the current official TOON benchmark, gemini-3.6-flash scored 69.3% retrieval accuracy on TOON versus 68.4% on JSON, while TOON used 42.6% fewer tokens overall and 58.7% fewer on flat tables. You get the same accuracy or slightly better for far fewer tokens in Gemini's large context window.
How Accurate Is TOON on Gemini?
The official TOON benchmark (current run, v4.1) made 5,856 LLM calls: 244 questions, six formats, and four models (claude-haiku-4-5-20251001, gemini-3.6-flash, gpt-5.4-nano, and grok-4.5). Token counts use the o200k_base tokenizer via gpt-tokenizer, with reasoning disabled. The datasets cover flat employee tables, time-series, GitHub repo metadata, e-commerce orders, event logs, deep configuration, feature flags, and nested contacts.
gemini-3.6-flash scored 69.3% on TOON and 68.4% on JSON. The benchmark reports Wilson 95% confidence intervals and notes that when two formats' intervals overlap, the difference is not statistically meaningful. So the honest reading is that TOON matches JSON on Gemini, not that it beats it. The real gain is cost, because TOON delivers that accuracy with far fewer tokens.
Across all four models, TOON scored 72.2% against JSON's 71.4% while using 42.6% fewer tokens. That works out to 29.2 accuracy points per 1,000 tokens for TOON versus 16.6 for JSON. An earlier run (early 2026) tested gemini-3-flash-preview and reported much higher absolute accuracy for Gemini. Those figures came from a different question set and model version, so they are not directly comparable with the current run.
See the full accuracy breakdown and methodology in our TOON benchmark deep-dive, or compare TOON against JSON token-by-token in the JSON vs TOON guide.
What Does Gemini Cost per Million Tokens?
As of October 2026, Google's Gemini API pricing page lists Gemini 3.8 Flash (and 3.7 and 3.6 Flash) at $0.75 per million input tokens and $3.75 per million output tokens on the paid tier through December 31, 2026. From January 1, 2027, those rates double to $1.50 input and $7.50 output, and context-cache reads move from $0.075 to $0.15 per million.
| Model (paid tier) | Input / 1M | Output / 1M | Cached input / 1M |
|---|---|---|---|
| Gemini 3.8 / 3.7 / 3.6 Flash (through Dec 31, 2026) | $0.75 | $3.75 | $0.075 |
| Gemini 3.8 / 3.7 / 3.6 Flash (from Jan 1, 2027) | $1.50 | $7.50 | $0.15 |
| Gemini 3.1 Pro Preview (up to 200k tokens) | $2.00 | $12.00 | $0.20 |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | $0.025 |
Here is what the January increase means for input costs. Take a workload that sends 1 billion input tokens of JSON context per month to Gemini 3.8 Flash. That costs $750 today and $1,500 from January. At the overall 42.6% reduction, the same data as TOON is about 574M tokens, roughly $431 today or $861 after the increase. Halving the token count offsets most of the price doubling.
How TOON Extends Gemini's Large Context Window
Gemini's large context window is one of its defining strengths. But a large window does not help if your input format wastes tokens on structural glyphs — repeated keys, braces, and quotes — that carry no information. TOON eliminates that overhead for uniform arrays of objects by declaring fields once in a header and then emitting only values per row.
The token savings versus JSON by data shape, from the current official benchmark, are:
- Flat-only track total: 58.7% fewer tokens
- Employee records (flat table): 60.7% fewer tokens
- Time-series: 59.0% fewer tokens
- GitHub repo data: 41.7% fewer tokens
- E-commerce orders (nested): 32.9% fewer tokens
- Mixed-structure track total: 32.7% fewer tokens
In practical terms: if your Gemini context window holds 1M tokens, encoding context as TOON instead of JSON effectively fits the data equivalent of roughly 2.4M tokens of JSON on flat tables. That is additional retrieved rows, more few-shot examples, or longer conversation history — without a model or pricing tier change.
For a complete cost analysis of how context format affects API spend, see our guide on optimizing API costs with TOON.
Gemini API Request: TOON Context vs JSON Context
The following example shows the same product catalog data sent to the Gemini API first as JSON, then as TOON. Both carry identical information; TOON is significantly more compact.
// ── JSON version ── (keys repeated on every row)
const jsonPayload = {
contents: [
{
role: "user",
parts: [
{
text: `Here is our product catalog in JSON:
[
{"id":1,"name":"Widget A","price":9.99,"stock":142,"category":"tools"},
{"id":2,"name":"Widget B","price":14.49,"stock":87,"category":"tools"},
{"id":3,"name":"Gadget X","price":49.00,"stock":23,"category":"electronics"},
{"id":4,"name":"Gadget Y","price":79.99,"stock":11,"category":"electronics"},
{"id":5,"name":"Part Z","price":2.50,"stock":500,"category":"supplies"}
]
Which products have fewer than 25 units in stock?`
}
]
}
]
};
// ── TOON version ── (same rows, keys declared once)
const toonPayload = {
contents: [
{
role: "user",
parts: [
{
text: `Here is our product catalog in TOON format
(fields declared once in the header; values follow as comma-separated rows):
products[5]{id,name,price,stock,category}:
1,Widget A,9.99,142,tools
2,Widget B,14.49,87,tools
3,Gadget X,49,23,electronics
4,Gadget Y,79.99,11,electronics
5,Part Z,2.5,500,supplies
Which products have fewer than 25 units in stock?`
}
]
}
]
};
// Send to Gemini API
const response = await fetch(
`https://generativelanguage.googleapis.com/v1beta/models/gemini-3.8-flash:generateContent?key=${API_KEY}`,
{
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify(toonPayload),
}
);The TOON version declares the schema once in the header (products[5]{id,name,price,stock,category}:) and then lists values without repeating keys. The encoder also writes numbers in canonical form, so 49.00 becomes 49 and 2.50 becomes 2.5. For five rows the saving is modest, but it grows with every row: on flat tables like this the benchmark measured 58.7% fewer tokens than JSON.
If you need to convert an existing JSON file to TOON before sending it to Gemini, the free JSON to TOON converter handles that client-side with no data leaving your browser. For a primer on TOON syntax, see What is TOON?
Gemini Use Cases: When to Use TOON vs JSON
Not every Gemini prompt benefits from TOON. The format pays off on large, uniform arrays; it adds overhead (the "prompt tax" of format instructions) on small or non-uniform payloads. A 2026 arXiv paper (2603.03306) confirmed TOON's efficiency advantage for comprehension and retrieval while noting that plain JSON outperforms TOON on generation tasks where the model must produce structured output.
| Gemini scenario | Recommended format | Rationale |
|---|---|---|
| Retrieval over a large product / user / log table (50+ rows) | TOON | 58.7% token reduction on flat data; gemini-3.6-flash scored 69.3% on TOON vs 68.4% on JSON |
| RAG context window — injecting retrieved records | TOON | Uniform schema from the vector DB maps directly to a TOON table block; fits more evidence per window |
| Time-series data (sensor readings, stock prices, metrics) | TOON | 59.0% token reduction on time-series data; ideal for 60-day or longer windows |
| Asking Gemini to produce a JSON API response | JSON | arXiv 2603.03306: JSON wins on generation accuracy; do not ask the model to output TOON |
| Single-object or tiny payload (<10 records) | JSON | Prompt tax for TOON format instructions erases the per-row savings at this scale |
| Deeply nested configuration tree | JSON or YAML | Savings shrink on irregular trees (deep config: 34.9% vs JSON but 6.7% more than compact JSON) |
| Long multi-turn conversation with repeated data context | TOON + context caching | Smaller TOON block is cheaper to cache and cheaper to re-read each turn; savings multiply |
TOON and Gemini Context Caching: A Multiplicative Effect
Prompt caching discounts the price per token. TOON reduces the number of tokens. These two effects multiply rather than add.
When you cache a stable TOON context block in Gemini — a product catalog, a user table, a metrics dataset — you write a smaller block to the cache and then re-read a smaller block on every subsequent turn. The caching discount applies to that already-reduced token count. At scale across hundreds of requests, the combined effect can reduce context costs by 70–80% compared to uncached JSON, depending on cache hit rates and data shape.
This pattern is particularly effective for agentic workflows where the same tabular data is referenced repeatedly across multiple tool calls or conversation turns. For a detailed cost breakdown of TOON combined with caching and batch APIs, see our API cost optimization guide.
Model Accuracy Varies: Always Test on Your Data
The 69.3% figure for gemini-3.6-flash is an average across all question types and data shapes. Accuracy varies a lot by task type. Here is the official breakdown across all four models (TOON vs JSON):
- Field retrieval: 97.8% vs 99.2%
- Structure awareness: 90.3% vs 84.0%
- Structural validation: 100.0% vs 50.0%
- Aggregation: 48.4% vs 48.4%
- Filtering: 38.0% vs 41.1%
If your Gemini integration mostly involves aggregation (sum, count, average across rows) or filtering (find rows matching a condition), accuracy is much lower for every format, and JSON is slightly ahead on filtering. For those workloads, consider pre-computing aggregates or filters before the prompt, or evaluate the full benchmark data against your specific question types.
The benchmark also tested a specific model version (gemini-3.6-flash). Newer releases such as Gemini 3.8 Flash may score differently. Run your own accuracy evaluation against a representative sample of your queries before committing to TOON in production — the converter lets you generate TOON from any JSON dataset in seconds.
Frequently Asked Questions
Does TOON work well with Google Gemini?
Yes. In the current official TOON benchmark (5,856 calls), gemini-3.6-flash scored 69.3% retrieval accuracy on TOON versus 68.4% on JSON, so TOON was slightly ahead while using 42.6% fewer tokens overall and 58.7% fewer on flat tables. That lets you fit more data into each Gemini request.
How much does TOON reduce tokens when used with Gemini?
According to the current official toonformat.dev benchmark, TOON uses 42.6% fewer tokens overall than JSON, 32.7% fewer on mixed-structure datasets, and 58.7% fewer on flat tables. On flat data, a 1M-token Gemini context window holds roughly as much as 2.4M tokens of JSON.
Which Gemini model was tested in the TOON benchmark?
The current official TOON benchmark tested gemini-3.6-flash alongside claude-haiku-4-5-20251001, gpt-5.4-nano, and grok-4.5. gemini-3.6-flash scored 69.3% on TOON and 68.4% on JSON. Newer releases such as Gemini 3.8 Flash were not part of that run, so test them on your own data.
Should I use TOON for Gemini output or only for input context?
Use TOON for input context — the data you feed into Gemini — not for output. A 2026 arXiv study (2603.03306) found that plain JSON has the best one-shot and final accuracy when models must produce structured output. TOON's advantage is in comprehension and retrieval, not generation.
Does TOON's token savings stack with Gemini's context caching?
Yes. Prompt caching discounts the price per token; TOON reduces the number of tokens. The two effects multiply: a smaller TOON-encoded context block is cheaper to write into the cache and cheaper to re-read on each turn. This makes the combination especially valuable for long-running conversations with large data payloads.
Recommended Reading
Gemini 3.8 Flash Doubles in Price in 2027: Cut Tokens Now
Gemini 3.8, 3.7 and 3.6 Flash prices double on January 1, 2027. Estimate the impact on your bill and how much sending fewer tokens with TOON can offset.
Feeding GraphQL Responses to LLMs with TOON
GraphQL trims which fields you fetch; TOON trims how you serialize them. Combine a precise selection set with TOON encoding to cut LLM input tokens twice over.
Using TOON with GPT-5 and the OpenAI API
TOON with GPT-5 and the OpenAI API: where TOON cuts input tokens, where to keep JSON for structured outputs, and how prompt caching stacks on top.