Gemini 3.8 Flash Doubles in Price in 2027: Cut Tokens Now
Gemini 3.8, 3.7 and 3.6 Flash prices double on January 1, 2027. Estimate the impact on your bill and how much sending fewer tokens with TOON can offset.
Gemini 3.8 Flash, 3.7 Flash and 3.6 Flash double in price on January 1, 2027: input goes from $0.75 to $1.50 per million tokens and output from $3.75 to $7.50. You cannot negotiate the rate, but you can send fewer tokens. Encoding structured input as TOON offsets part of the increase.
What exactly changes in Gemini Flash pricing?
Google's Gemini API pricing page (checked October 1, 2026) lists a dated schedule for gemini-3.8-flash: input is "$0.75 through December 31, 2026. $1.50 starting January 1, 2027." The same schedule applies to gemini-3.7-flash and gemini-3.6-flash. All prices are per million tokens, standard tier.
| gemini-3.8 / 3.7 / 3.6-flash | Through Dec 31, 2026 | From Jan 1, 2027 |
|---|---|---|
| Input | $0.75 | $1.50 |
| Output | $3.75 | $7.50 |
| Context caching | $0.075 | $0.15 |
| Batch input | $0.375 | $0.75 |
| Batch output | $1.875 | $3.75 |
Every line doubles, so no billing mode escapes it. For comparison, the same page lists gemini-3.5-flash at $1.50 input and $9.00 output with no dated change, and gemini-3.1-flash-lite at $0.25 input and $1.50 output for text. After January 1, the newer Flash models cost the same as 3.5 Flash for input and less for output.
How much will the Gemini Flash increase cost you?
Since input and output both double, your Flash bill doubles too, unless usage changes. Take an example workload of 500 million input tokens and 50 million output tokens per month on gemini-3.8-flash:
| Scenario (per month) | Input cost | Output cost | Total |
|---|---|---|---|
| 2026 prices, JSON input | $375 | $187.50 | $562.50 |
| 2027 prices, JSON input | $750 | $375 | $1,125 |
| 2027 prices, TOON for structured input | $570 | $375 | $945 |
The TOON row assumes 60% of input tokens are structured data (tool results, retrieved rows, catalogs) and that TOON shrinks that part by 40%. That removes 120 million input tokens, saving $180 a month at the 2027 rate, or about a third of the $562.50 increase. Plug in your own split. If your prompts are mostly prose, the offset is small. If they are mostly tables, it is larger. Our LLM cost calculator guide walks through the same math.
How much does TOON save on Gemini input?
The official TOON benchmark (spec v4.1, 5,856 LLM calls) reports 42.6% fewer tokens than JSON, with 72.2% accuracy vs 71.4% for JSON (retrieval-accuracy results). Savings depend on data shape: 58.7% on flat tables and 32.7% on mixed structures (token-efficiency results). Those counts use OpenAI's o200k_base tokenizer. Gemini tokenizes differently, and the benchmark notes that relative differences "hold directionally." Measure on Gemini's own token counter before you budget.
Here is a small product list, as pretty JSON and as TOON (4.1.1 encoder output):
{
"products": [
{
"sku": "SKU-1042",
"name": "Trail Running Shoe",
"price": 129.99,
"stock": 42,
"rating": 4.6
},
{
"sku": "SKU-1043",
"name": "Merino Base Layer",
"price": 79.5,
"stock": 118,
"rating": 4.8
},
{
"sku": "SKU-1044",
"name": "Packable Rain Shell",
"price": 189,
"stock": 7,
"rating": 4.4
}
]
}products[3]{sku,name,price,stock,rating}:
SKU-1042,Trail Running Shoe,129.99,42,4.6
SKU-1043,Merino Base Layer,79.5,118,4.8
SKU-1044,Packable Rain Shell,189,7,4.4With o200k_base we counted 153 tokens for the pretty JSON, 92 for minified JSON and 73 for TOON. Even against minified JSON, the TOON version is about a fifth smaller, and the gap grows with more rows.
Does Gemini read TOON accurately?
gemini-3.6-flash is one of the four models in the current benchmark. It scored 69.3% with TOON and 68.4% with JSON across 244 questions (source). The intervals (about ±5.8 points) overlap, so the fair reading is "no accuracy loss," not "TOON is more accurate." Minified JSON scored lower for this model, at 63.5%. Our TOON with Google Gemini guide covers SDK setup and prompt wording.
What else cuts the Gemini bill before January?
- Batch API. Batch is 50% of standard on every listed model. Moving async jobs to batch roughly cancels the increase for those jobs.
- Context caching. Cached input will cost $0.15 per million, a tenth of the new input rate. Caching and TOON stack: a smaller cached block is cheaper to write and to re-read. See TOON for prompt caching.
- Route simple tasks to Flash-Lite.
gemini-3.1-flash-liteis listed at $0.25 input and $1.50 output. - Trim fields before encoding. A column the model never uses costs tokens in any format.
Keep TOON's limits in mind. It shrinks input, not output, so output-heavy workloads such as long-form generation gain little. On deeply nested or semi-uniform data, minified JSON can be smaller. If your Flash calls are mostly about reading tables, rows and tool results, now is a good time to switch, while the cheaper rate still applies and you can measure the difference.
Frequently Asked Questions
When does the Gemini 3.8 Flash price increase take effect?
On January 1, 2027. Google's pricing page lists gemini-3.8-flash input at $0.75 per million tokens through December 31, 2026 and $1.50 starting January 1, 2027. Output goes from $3.75 to $7.50 and context caching from $0.075 to $0.15. Gemini 3.7 Flash and 3.6 Flash follow the same schedule.
Does the Gemini Batch API avoid the price increase?
No, but it halves it. Batch prices for gemini-3.8-flash also double, from $0.375 to $0.75 per million input tokens and from $1.875 to $3.75 per million output tokens. Batch stays 50% cheaper than standard, so moving async jobs to batch roughly cancels the increase for those jobs.
How much of the Gemini price increase can TOON offset?
TOON only cuts input tokens, and only the structured-data part of them. In the official benchmark TOON used 42.6% fewer tokens than JSON overall. In our example workload, where structured data is 60% of input, TOON offsets about a third of the increase. Output-heavy workloads see less benefit.
Does Gemini understand TOON as well as JSON?
In the current official TOON benchmark, gemini-3.6-flash scored 69.3% accuracy with TOON and 68.4% with JSON on 244 retrieval questions. The confidence intervals overlap, so treat that as parity rather than a win. Test on your own prompts before switching production traffic.
Recommended Reading
Using TOON with Google Gemini
Gemini scored the highest TOON retrieval accuracy in the official benchmark. Learn how to feed TOON context to the Gemini API and make the most of its large context window.
How to Calculate Your LLM Token Savings from TOON
A simple cost model for estimating real savings from TOON: tokens times price, minus the prompt tax, stacked with batch and caching discounts. With a worked example.
Halving Costs Twice: The OpenAI Batch API Plus TOON
The Batch API takes 50% off input and output tokens; TOON removes 40-60% of them first. Here's how to combine async batching and token-efficient formatting for bulk LLM jobs.