9 min read

How TOON Handles Nested and Non-Uniform Data

TOON shines on uniform arrays, but real data nests. Learn how TOON represents nested objects and mixed structures, where savings drop, and when JSON or YAML wins.

By json2toon.co

TOON handles nested data with YAML-style indentation for objects and CSV-style tabular rows for uniform arrays. Since spec 4.0, uniform nested objects can stay inside a table through nested field groups. Savings are highest on flat or uniformly nested data (58.7% to 66.5% fewer tokens than JSON) and shrink on irregular nesting: 32.9% for e-commerce orders and 15.0% for event logs.

How TOON Represents Nested Data

TOON was designed by Johann Schopplich and contributors around two complementary syntaxes. For uniform arrays of objects, the case where TOON excels, it uses a CSV-style tabular block: a single header line that declares the array length and field names, followed by one plain-value row per object. For nested objects and non-uniform structures, it falls back to YAML-style indentation.

The header line for a tabular block looks like items[3]{sku,qty,unit_price}:. That one line gives the LLM an explicit schema and a row count to validate against. On purely flat tables, the TOON README puts the cost of this structure at roughly 5 to 10% more tokens than CSV, traded for higher extraction accuracy. When an object property is itself an object rather than a scalar, TOON indents it beneath the parent key using the same whitespace-significant convention as YAML.

This design means TOON is not a single homogeneous format but a hybrid that picks the right representation for each sub-structure. The practical consequence is that savings depend on how much of your payload is uniform-array versus irregular-object.

JSON vs TOON on a Nested Object: A Concrete Example

Consider an order object that contains a uniform array of line items. Here is the data as JSON:

{
  "order_id": "ORD-9041",
  "customer": "Priya Nair",
  "status": "shipped",
  "items": [
    {"sku": "HDMI-4K", "qty": 2, "unit_price": 14.99},
    {"sku": "USB-C-HUB", "qty": 1, "unit_price": 39.99},
    {"sku": "PATCH-CAT8", "qty": 3, "unit_price": 8.49}
  ]
}

And the same data as TOON, generated with encode() from @toon-format/toon 4.1.1:

order_id: ORD-9041
customer: Priya Nair
status: shipped
items[3]{sku,qty,unit_price}:
  HDMI-4K,2,14.99
  USB-C-HUB,1,39.99
  PATCH-CAT8,3,8.49

The outer object gets YAML-style key-value pairs. TOON does not try to tabularize it because there is only one object, not a repeating array. The inner items array is uniform (three objects, same three primitive fields), so it becomes a tabular block. That block is where most of the savings come from. The outer object saves less because there is nothing to amortize the key names across.

For a deeper look at how TOON compares to JSON on flat data, see the JSON vs TOON token comparison.

What Are Nested Field Groups in TOON v4?

Before spec 4.0, any nested object inside an array item pushed the whole array out of tabular form and into an expanded list. Spec 4.0 added nested field groups: when every row has a nested object with the same primitive fields, the header lists those fields in braces and the row carries the values inline. Here is a contacts array with nested address and phone objects:

contacts[2]{id,name,address{city,zip},phone{mobile,work}}:
  1,Ada Lovelace,London,NW1,555-0101,555-0201
  2,Alan Turing,Wilmslow,SK9,555-0102,555-0202

This decodes back to the original objects without loss. The official benchmark's contacts dataset, which uses this shape, comes out 66.5% smaller than JSON, the largest saving of any dataset. Spec 4.0 also added a keyed tabular form for maps of uniform objects, such as feature flags keyed by name:

flags[3:]{enabled,rollout}:
  darkMode: true,50
  newCheckout: false,0
  betaSearch: true,10

The feature-flags dataset (keyed) saves 54.6% versus JSON. Both forms are emitted automatically by the 4.x encoder. The older key folding feature (a.b.c: 1) was removed in spec 4.0, and the library now ignores the keyFolding option.

Token Savings by Data Shape: What the Benchmark Shows

The current official toonformat.dev benchmark (v4.1) ran 5,856 LLM calls across 244 questions, six formats and four models. Token reduction versus formatted JSON, per the token-efficiency results, breaks down by data shape as follows:

  • Contacts (nested field groups): 66.5%
  • Employees (flat uniform table): 60.7%
  • Time-series: 59.0%
  • Feature flags (keyed table): 54.6%
  • GitHub repo data: 41.7%
  • Deep configuration: 34.9% (but 6.7% larger than compact JSON)
  • E-commerce orders (nested): 32.9%
  • Event logs (semi-uniform): 15.0% (but 19.9% larger than compact JSON)

Across the two tracks, flat-only datasets total 58.7% fewer tokens than JSON, and mixed-structure datasets total 32.7%. The mixed track is only 1.6% smaller than compact JSON overall, which is the honest picture for nested data: most of TOON's advantage over minified JSON comes from its tables.

E-commerce orders are real-world nested data: an array of orders, each containing a nested customer object and a nested line-items array. Because items contain a sub-array, the outer array cannot be a table, but every line-items array inside it still is. Event logs compress the least because events of different types carry different fields, so TOON cannot exploit tabular repetition.

Overall accuracy in the same run was 72.2% for TOON versus 71.4% for JSON, with an efficiency of 29.2 accuracy points per 1,000 tokens versus JSON's 16.6. The 95% confidence intervals overlap, so the accuracy difference is not statistically meaningful. The token saving is.

Data Shape, TOON Representation, and Token Reduction: Full Comparison

Data shapeHow TOON represents itToken reduction vs JSONBetter alternative (if any)
Flat uniform array (same fields every row)Tabular block: header once, delimited rows60.7% (employees)CSV (if purely flat, zero nesting)
Uniform rows with nested objects of the same shapeTabular block with nested field groups (v4)66.5% (contacts)None. TOON is optimal
Time-series (uniform rows, timestamp + value)Tabular block59.0%CSV is slightly smaller if no nesting
Map of uniform objects (feature flags)Keyed tabular block [N:] (v4)54.6%None at this saving level
GitHub repo / API response (moderate nesting)Tabular blocks for repeated sub-arrays; indentation for unique fields41.7%None at this saving level
E-commerce orders (nested customer + line items)Outer array as expanded list; each line-items array as a table32.9%None unless nesting is very deep
Very deep, non-repetitive nesting (config tree)YAML-style indentation only34.9%, but 6.7% larger than compact JSONCompact JSON or YAML
Semi-uniform records (event logs)Expanded list with indentation; limited tables15.0%, but 19.9% larger than compact JSONCompact JSON
Single scalar object (few fields, no array)YAML-style key-value pairsLow; prompt tax can exceed savingsJSON (universally understood)

Source: token reduction figures from the official toonformat.dev benchmark (v4.1). The "better alternative" column reflects the "When Not to Use TOON" guidance in the TOON GitHub repository and the arXiv paper 2603.03306.

When Does JSON Beat TOON on Nested Data?

Three conditions push the balance toward JSON. First, highly non-uniform data: if the objects in your array do not share a consistent field set (think a heterogeneous event log where each event type has different keys), TOON cannot form a tabular block and falls back to an expanded list. The event-log dataset shows what happens: 15.0% smaller than formatted JSON, but 19.9% larger than compact JSON.

Second, small payloads. A February 2026 arXiv study (2603.03306) on TOON vs JSON confirmed a scaling threshold: below roughly 10 objects, the token cost of format instructions (explaining TOON's header syntax to an unfamiliar model) can exceed the per-row savings TOON provides. For a three-item nested cart, stick with JSON.

Third, very deep, non-repetitive nesting. A six-level-deep configuration tree where every sub-key is unique offers TOON nothing to tabularize. The benchmark's deep-config dataset is 6.7% larger in TOON than in compact JSON. YAML also handles this naturally and requires no format instructions because every current LLM has seen extensive YAML in training. The YAML vs TOON comparison covers the crossover point. For more cases like these, see when not to use TOON.

TOON Specification: How Nesting Is Defined Formally

The TOON specification (version 4.1) defines a few structural primitives. A tabular array opens with a header like name[count]{field1,field2}: followed by indented delimited rows. A key-value pair uses key: value, as in YAML. When a value is an object, its keys are indented one level and follow the same rules recursively. When a value is a uniform array, it becomes a tabular array at the next indentation level. Anything else becomes an expanded list with one - item per element.

One rule matters most for nested data: table rows contain only primitive cells. A row can never hold another table. If items in an array contain a sub-array, the outer array becomes an expanded list, and each item's sub-array becomes its own table:

orders[2]:
  - id: ORD-1
    customer:
      name: Ada
      city: London
    items[2]{sku,qty}:
      A1,2
      B2,1
  - id: ORD-2
    customer:
      name: Bob
      city: Austin
    items[1]{sku,qty}:
      C3,5

This is the shape of the benchmark's e-commerce data, and it is why that dataset saves 32.9% rather than 60%. If the orders had no items array, the encoder would fold customer into a nested field group and keep the whole array tabular: orders[2]{id,customer{name,city}}:.

For the full syntax grammar and edge cases (null values, empty arrays, quoting), see the official SPEC.md and the toonformat.dev documentation.

Practical Guidance: Choosing a Format by Data Shape

The benchmark numbers map directly onto a decision rule. If your data contains at least one uniform array of objects with ten or more rows, TOON will save a meaningful number of tokens, likely 30% or more versus formatted JSON. If the top-level structure is a single object with a handful of scalar fields and one small nested array, the savings are modest and may not justify the integration cost.

A useful diagnostic: count the number of times the same key repeats across your payload. In JSON, every key in every object costs tokens. If the word customer_id appears 500 times in a response, that alone is several hundred tokens of pure overhead. TOON collapses that to a single declaration in the header. The more repetitive your keys, the steeper the savings curve. If repetition is hidden under a sub-array, consider flattening the data in your application before encoding so that more of it fits into tables.

For HTML tables scraped from web pages, which are typically uniform and flat, see the companion post on converting HTML tables to TOON for LLM extraction. For a format-agnostic comparison covering JSON, TOON, YAML, CSV, and TONL, the TOON format comparison guide is the most complete reference.

If your nested data also requires schema validation, streaming, or complex queries, consider TONL (Token-Optimized Notation Language), which adds a query API, type hints and streaming on top of similar token savings.

Frequently Asked Questions

How does TOON handle nested or non-uniform data?

TOON uses YAML-style indentation for nested objects and CSV-style tabular rows for uniform arrays of objects. Since spec 4.0, uniform nested objects can stay inside a table through nested field groups. Nested data still saves fewer tokens than flat data: e-commerce orders save 32.9% versus JSON and event logs only 15.0%. For highly non-uniform or deeply nested data, compact JSON or YAML may be a better choice.

What data shapes does TOON save the most tokens on?

In the current official toonformat.dev benchmark (v4.1), TOON saves the most on contacts with uniform nested objects (66.5% versus JSON), employee tables (60.7%), time-series (59.0%) and keyed feature flags (54.6%). Savings drop to 41.7% on GitHub repo data, 32.9% on e-commerce orders and 15.0% on event logs. The mixed-structure track totals 32.7%.

When does JSON beat TOON on nested data?

JSON wins on highly non-uniform data where objects in the same array have different field sets, and on payloads with fewer than roughly 10 objects where TOON's format-instruction overhead outweighs per-row savings. A 2026 arXiv paper (2603.03306) documented this scaling threshold as the prompt tax. On the official benchmark, compact JSON is already smaller than TOON for event logs and deep configuration.

Does TOON handle arrays of objects with a nested sub-array?

Yes, but the outer array is not tabular in that case. Table rows can only hold primitive cells, so an array whose items contain a sub-array is written as an expanded list with one dash per item, and each item's uniform sub-array becomes its own nested table. Nested objects with the same primitive fields in every row can stay tabular through nested field groups.

When should I use YAML instead of TOON for nested data?

Use YAML or compact JSON over TOON for very deep, non-repetitive nesting where there is no uniform array to exploit, such as a configuration tree or a document with many unique sub-keys. YAML's indentation syntax is familiar to all models and requires no format instructions, while TOON's header-plus-rows syntax offers little benefit when rows do not share a schema.

Recommended Reading

TOONNested DataData FormatSpecificationToken EfficiencyLLM