Validating TOON Data: Schemas, Round-Trips, and Lint Checks
TOON's header declares row count and fields, giving models and tools a built-in checksum. Learn practical ways to validate TOON, from round-trips to TONL schema validation.
Validate TOON data by checking that the header's declared row count matches the actual rows, running a JSON-to-TOON-to-JSON round-trip to confirm lossless conversion, and — for production pipelines — adding TONL schema validation for type-level guarantees. Do not rely on the LLM alone: structural validation accuracy in the official benchmark was only 70.0%.
Why TOON Has a Built-In Validation Hook
Most data formats carry no structural metadata in the payload itself. A JSON array of 200 objects gives a parser no way to know whether a truncated response delivered 180 or 200 rows — you have to count. TOON's design includes a deliberate fix for this: the header line declares both the expected count and the field schema before any data appears.
The header syntax looks like this:
users[3]{id,name,role}:
1, Alice, admin
2, Bob, viewer
3, Carol, editorThe users[3] segment declares the array name and expected row count. The {id,name,role} segment declares the ordered field list. Both are checkable without parsing the full payload. According to the TOON reference implementation, this header gives the LLM — and your tooling — an explicit schema and count to validate against, reducing ambiguity about what a well-formed payload looks like.
For a deeper explanation of TOON's syntax and design rationale, see What is TOON?
Why You Cannot Rely Solely on the LLM for Validation
The TOON header gives models a reference point, but model-side structural validation is less reliable than it might appear. The official toonformat.dev benchmarks — 5,016 LLM calls across 209 questions, six formats, and four models — measured structural validation accuracy at 70.0% for TOON. That sounds reasonable until you compare it to field retrieval accuracy, which sat at 99.6%.
The 29.6-percentage-point gap between those two numbers is the practical case for programmatic validation. Models are highly reliable at answering "what is the value of field X in row Y?" but meaningfully less reliable at answering "does this block have the right number of rows and the correct set of fields?" A missing row or a silently dropped column is exactly the kind of error that slips past a model and causes downstream bugs.
Add programmatic checks. The LLM is a consumer of your TOON data, not a validator of it.
Validation Layer 1: Header-Count Check
The cheapest validation you can do is a header-count check: parse the declared count from the header line and compare it to the number of non-empty data rows that follow. A mismatch means the data was truncated, a row was duplicated, or the header was edited without updating the count.
// JavaScript — minimal header-count validator
function validateToonBlock(toon) {
const lines = toon.trim().split('
');
const headerMatch = lines[0].match(/[(d+)]/);
if (!headerMatch) throw new Error('No row count in TOON header');
const declaredCount = parseInt(headerMatch[1], 10);
// Data rows start after the header line; skip blank lines
const dataRows = lines.slice(1).filter(l => l.trim().length > 0);
if (dataRows.length !== declaredCount) {
throw new Error(
`Row count mismatch: header declares ${declaredCount}, found ${dataRows.length}`
);
}
return true;
}
// Example usage
const toon = `users[3]{id,name,role}:
1, Alice, admin
2, Bob, viewer
3, Carol, editor`;
validateToonBlock(toon); // passes
This check is fast, has no dependencies, and catches the most common failure mode: a response that was truncated mid-stream. It does not verify field names or value types — that requires the next two layers.
You can also validate the field list by parsing the header's {field1,field2,...} segment and comparing it against your expected schema. Any added, removed, or reordered field name is a signal that the source data changed shape.
Validation Layer 2: JSON-to-TOON-to-JSON Round-Trip
A round-trip test is the most thorough quick check available: convert your source JSON to TOON, then convert the TOON back to JSON, and perform a deep equality comparison between the original and the reconstructed object.
// Pseudocode — round-trip validation pattern
const originalJson = [
{ id: 1, name: "Alice", role: "admin" },
{ id: 2, name: "Bob", role: "viewer" },
{ id: 3, name: "Carol", role: "editor" },
];
// Step 1: convert to TOON (e.g. via @toon-format/toon package or json2toon.co)
const toon = toToon(originalJson);
// => users[3]{id,name,role}:
// 1, Alice, admin
// 2, Bob, viewer
// 3, Carol, editor
// Step 2: convert back to JSON
const reconstructed = fromToon(toon);
// Step 3: deep-equal check
const isValid = JSON.stringify(originalJson) === JSON.stringify(reconstructed);
if (!isValid) {
console.error('Round-trip failed — data loss or type coercion detected');
}Any field that was silently dropped, any value that was coerced from a number to a string, and any row that went missing will surface as a diff between the two serializations. Run this check whenever you change the source data schema, update the TOON library version, or modify your serialization configuration (delimiter, indentation, or quoting rules). The free converter makes it easy to run an ad-hoc round-trip manually.
Round-trip testing is also covered in the TOON specification as the canonical way to confirm format conformance.
Validation Layer 3: TONL Schema Validation
For production pipelines that need type-level guarantees — not just structural ones — plain TOON's header is not enough. TOON declares field names and counts; it does not declare types. A port field could silently hold a string instead of an integer and the header-count check would not catch it.
TONL (Token-Optimized Notation Language) extends the TOON concept with explicit type hints and a full schema validation layer. A TONL schema block looks like this:
// TONL with type hints — schema is machine-readable and auto-generates TypeScript
servers[3]{host:str, port:u32, region:str}:
api-1.example.com, 443, us-east
api-2.example.com, 443, eu-west
api-3.example.com, 443, ap-southThe type hints (str, u32, bool) add roughly 20 tokens to the block, but they unlock schema validation and auto TypeScript generation. Critically, the token cost remains competitive: TONL is still approximately 32% smaller than JSON even with type hints included, according to tonl.dev. For production data pipelines, that is a strong trade-off — you get type safety and save tokens simultaneously.
TONL also ships with 2,300+ passing tests and zero runtime dependencies, which makes it suitable for validation in constrained environments. For a full comparison of TOON and TONL capabilities, see TOON vs TONL.
Validation Approach Comparison
| Validation approach | What it catches | What it misses | Cost |
|---|---|---|---|
| Header row-count check | Truncation, duplicate rows, count drift after schema edits | Type errors, wrong field names, value corruption | Near-zero — single regex + integer comparison |
| Header field-list check | Added, removed, or reordered fields vs expected schema | Type errors, value-level corruption | Near-zero — string comparison of field names |
| JSON-to-TOON-to-JSON round-trip | Data loss, value coercion, row reordering, serialization bugs | Schema drift in live data (only validates a snapshot) | Low — two conversions + deep-equal; runs in milliseconds |
| TONL schema validation | Type errors, constraint violations, schema drift, all structural issues | Business-logic rules not expressed in the schema | Low — zero runtime deps; ~20 extra tokens per block for type hints |
| LLM structural check (model reads header) | Gross structural problems in simple cases | Subtler errors — only 70.0% structural validation accuracy in benchmarks | High — consumes tokens and adds latency; unreliable |
The practical recommendation is to stack the first three approaches. Header-count and field-list checks are so cheap that there is no reason to omit them. Round-trip tests belong in your CI pipeline and in any pre-flight check before constructing a prompt. TONL schema validation is the right upgrade when your data has non-trivial types or when a validation failure could cause a downstream production incident.
Validating TOON in a Real Pipeline
A concrete end-to-end pattern for an API pipeline that fetches database rows and passes them to an LLM:
// 1. Fetch rows from your data store
const rows = await db.query('SELECT id, name, role FROM users LIMIT 200');
// 2. Convert to TOON
const toon = toToon(rows, { arrayName: 'users' });
// => users[200]{id,name,role}:
// 1, Alice, admin
// ...
// 3. Header-count validation (always run this)
validateToonBlock(toon); // throws if count mismatch
// 4. Optional: round-trip check in staging / canary
if (process.env.NODE_ENV !== 'production') {
const back = fromToon(toon);
assert.deepStrictEqual(rows, back, 'Round-trip mismatch');
}
// 5. Build the prompt — TOON block goes in the context window
const prompt = `Answer the following question using only the data below.
${toon}
Question: Which users have the admin role?`;
const response = await llm.complete(prompt);Steps 3 and 4 add less than a millisecond of latency. They are not optional in production — the official benchmark's 70.0% structural validation accuracy means roughly one in three structural errors goes undetected if you rely on the model.
For best practices around prompt construction, error handling, and larger payloads, see TOON best practices and troubleshooting.
When to Upgrade to TONL for Stronger Guarantees
Plain TOON with the three validation layers above covers most retrieval and comprehension use cases. Upgrade to TONL when any of the following apply:
- Type safety is a hard requirement. If a
portbeing a string instead of an integer would cause a runtime failure, TONL's type hints and schema validator are the right tool. The validator enforces types programmatically, not probabilistically. - You want TypeScript interfaces generated from your data schema. TONL can auto-generate TypeScript types directly from the TONL schema, keeping your runtime types and your prompt schema in sync.
- Your dataset is large enough to need streaming. TONL supports streaming files larger than 50GB in under 100MB of memory. At that scale, a round-trip check on the full dataset is impractical; TONL's streaming validator handles incremental validation.
- You need a query API alongside validation. TONL ships a SQL-like query API with indexed lookups under 0.1ms. If you are filtering or aggregating before passing data to the LLM, running that logic inside TONL is more reliable than asking the model to do it — recall that TOON's filtering accuracy in the benchmark was only 56.8%.
The token cost difference is small. TONL with type hints is still roughly 32% smaller than JSON. The full capability comparison is in the TOON vs TONL guide.
If you are approaching TOON for the first time and want to understand format trade-offs before adding validation, TOON vs JSON5 and HJSON covers why the format was designed for machines rather than humans — which is the same reason programmatic validation matters more than human review.
Frequently Asked Questions
How do I validate TOON data?
Validate TOON data at three levels: first check that the header row count matches the actual number of data rows; second run a JSON-to-TOON-to-JSON round-trip to confirm lossless conversion; third use TONL schema validation if you need type-level guarantees and auto TypeScript generation. Do not rely solely on the LLM to catch structural errors.
What does the TOON header declare?
The TOON header line declares the array name, the expected row count, and the ordered list of field names — for example, users[3]{id,name,role}:. This gives both tools and LLMs an explicit schema and count to validate against. A mismatch between the declared count and actual rows indicates truncation or a parsing error.
Can an LLM validate TOON data reliably?
Partially. The official toonformat.dev benchmarks (5,016 LLM calls, four models) found structural validation accuracy of only 70.0% for TOON — well below field retrieval accuracy of 99.6%. That gap means you should run programmatic header-count and round-trip checks rather than trusting the model to catch every structural error.
What is a TOON round-trip test?
A round-trip test converts JSON to TOON, then converts the TOON back to JSON, and compares the two JSON objects for equality. Any loss of fields, reordering of values, or type coercion will surface as a diff. This is the fastest way to confirm that your TOON block faithfully represents the original data before using it in a prompt.
When should I use TONL schema validation instead of TOON?
Use TONL schema validation when you need type-level guarantees (u32, str, bool) and want auto-generated TypeScript interfaces from your data schema. TONL delivers roughly 32% fewer tokens than JSON even with type hints, adds schema validation with 2,300+ passing tests, and handles streaming files over 50GB. For simple retrieval tasks, plain TOON with header-count checks is sufficient.
Recommended Reading
How TOON Handles Nested and Non-Uniform Data
TOON shines on uniform arrays, but real data nests. Learn how TOON represents nested objects and mixed structures, where savings drop, and when JSON or YAML wins.
How to Calculate Your LLM Token Savings from TOON
A simple cost model for estimating real savings from TOON: tokens times price, minus the prompt tax, stacked with batch and caching discounts. With a worked example.
Feeding Financial and Market Data to LLMs with TOON
Prices, trades, and time-series are dense uniform tables—TOON's best case, saving up to 59%. Learn how to format financial data for LLM analysis without blowing the token budget.