โ† Back to ConceptsVerified: 2026-07-30
Structured OutputWhen Text Isn't Enough

Structured Output

LLMs are trained to produce natural language, but software speaks in types, fields, and constraints. Structured output bridges that gap: instead of parsing free text and hoping it's valid, you tell the model exactly what shape you need, and the model fills it in.

LLM OUTPUTSure! Here's thedata you asked for:{ "name": "Alice", "age": 34, "email": "alice@, "role": "admin"}Hope that helps!Mixed with prose,broken valuesJSON MODE~80% reliableFUNCTION CALLING~95% reliableCONSTRAINED GEN~99.9% reliableTYPEDOBJECTValid, validated,ready to useReliability increases with tighter schema enforcement

Three approaches, one goal: turn an LLM's natural-language output into a predictable, validated data structure your application can act on immediately.

01The Problem

"Please return JSON" works until it doesn't.

LLMs output text. Apps need data: JSON objects with typed fields, enums, dates in ISO format, arrays of known shapes. The naive approach: prompt the model with "return JSON" and it works about 80% of the time. The other 20% introduces errors that break production: trailing commas, unclosed brackets, hallucinated field names, markdown code fences wrapping the JSON, or extra prose like "Here you go!" mixed in.

Those 20% failures aren't random noise. They spike under exactly the conditions where structured output matters most: long responses, deeply nested objects, edge cases the model is uncertain about. When your payment processing pipeline depends on a valid JSON schema, 80% reliability isn't close to enough.

02How It Works

Three approaches, each tighter than the last.

Tier 1

JSON Mode

You instruct the model to output JSON, the API enforces that tokens form valid JSON syntax, and you parse the result. If parsing fails, you retry with an error message. Simplest to set up, least reliable. The model can still hallucinate field names, omit required fields, or produce values of the wrong type. Good enough for low-stakes internal tools.

Tier 2

Function Calling

You declare a typed function schema with named parameters, types, and descriptions. The API tells the model it must call this function with valid arguments. Because the schema is part of the API contract rather than a prompt instruction, enforcement moves from soft guidance to hard constraint. The model still makes mistakes, but far fewer.

Tier 3

Constrained Generation

Instead of hoping the model produces valid JSON and catching errors after the fact, the token sampler itself is restricted: it can only emit tokens that stay within the valid JSON grammar for your schema. Every single token is guaranteed to be syntactically and structurally valid. Zero parse errors. This is the gold standard, available via libraries like Outlines and Guidance.

03The Tool Stack

Define once, validate everywhere.

Python

Pydantic

Define your schema as a Python class with typed fields, validators, and defaults. Pydantic handles serialization, deserialization, and validation in one place. Define it once and get JSON Schema, strict validation, and clear error messages for free. The standard for typed data in Python.

Python

Instructor

Wraps OpenAI, Anthropic, and other API clients. You pass a Pydantic model as the response type and Instructor handles schema extraction, API call with function-calling mode, and parsing. Returns a fully typed and validated Pydantic instance directly. No manual parsing step needed.

TypeScript

Zod

The TypeScript equivalent of Pydantic. Define schemas with chained methods, infer TypeScript types directly from the schema definition, and validate runtime data. Pairs with libraries like Vercel AI SDK to get typed responses from LLM calls with automatic parsing and validation.

Advanced

Outlines / Guidance

Constrained generation libraries that modify the token sampling process itself. Define your schema as a grammar, regex, or Pydantic model, and the library guarantees every generated token is valid. Zero parse failures by construction. More setup overhead than JSON mode or function calling, but the reliability gain is absolute.

Define schema (Pydantic / Zod)โ†’Call model with schema as constraintโ†’Receive typed, validated objectโ†’Use it

The pattern is the same across every stack: define the shape, constrain the generation, validate the result. The difference is only where the constraint lives: in the prompt, in the API, or in the token sampler.

04When It Fails

Even structured output breaks. Here's how.

Pitfall 1

Hallucinated Fields

The model invents fields you didn't ask for. Your schema says {name, age} and the model returns {name, age, occupation, favorite_color}. JSON mode won't catch this because the output is syntactically valid. Only post-parse validation or constrained generation prevents it.

Pitfall 2

Schema Drift

You add a required field to your Pydantic model, deploy the change, but your prompt still describes the old schema. The field is now required by validation but the model has no instruction to provide it. Parse errors spike and no one knows why. Keep prompts and schemas in lockstep.

Pitfall 3

Enum Case Sensitivity

Your schema says enum: ["Low", "Medium", "High"]. The model returns "medium" or "MEDIUM" instead of "Medium". JSON is valid, the value is semantically correct, but your parser rejects it. Normalize enum values or use case-insensitive matching in validation.

Pitfall 4

Deep Nesting

A schema with objects inside objects inside arrays inside objects is exponentially harder for models to produce correctly. Every nesting level multiplies the chance of a structural error. Flatten where possible: prefer two levels of nesting with references over four levels of inline nesting.

05Production Patterns

Five practices that separate prototype from production.

1. Always retry on parse failureGive the model the exact parse error and let it self-correct. Works ~90% of the time on retry 1. If it fails twice, fail loudly.
2. Use enums, not free-text, for categories"sentiment: positive | negative | neutral" is a closed set the model can map to reliably. "sentiment: describe the tone" invites drift.
3. Validate post-parseThe model said it's a date. Is it actually parseable as a date? Is the age a positive integer? Schema validation catches structural errors; business-logic validation catches semantic ones.
4. Keep schemas flatTwo levels of nesting max. Deeply nested JSON is exponentially harder for models. Use top-level arrays of flat objects with reference IDs over nested trees.
5. Test with adversarial inputsFeed your pipeline edge cases: empty strings, nulls, extremely long values, Unicode, markdown injection. The model will surprise you. Find the surprises before your users do.

Structured output isn't a feature you bolt on at the end. It's a pipeline: schema design, generation, parsing, validation, retry. Each stage can fail. Production systems handle every stage explicitly.

"Natural language is for humans. Structured output is how software talks to software."

01

Structured output transforms LLMs from conversation partners into software components. Without it, every integration is a fragile parsing layer waiting to break.

02

The reliability gap between JSON mode (~80%) and constrained generation (~99.9%) is the difference between a demo and a production payment pipeline.