Structured Output
LLMs are trained to produce natural language, but software speaks in types, fields, and constraints. Structured output bridges that gap: instead of parsing free text and hoping it's valid, you tell the model exactly what shape you need, and the model fills it in.
Three approaches, one goal: turn an LLM's natural-language output into a predictable, validated data structure your application can act on immediately.
"Please return JSON" works until it doesn't.
LLMs output text. Apps need data: JSON objects with typed fields, enums, dates in ISO format, arrays of known shapes. The naive approach: prompt the model with "return JSON" and it works about 80% of the time. The other 20% introduces errors that break production: trailing commas, unclosed brackets, hallucinated field names, markdown code fences wrapping the JSON, or extra prose like "Here you go!" mixed in.
Those 20% failures aren't random noise. They spike under exactly the conditions where structured output matters most: long responses, deeply nested objects, edge cases the model is uncertain about. When your payment processing pipeline depends on a valid JSON schema, 80% reliability isn't close to enough.
Three approaches, each tighter than the last.
JSON Mode
You instruct the model to output JSON, the API enforces that tokens form valid JSON syntax, and you parse the result. If parsing fails, you retry with an error message. Simplest to set up, least reliable. The model can still hallucinate field names, omit required fields, or produce values of the wrong type. Good enough for low-stakes internal tools.
Function Calling
You declare a typed function schema with named parameters, types, and descriptions. The API tells the model it must call this function with valid arguments. Because the schema is part of the API contract rather than a prompt instruction, enforcement moves from soft guidance to hard constraint. The model still makes mistakes, but far fewer.
Constrained Generation
Instead of hoping the model produces valid JSON and catching errors after the fact, the token sampler itself is restricted: it can only emit tokens that stay within the valid JSON grammar for your schema. Every single token is guaranteed to be syntactically and structurally valid. Zero parse errors. This is the gold standard, available via libraries like Outlines and Guidance.
Define once, validate everywhere.
Pydantic
Define your schema as a Python class with typed fields, validators, and defaults. Pydantic handles serialization, deserialization, and validation in one place. Define it once and get JSON Schema, strict validation, and clear error messages for free. The standard for typed data in Python.
Instructor
Wraps OpenAI, Anthropic, and other API clients. You pass a Pydantic model as the response type and Instructor handles schema extraction, API call with function-calling mode, and parsing. Returns a fully typed and validated Pydantic instance directly. No manual parsing step needed.
Zod
The TypeScript equivalent of Pydantic. Define schemas with chained methods, infer TypeScript types directly from the schema definition, and validate runtime data. Pairs with libraries like Vercel AI SDK to get typed responses from LLM calls with automatic parsing and validation.
Outlines / Guidance
Constrained generation libraries that modify the token sampling process itself. Define your schema as a grammar, regex, or Pydantic model, and the library guarantees every generated token is valid. Zero parse failures by construction. More setup overhead than JSON mode or function calling, but the reliability gain is absolute.
The pattern is the same across every stack: define the shape, constrain the generation, validate the result. The difference is only where the constraint lives: in the prompt, in the API, or in the token sampler.
Even structured output breaks. Here's how.
Hallucinated Fields
The model invents fields you didn't ask for. Your schema says {name, age} and the model returns {name, age, occupation, favorite_color}. JSON mode won't catch this because the output is syntactically valid. Only post-parse validation or constrained generation prevents it.
Schema Drift
You add a required field to your Pydantic model, deploy the change, but your prompt still describes the old schema. The field is now required by validation but the model has no instruction to provide it. Parse errors spike and no one knows why. Keep prompts and schemas in lockstep.
Enum Case Sensitivity
Your schema says enum: ["Low", "Medium", "High"]. The model returns "medium" or "MEDIUM" instead of "Medium". JSON is valid, the value is semantically correct, but your parser rejects it. Normalize enum values or use case-insensitive matching in validation.
Deep Nesting
A schema with objects inside objects inside arrays inside objects is exponentially harder for models to produce correctly. Every nesting level multiplies the chance of a structural error. Flatten where possible: prefer two levels of nesting with references over four levels of inline nesting.
Five practices that separate prototype from production.
Structured output isn't a feature you bolt on at the end. It's a pipeline: schema design, generation, parsing, validation, retry. Each stage can fail. Production systems handle every stage explicitly.
"Natural language is for humans. Structured output is how software talks to software."
Structured output transforms LLMs from conversation partners into software components. Without it, every integration is a fragile parsing layer waiting to break.
The reliability gap between JSON mode (~80%) and constrained generation (~99.9%) is the difference between a demo and a production payment pipeline.