
Constrained decoding is the technique that powers every modern "structured output" mode, but the internal mechanics differ sharply between libraries. Outlines, an open-source structured generation library, takes a notably explicit route: it compiles a schema into a finite state machine and deterministically rebuilds that FSM over tokenizer tokens. This article traces that full pipeline for engineers who already use structured output and want to understand, debug, or extend what happens between schema definition and emitted JSON. It also examines the formal-language-theory ceiling that determines which constraints a system based on regular expressions can actually enforce. The focus is on the JSON Schema path of Outlines; the library's context-free grammar path shares the same scaffolding but operates at a different level of the Chomsky-Schützenberger hierarchy.

outlines.generate.json and StructuredGeneratorWhen you call outlines.generate.json(model, Character) or instantiate outlines.StructuredGenerator, the library does not inspect your Pydantic class directly. The first concrete hop is a delegation to Pydantic's own serializer: Character.model_json_schema(). This produces a JSON Schema document — a declarative specification built from $defs, properties, required, enum, and type keywords — that fully describes the data you want the model to emit.
Consider the Character example used in the Outlines blog:
from enum import Enum
from pydantic import BaseModel
class Name(str, Enum):
john = "John"
paul = "Paul"
class Age(int, Enum):
twenty = 20
thirty = 30
class Character(BaseModel):
name: Name
age: Age
Calling Character.model_json_schema() returns:
{
"$defs": {
"Age": {"enum": [20, 30], "title": "Age", "type": "integer"},
"Name": {"enum": ["John", "Paul"], "title": "Name", "type": "string"}
},
"properties": {
"name": {"$ref": "#/$defs/Name"},
"age": {"$ref": "#/$defs/Age"}
},
"required": ["name", "age"],
"title": "Character",
"type": "object"
}
Several JSON Schema features appear here: nested $defs for reused enums, $ref pointers resolving back into $defs, an explicit required list, enum membership for allowed string and integer values, and type discriminators (object, string, integer).
Outlines does not interpret JSON Schema generically. Instead, it transforms the schema object into one composite regular expression that, if matched, guarantees the output is both valid against the schema and round-trippable through Pydantic. The documented entry point is outlines.fsm.json_schema.build_regex_from_object, which takes the JSON-serialized schema:
import outlines.fsm as fsm
import json
regex_str = fsm.json_schema.build_regex_from_object(json.dumps(json_schema))
# '{"name":("John"|"Paul"),"age":(20|30)}'
If the model emits any string matching this regex, the result is guaranteed schema-conformant. The downstream stages — regex-to-FSM compilation via the interegular library and per-token state tracking — operate exclusively on this regex.
Because the converter only understands a documented subset of JSON Schema keywords, anything outside that subset either raises an error or, in some cases, silently drops. Common offenders include certain oneOf constructs, unbounded numeric ranges (minimum/maximum without a co-occurring enum), and custom string format validators. This is the most frequent debugging trap: a schema looks valid to Pydantic, yet the regex conversion fails or weakens the constraint. The Outlines repository and the tracking issue #215 describe the current support matrix and are the most reliable reference when a schema element refuses to round-trip.

The second stage of the Outlines pipeline is a single, mostly string-emitting function: outlines.fsm.json_schema.build_regex_from_object. It walks the JSON Schema dictionary produced by model.model_json_schema() and returns one Python string that, once parsed, defines the entire space of legal completions.
The canonical Character model has two fields, name (enum: "John", "Paul") and age (enum: 20, 30). After Pydantic serialization the schema is a small dictionary whose properties only reference $defs blocks for Name and Age. Calling:
fsm.json_schema.build_regex_from_object(json.dumps(json_schema))
returns the composite regular expression:
\{("name":("John"|"Paul"),"age":(20|30)\}
Every character of the example output — the escaped braces, the quoted keys, the colons, the comma, and the value groups — is produced by the walker (blog.dottxt.ai — Coalescence).
The walker applies a deterministic rule per JSON Schema node:
type: object emits the literal { and }, and for each property in properties (restricted to required) it emits a quoted key, a :, and the regex of the value type, joined by ,.enum (for strings or integers) becomes a parenthesized alternation. Each member is emitted in its natural form — "John" for a string, 20 for an integer — separated by |.type: integer compiles to a digit class such as [0-9]+, so any length of digits is permitted rather than only the enumerated values.type: array becomes a parenthesized group containing the element regex, optionally followed by (, ...)* so multiple elements separated by , are allowed.$ref pointers (e.g. #/$defs/Name) are resolved and inlined at the call location, flattening what looks like nested definitions into a flat regular expression. This is why the Character output has no group naming or back-references — the $defs blocks have been substituted in place.Any deviation from the schema flavor Outlines recognizes falls into the same string emission loop: an unknown keyword is either passed through as a regex literal (with regex metacharacters escaped) or, if it cannot be expressed as one, surfaces as an Unsupported schema feature error at construction time.
Because the regex is read once at generator instantiation and never reinterpreted, schema validation is effectively a one-shot transformation. There is no runtime reflection on the Pydantic model — the FSM the next stage builds operates on the regex, not on the schema.
The same property is what enables Outlines' near-zero per-token overhead. Building the regex and compiling the FSM is a fixed cost paid per generator call; per-token work reduces to a lookup of allowed transitions in the current FSM state (blog.dottxt.ai — Coalescence). That is the empirical basis for the claim that structured generation in Outlines is "free" relative to vanilla sampling: the construction cost is amortized across all tokens of a completion.

The pipeline now holds a regular expression string describing the JSON Schema's surface form. Stage 3 is the step that turns that string into something a generation loop can walk through one symbol at a time: a finite automaton.
The pivot is a textbook equivalence from automata theory. Regular expressions and finite state machines recognize exactly the same class of languages — the regular languages. Given any regex, there exists an FSM whose accepted strings are precisely the strings matched by the regex, and vice versa. This is the lever Outlines pulls to make validity "free": if generation is constrained to follow transitions of an FSM built from the schema's regex, every emitted string is in the language the regex defines, and therefore in the language the schema accepts.
Outlines performs this step by delegating to the third-party interegular library:
import interegular
fsm = interegular.parse_pattern(regex_str).to_fsm()
parse_pattern parses the regex into an internal representation; .to_fsm() materializes a nondeterministic finite automaton over the Unicode alphabet of single characters. The result is a labelled transition system whose accepting set exactly matches the regex's language.
The runtime loop over a character FSM is small and worth reading end-to-end:
Validity is free in this loop because of the equivalence above. Any path from the start state to an accepting state, traced one labelled character at a time, spells out a string in the language of the regex. Conversely, anything not in the language cannot reach an accepting state, so the loop physically cannot produce an invalid string.
The FSM that to_fsm() returns is generally nondeterministic: a single (state, character) pair may have multiple outgoing transitions, including ε-transitions left over from regex compilation. A loop that needs to answer "which characters are legal here?" against an NFA must keep a set of simultaneously active states and compute the union of their outgoing labels — a parallel simulation that grows with the NFA's ambiguity.
Outlines determinizes the result with outlines.fsm.regex.make_deterministic_fsm, which performs classic subset-state (Rabin–Scott) determinization. Each DFA state corresponds to a set of NFA states; transitions are lifted from the NFA by taking the union of all transitions out of the active set on each input character. The result is an equivalent automaton in which every (state, character) pair has at most one successor, so the generation loop reduces to a single dictionary lookup per character rather than a set simulation.
That determinization is what makes Stage 4's token-level rebuild possible at all: the next stage re-derives a token-indexed transition table by composing the DFA with the tokenizer, and it does so cheaply precisely because the underlying character FSM is deterministic.
Source: Coalescence in LLM Structured Generation (dottxt).

The character-level FSM that interegular produces cannot be used directly to drive a BPE or SentencePiece tokenizer, because the model does not emit characters — it emits token IDs. Outlines bridges that gap with a deterministic compilation step implemented in outlines.fsm.regex:
from outlines.fsm.regex import make_deterministic_fsm, create_fsm_index_tokenizer
new_fsm, _ = make_deterministic_fsm(fsm)
index, _ = create_fsm_index_tokenizer(new_fsm, tokenizer)
create_fsm_index_tokenizer walks every transition of the deterministic character FSM and asks the tokenizer, for each state, which vocabulary tokens can be legally consumed from that state. The result is the index — a Python dictionary whose keys are FSM state integers and whose values are inner dictionaries mapping each allowed token (or token ID) to the resulting next state:
index = {
0: {'{"': 2, '{': 1}, # tokens that start a JSON object
2: {'"name': 6, '"id': 6},
6: {'":': 7, '" :': 7},
# ...
}
The generation loop then becomes a tight, four-step procedure at every decoding step:
0. Look up index[0] to obtain the set of token IDs currently valid from this state.index[0], then re-normalize and sample as usual.new_state = index[current_state][sampled_token]. Append the token's text to the running output and repeat.The cost model is the key to understanding why constrained decoding in Outlines has roughly the same throughput as unconstrained generation: every emitted token still requires exactly one full LLM forward pass. The FSM does not skip the model, does not merge tokens, and does not bypass the attention computation — it only restricts the sampling distribution. The only per-step overhead is an integer-keyed dictionary lookup and a vocabulary-sized logit mask, both of which are negligible relative to a transformer forward pass.
This accounting has a direct architectural consequence. Because masking happens on the raw logits before sampling, Outlines requires white-box access to the model's logits. Closed inference endpoints such as OpenAI's hosted completions, Anthropic's Messages API, and most managed cloud LLM services do not expose logits to the caller, so Outlines cannot enforce a JSON schema against them. The library is therefore restricted to self-hosted open-weight models — typically loaded through Hugging Face transformers, vLLM, or an equivalent runtime that hands the logits back to user code (blog.dottxt.ai).

A character-level FSM has a unique next state for every emitted string, but a token-level FSM does not. BPE and SentencePiece vocabularies are built so that, whenever a multi-character token such as "name" is present, the individual letters n, a, m, e and the prefixes na, nam, ame, me are also present in the vocabulary. That redundancy means the string name can be produced by at least eight distinct token sequences: ["name"], ["n","a","m","e"], ["na","m","e"], ["nam","e"], ["n","am","e"], ["n","ame"], ["na","me"], and ["n","a","me"] (Outlines: Coalescence in LLM Structured Generation).
Each of those sequences arrives at the same character state, so the FSM is happy with all of them and the JSON validator cannot tell them apart. The model, however, can. After emitting ["name"] the conditional distribution over the next token is concentrated on continuations that follow a single whole-word; after ["n", "a", "m", "e"] the distribution is one that has just produced four separate letter tokens, which shifts the probability mass in measurable ways. In the running example of choosing between John and Paul, that shift can flip the answer, so two structurally identical JSON strings — {"name":"John",...} and {"name":"John",...} — can be generated under materially different distributions. Structured decoding guarantees syntactic validity; it does not guarantee distributional identity.
This is precisely the observation Outlines formalizes as coalescence. The rule is simple: when several valid tokens at a state share the same prefix, keep only the transition for the longest token. In a small JSON object with nine content tokens, that rule collapses the transition table so that the model is only called twice, while the remaining seven tokens are appended directly, yielding the roughly5× speedup the coalescence blog reports. The emphasis there is deliberate: coalescence is a speed optimization, not a correctness one.
That distinction leaves an open research question. When the longest-token path and a shorter-token path assign noticeably different probabilities to the next continuation, which should the library prefer? Temperature, top-p, and even the prompt's trailing characters shift which token the model wants to emit first, so the answer is not fixed by the schema. Coalescence, as currently implemented in Outlines, defaults to the longest path; whether that is the path that best matches the model's intent under realistic sampling regimes is an empirical question the structured-generation community has not yet settled.

Outlines' JSON Schema path is anchored at type-3 of the Chomsky-Schützenberger hierarchy. The pipeline — Pydantic model → JSON Schema → regex string → FSM via the interegular library — only makes sense because every regular expression recognizes a regular language and every regular language is accepted by a finite state machine. Because the schema is converted to a regex first and only then to a state machine, anything the FSM can guarantee is, by construction, something a regular expression could have matched at the character level.
For the schemas Outlines compiles in practice, the situation is even narrower. The full space of all valid JSON documents is context-free (the matching of {/} and [/] across nested arrays and objects requires a stack), but a Pydantic- or OpenAPI-style schema with a fixed, finite set of keys and bounded depth can be unrolled into a finite set of type-3 patterns. The result is a language that is type-1 (context-sensitive) on the hierarchy but trivially so, because the key universe and depth bounds are known at compile time — the observation that lets Guardrails AI drop back to plain greedy decoding on top of the FSM (Guardrails AI, 2023).
The cost of staying at type-3 is the absence of any unbounded memory. A finite automaton cannot count to an arbitrary depth, so it cannot enforce:
[[[[…]]]] of arbitrary length — push the regex path past its limits. A fixed-depth bound works; an open-ended one does not.<b><i>x</b></i> cannot be distinguished from <b><i>x</i></b> without a stack.When a schema is "essentially a regular expression written in JSON" — flat objects, enums, bounded nesting, primitive types — Outlines' JSON path is the right tool, and the FSM guarantees well-formed output at essentially no inference overhead (dottxt blog). When the schema requires unbounded recursion or matching pairs, three practical escapes exist:
CFG path, which operates one level up the hierarchy (type-2, context-free) and supports grammars like SQL.xgrammar and accepts either regex or CFG targets against the same FSM/grammar machinery.The takeaway for engineers is that the schema's structure, not its size, determines whether regex-constrained decoding is sufficient. If the validation rules can be written down as a finite set of patterns over a fixed alphabet, Outlines' JSON path will enforce them exactly. If the rules demand a stack, expect to either bound the input or move up the hierarchy.

Before touching the FSM, dump the schema Outlines will actually receive:
import json
print(json.dumps(Character.model_json_schema(), indent=2))
This is the single most common failure surface. Pydantic field aliases, custom Field(alias=...), model_config = ConfigDict(populate_by_name=True), field_validator transformations, and Annotated metadata all flow into the generated schema. A field that "exists" on the Python class can vanish from the schema, become optional when you expected it required, or appear under a different key. If the schema looks wrong here, no amount of FSM debugging downstream will recover valid output.
Pass the dumped schema to the regex builder and read what it returns:
import outlines.fsm.json_schema as jsfsm
regex_str = jsfsm.build_regex_from_object(json.dumps(Character.model_json_schema()))
print(regex_str)
You are looking for two pathologies. First, alternation explosion: an enum with hundreds of string values, or a Union of dozens of object types, produces a regex of the form (a|b|c|...) whose size scales linearly with the enumeration. Second, deeply nested optionals and anyOf constructs, which interegular handles but can push the FSM state count into the thousands. If the regex string is megabytes long, expect slow first-call compilation and a tokenizer index that does not fit in cache.
Build the FSM directly and step through transitions to confirm the expected strings are accepted:
from outlines.fsm.regex import make_deterministic_fsm
fsm, _ = make_deterministic_fsm(regex_str)
state = 0
for ch in '{"name":"John","age":20}':
state = fsm.next_state(state, ch)
assert state in fsm.finals
This isolates the bug to either the schema-to-regex conversion, the regex-to-FSM conversion, or a model/tokenizer issue downstream. It also gives you a ground truth against which the token-level index can be cross-checked.
A healthy constrained generator at state 0 admits only a small set of tokens (typically the few whose decoded text starts with { or, after the opening brace, a " for the first key). Build the tokenizer index and compare:
from outlines.fsm.regex import create_fsm_index_tokenizer
index, _ = create_fsm_index_tokenizer(fsm, tokenizer)
allowed_at_start = set(index[0].keys())
print(len(allowed_at_start), "tokens admitted at start; vocab size:",
len(tokenizer.get_vocab()))
If the admitted set is comparable to the full vocabulary, the index was built against the wrong FSM state, the regex was accepted by interegular but is not actually constraining (e.g., .*), or the tokenizer passed in does not match the model being used.
Wrap the generator loop to record, for every emitted token, the FSM state it left and the set of tokens that would have been legal next:
state = 0
for token in generator:
legal_next = set(index[state].keys())
top_k = sample_top_k(logits, k=5)
if not any(t in legal_next for t in top_k):
logger.warning("FSM forced a low-probability token at state %d", state)
state = index[state][token]
Two failure modes are now distinguishable. If the model repeatedly requests tokens that fall outside legal_next, the logit mask is functioning correctly but the model is out-of-distribution for the schema, so the fix is at the prompt or model layer. If a state admits exactly one legal token and the model assigns it low probability, you are observing a forced-but-undesired completion; relax the schema or warm up the model on similar structures.