- Structured outputs make a model return data that matches a schema, usually JSON. Constrained decoding can guarantee the format, but not that the values are correct or allowed.
- Treat model output like untrusted user input: validate it at the point it enters your system, with both a schema check and business-rule checks.
- On failure, retry once or twice with the specific validation error in the prompt. Models fix precise errors reliably; they do not fix a generic “try again”.
- When validation keeps failing, fall back to a safe default or a human, and log the case as a new eval example.
The first time a model returns clean JSON, it feels like the integration problem is solved. Parse it, pass it on. Then a refund amount arrives as a string, a date is in the wrong year, an enum value is one the code has never seen, and a downstream service does something nobody intended.
01What do LLM structured outputs actually guarantee?
Structured outputs ask a model to return data in a fixed shape, usually JSON that matches a schema. Providers offer this in a few forms: function or tool calling, JSON mode, and constrained decoding, where the model can only produce tokens that keep the output valid against the schema.
Constrained decoding is a real improvement. It removes the class of bug where the JSON does not parse. It does not tell you whether:
- the
amount_centsis within the order total, - the
order_idexists, - the chosen
actionis allowed for this customer, - two fields are consistent with each other.
A schema describes shape. Correctness lives in rules only your system knows.
02The decision
03What does the boundary look like?
Define the output as a type with validators. Everything the rest of the system touches has passed through it.
from pydantic import BaseModel, Field, ValidationError, model_validator
from typing import Literal
class TriageResult(BaseModel):
action: Literal["refund", "replace", "escalate", "reply_only"]
order_id: str = Field(pattern=r"^ORD-\d{4,}quot;)
amount_cents: int = Field(ge=0)
reply: str = Field(max_length=600)
@model_validator(mode="after")
def refund_needs_amount(self):
if self.action == "refund" and self.amount_cents == 0:
raise ValueError("action is refund but amount_cents is 0")
return self
def check_rules(r: TriageResult, order: Order) -> list[str]:
errors = []
if r.order_id != order.id:
errors.append(f"order_id {r.order_id} does not match the ticket's order {order.id}")
if r.amount_cents > order.total_cents:
errors.append(f"amount_cents {r.amount_cents} exceeds order total {order.total_cents}")
return errors04Why retry with the error, not just retry?
A plain retry samples again and hopes. Often it produces the same mistake, because nothing in the prompt changed. A retry that includes the precise validation error is a different request: the model now knows exactly what was wrong.
def triage(ticket: Ticket, order: Order, max_attempts: int = 3) -> TriageResult | None:
messages = build_messages(ticket, order)
for attempt in range(max_attempts):
raw = model.complete(messages, response_format=TriageResult)
try:
result = TriageResult.model_validate_json(raw)
errors = check_rules(result, order)
except ValidationError as e:
errors = [err["msg"] for err in e.errors()]
if not errors:
return result
messages += [assistant(raw), user("Fix these problems and answer again:\n- " + "\n- ".join(errors))]
log.warning("triage_invalid", attempt=attempt, errors=errors)
return None # caller escalates to a humanKeep the retry count small. If two corrected attempts still fail, the input is genuinely hard or the prompt is wrong, and a third attempt mostly adds cost and latency.
05What happens when it still fails?
Decide the fallback per call, in advance:
- Escalate to a person with the raw output and the errors attached.
- Degrade to a safe default, such as
reply_onlywith a holding message, when a reply is better than nothing. - Refuse the action entirely when the call controls something irreversible.
Never “fix” invalid output silently, for example by clamping an amount to the order total. That hides a real failure and makes the system look more reliable than it is.
Constrained decoding guarantees the answer parses. Only your system can guarantee it makes sense.
06Where this sits
The boundary belongs between L2 reasoning and L4 actions: nothing a model decides reaches a tool without passing it. It pairs naturally with effect classes on tools, which decide what happens after a valid decision is executed, and with evals on every prompt change, which track how often the boundary has to say no.