A user in acceptance testing opened a project charter and saw this on the “Key Risks” card: {'risk_name': ..., 'risk_description': ...}. The model had not failed. It had produced a perfectly reasonable risk object. Everything after the model had failed: the parser, the storage layer, the formatter, and the UI. Over a few months we hit three distinct classes of this problem in a Semantic Kernel application running on Bedrock. Each one looked like “the LLM is unreliable” from the outside. Each one was actually our code.
Failure 1: the leak
The intake agent produces a project charter from uploaded documents. It is a free-text LLM call parsed into an untyped dict. We had a Pydantic model describing the canonical shape, but it was documentation. Nothing enforced it at runtime.
So when the model emitted a risk as a dict with keys the formatter didn’t recognise, the formatter fell back to str(dict). That string went into the database. The frontend rendered it faithfully.
The fix had two layers, and I think both are necessary.
The first layer is a display chokepoint. On the backend, a humanize_value() helper collapses any value into clean text. On the frontend, a sanitizeDisplayText utility mirrors it and is applied in one place: the Markdown renderer that every section passes through. If a dict ever reaches the UI again, the user sees prose, not braces.
The second layer cleans the data at the source. A normalize_charter() function runs on every charter write, both the extraction path and the manual section edit path. Its docstring states three rules:
- Never raise. Any field that can’t be coerced is left exactly as-is.
- Never drop data. Unknown keys on structured items are preserved.
- Idempotent. Running twice yields the same result.
The idempotency rule matters more than it looks. It means we could apply the normaliser on every write without worrying about double-processing, and we could run it over existing records safely. The commit that shipped it reports running it across 372 live charters: zero raised, zero non-idempotent, 46 cleaned, zero leaks.
The shape of the coercion is simple. Aliases map to canonical keys, and money becomes a structured amount:
_RISK_ALIASES = {
"risk_name": ("risk_name", "name", "title", "risk", "risk_title"),
"risk_description": ("risk_description", "description", "desc", "detail"),
...
}
def _to_money(value):
if isinstance(value, dict):
money = dict(value) # preserve extra keys
money.setdefault("currency", "USD")
money.setdefault("value", None)
return money
if isinstance(value, (int, float)):
return {"value": float(value), "currency": "USD"}
return value # leave anything else untouched
We also pinned the element shape in the extraction prompt so the model emits the canonical form in the first place. Prompt, normaliser, renderer. Three places, because any one of them alone will eventually let something through.
Failure 2: the truncation
The charter has 14 sections. For a content-rich document, the full JSON exceeded 4,096 output tokens. The response was cut mid-object. The parser returned None. The code treated None as “nothing to save” and skipped the charter write. From the user’s side: uploads did nothing.
Nobody saw an error. The failure was a quiet early return.
We raised the output cap to 8,192. That helped and then overflowed again on richer documents. It went to 16,384, and finally to a named constant at 32,768, which sits well under the model’s 64K output ceiling.
But the output cap was only half the story. The input cap had been copied from the embedding service, which limits text to 8,000 characters. That limit was fine for embeddings and wrong for extraction. It silently dropped most of a real client workbook, so the intake cards came back empty. The extraction cap is now 60,000 characters per document, and the assembled multi-document context has its own 60,000-character lever.
One more detail I would not have predicted. When several documents are concatenated and then sliced to the window, order decides what survives. We had been loading documents in insertion order, which meant newly attached files fell past the window and never reached the extractor. The fix was to order newest-first:
stmt = (
select(Document)
.where(Document.project_id == project_id)
.order_by(Document.created_at.desc(), Document.id.desc())
)
The id.desc() tiebreak keeps the ordering deterministic when two documents share a timestamp. A small thing, but without it the same project could extract differently on two runs.
Failure 3: the timeout cliff
The RFP agent generates a Terms and Conditions section as JSON. It originally went through the tool-enabled agent chat path, which is capped at 4,096 tokens. Same symptom as before: truncated JSON, silent empty success.
The obvious fix was to double the budget to 8,192. We did. The model then generated a very long, comprehensive T&C document that took more than 120 seconds to produce and timed out, or got cut off at the new cap into unparseable JSON. A bigger budget had moved the failure, not removed it.
The fix that stuck bounds the prompt rather than the output. The prompt now says: at most 4 categories, at most 3 clauses each, 2 to 3 sentences per clause. That is roughly 1.5 to 2.5K tokens, so a 4,096-token budget has headroom. The whole call sits inside an explicit timeout:
try:
generated_data = await asyncio.wait_for(
rfp_agent._generate_json_response(prompt, max_tokens=4096),
timeout=120,
)
except asyncio.TimeoutError:
raise HTTPException(status_code=504, detail="Terms & Conditions generation timed out.")
A 504 the user can see beats a 200 with an empty body.
There was a second bug hiding in the same function. The generator was decorated with a JSON-validating helper, so it already returned a parsed dict. The caller re-parsed that dict as if it were a string, which raised a TypeError on every successful generation. A broad except turned that into an HTTP 500. The model was fine. The retry was fine. We were failing on success.
The shared toolkit
By the time we had hit all three, the codebase had three divergent copies of “parse JSON from an LLM response”. They were replaced with one function:
def parse_llm_json_response(response: str):
# 1. direct parse
# 2. ```json ... ``` fenced block
# 3. first {...} or [...] anywhere in the text
# returns None on failure, never raises
On top of it sits a decorator, @validate_json_output(max_retries=2), which re-runs the LLM call when the result can’t be parsed. Structured flows fall back to typed defaults when all attempts fail.
One warning about that stack. The low-level Bedrock client also retries, with tenacity, up to 3 attempts with exponential backoff on any ClientError. Stacked with the JSON decorator’s 3 attempts, a single logical call can become 9 model invocations in the worst case. Know that number before you set a request timeout.
I should be plain about what this is. There is no native structured output or JSON-schema enforcement anywhere in this pipeline. It is prompt, parse, retry, defaults. For a 14-section document with prose fields, that has been enough. For anything where a malformed field has a cost, I would want schema enforcement at the model boundary and I would not accept “parse and hope” as the final answer.
What I’d tell you to do
- Treat
Nonefrom a parser as an error, not an empty result. Every one of our silent failures was an early return onNone. - Put a display chokepoint in front of users and a normaliser behind storage. One layer is never enough.
- When you raise an output cap, check the input cap too, and check what order your inputs are concatenated in.
- Bound the prompt before you raise the budget. A bigger budget often just moves the failure to a timeout.
- Count your stacked retries. Two “3 attempts” layers is 9 calls.