All writing

Shipping LLM systems · 9 of 9

Disclosed limitations beat confident summaries

Design docs said inputs were sanitized and the code said otherwise. Why drift is normal, and the habit of writing limitations down.

Fortan Pireva 4 min read llmdocumentationsecurity

Our security design document says all inputs received by agents are filtered and sanitized, and that only the orchestration agent invokes downstream agents. Our developer guide lists an ObservabilityAgent. The code concatenates the user’s message straight into the prompt, the orchestration agent is flagged off, and the agent that actually exists is called CommandCenterAgent. None of this is anyone’s fault. It is what a fast-moving AI codebase does to its documentation, and the habit that fixes it is cheaper than the habit that prevents it.

A short catalogue of drift

I collected these over a few hours of reading docs against code on a Semantic Kernel platform I help lead. Each one was true when written.

The developer guide says Azure OpenAI is the preferred provider. The config defaults USE_AWS_BEDROCK on, and every production call goes to Bedrock. The same guide says the default embedding dimension is 384, for a sentence-transformers model that was replaced months ago by a 1024-dimension Titan model. It points to an implementations.py file that became a package of six subfolders.

The architecture doc and the guide both list an ObservabilityAgent. The registry constructs a CommandCenterAgent. Same slot, different name, different job.

The security document’s threat table says the orchestration agent is the only path to downstream agents. The product overview, written later, describes orchestration as “an optional, flagged-off capability rather than something active by default.” Both documents are in the same folder.

A commit titled “temporarily disable web grounding” landed on 22 April. It removed the block that attached the Nova grounding system tool to a request. The service’s docstring still says it “specifically supports Amazon Nova models and Web Grounding,” and the enable_grounding=True parameter is still in every signature. The parameter is accepted and does nothing. Every feature downstream that calls itself “web research” has run on the model’s own knowledge since spring.

A pull request that introduced the global guardrail prompt also described a response cache: identical query, no context change, serve the cached answer, zero tokens. The cache was removed before merge, 36 lines out of the base agent, in a commit titled “lint isssues.” The pull request description, which is what people read, still advertises it.

Why this is normal, not negligent

Every one of those changes was a good decision made quickly. The provider switch was a cost and compliance call. Grounding was disabled because it caused the model to hallucinate anonymised supplier names. The cache was removed because it masked context changes. The agent rename reflected what the thing actually did.

What none of those changes had was a step that said: now find every sentence that was true before this commit and is false after it. That step is expensive when done by hand, and in a codebase where an agent can make the change in minutes, the documentation lag is not weeks. It is the same afternoon.

So the question is not how to keep docs perfectly current. It is how to make the gap visible and cheap to close.

The habit the team adopted

On the parallel build I was part of, the onboarding note for a new engineer carried one line that became a rule: this project rewards disclosed limitations over confident summaries. A later ruling made it concrete for status: code is the authority for “done,” and the repo outranks the tracker and the spreadsheet.

In practice that meant three things.

Every status claim links to something checkable. A test name, a commit hash, a file and line. A claim with no link is a lead to verify, not a fact to repeat.

Limitations are written next to the feature, not in a separate risks section. The action gate that controls agent side-effects has, in its own model docstring, the sentence that it “is NOT a guarantee that a human clicked” anything. The retrieval scorer’s docstring separates the four metrics that carry real signal from the seven that are structural smoke checks. Those sentences are the most valuable documentation in the repository because they are where a reader will actually look.

And when a document is found wrong, the correction stays in the document with the old text struck through and the date of the fix. The implementation plan has a line that reads, in capitals, do not quote the 7 figure, it is stale. A reader who finds the old number also finds the warning.

What I would add

If I were starting a new AI codebase today I would add one thing to the definition of done for any pull request that changes behaviour: a grep. Search the docs folder for the feature name, the flag name, the model id. Paste the hit list into the pull request. Either fix the hits or write “known stale, see issue.” The grep takes a minute. The confident summary that goes stale takes a reader much longer to unlearn.

What I’d tell you to do

  • Assume the docs are wrong in proportion to how fast the code moves. Read them as leads.
  • Put limitations in docstrings next to the feature. That is where people look.
  • Make “code is the authority for done” a stated rule, so nobody has to argue it.
  • Add a doc grep to the pull request template for behaviour changes.
  • Reward the engineer who writes “this parameter does nothing” over the one who writes “supports web grounding.”