All writing

Shipping LLM systems · 7 of 9

Thresholds are config, not code

Why approval limits, tolerances and cut-offs live in a per-tenant config service, and the bugs that taught us to keep two read paths.

Fortan Pireva 6 min read architectureconfigcompliance

The product requirements for our purchasing compliance build carried one rule that looked bureaucratic on the first read: no threshold-like value may be hardcoded downstream. Every approval limit, tolerance band, confidence cut-off and page cap had to be read from a config service, per tenant, at the moment of the decision. Six weeks later that rule had produced a 2,400-line registry document, thirty-three namespaces, one bug class I now watch for everywhere, and a test that exists because a numbered list in a document was forgotten four times. This is what I learned about building the thing that holds the numbers.

What the service actually is

The config service is a versioned key store, one row per (tenant, namespace). A namespace is a group of related keys, such as the thresholds for one compliance check, or the module kill switches. Each namespace is a Pydantic model registered in code with extra="forbid" and bounds on every field. A payload that contains a key the model does not know is rejected. A namespace that is not registered is rejected as unknown_namespace. The registry document is the frozen source of truth, and the code enforces it rather than trusting the convention.

Writes are deliberately heavy. The caller supplies the version it last read. The service takes a row lock, compares versions, and raises a stale-version error on mismatch. Two concurrent first writers cannot be caught by a row lock on a row that does not exist yet, so the unique constraint on (tenant, namespace) is the backstop, and the loser’s integrity error is converted into the same 409 the update path raises. Every successful write appends an immutable audit row, so history survives independently of the current value.

That is a lot of machinery for a dictionary of numbers. It is justified by what the numbers do: they decide whether a purchase requisition is blocked or passes. A silently lost write on that table is a compliance incident.

The bug that taught me to have two read paths

The first version had one read function. It returned the stored row verbatim. It worked through code review and through CI. Then someone added a key to an existing namespace.

Tenants whose row predated the key had no such key in their stored payload. The seeder never backfilled existing rows. A consumer did cfg[key] and got a KeyError. A fresh database, which is what CI uses, had the key from the start and passed every test. The service docstring now describes it as “Green through CI and review, broken on deploy.”

The fix was a second read path with a different name and a different promise:

async def get(client_id, namespace) -> dict:
    """What is STORED. Verbatim, unmerged."""
    ...

async def get_effective_config(client_id, namespace) -> dict:
    """What APPLIES. Stored payload validated through the
    namespace schema, so a key added after the row was
    written resolves to its registered default."""
    ...

Why not just merge defaults inside get()? Because put() is a whole-payload replace, and the operating runbook requires read-modify-write for every edit. If the admin plane read a merged payload and wrote it back, today’s defaults would land in the tenant’s row as if someone had chosen them. A later change to a default would never reach that tenant. The fix would have caused per-tenant drift. So the merge lives only on the consumer path, and the names say which is which.

There is a consequence that we accepted on purpose. Because the models forbid extra keys and the consumer path validates on read, removing a key from a schema turns every stored row that still contains it into a hard read failure. The docstring calls this “the intended behaviour, not a limitation”. Removing a key from a frozen registry should require a deliberate data migration, not a silent reinterpretation of what tenants’ rows mean.

A numbered list is not a control

The registry document has a process section for adding a namespace or a key. Two steps: PR the doc, then register the Pydantic model. Step two was forgotten four times. Each time the keys were documented, the board recorded them as available, and they were unreadable, because the service rejects unregistered namespaces and forbids unregistered keys.

After the fourth time we stopped treating it as individual carelessness. The test file that replaced the process step opens with the reasoning:

Four instances of one mistake is a design signal. A numbered list in a document is not a control; this file converts it into a property.

The test has one feature I would copy into any project. The check that every documented namespace is reachable runs in a subprocess that imports only the schema module. An in-process check is satisfied by whatever earlier tests happened to import, so it would have passed against all four bugs. The subprocess measures what a consumer reaching for the config service alone actually sees. The remaining checks compare documented keys against registered model fields in both directions, in-process, which is safe once reachability is known.

Flags: absent does not mean off

The feature-flag layer sits on top of the config service as a thin read module. Its one interesting decision is what an absent key means, and the answer is “the registered default”, not False.

That sounds like a detail until you look at the two kinds of flag that share the namespace. A module flag is a kill switch: absent must mean off. A check flag gates a compliance rule: absent meaning off would silently skip the check, which is fail-open in the only sense that matters for a compliance product. Module flags default to False in the registry, check flags default to True, and resolving to the registered default is the one rule that is safe for both. It also handles the real production case: tenants with a flags row written before the module flags existed, where every module flag is missing and must read as off without anyone backfilling anything.

Two more rules from the same module. An unknown flag name is a loud 500, never a quiet False, because a typo that disables a feature looks exactly like a deliberate kill switch. A disabled feature returns 503 rather than 404 or 403, because the feature exists and the caller is probably allowed to use it; it is just switched off.

There is no cache and no poll interval. A flip is visible on the next read. The cost is one indexed select per boundary check, which the module argues is the right trade for a kill switch: you reach for one during an incident, and “the switch takes up to N seconds” is the property you least want then. A database failure propagates instead of reading as off, because the caller was about to use that same database anyway, and hiding the fault behind a fake disable would only delay the diagnosis.

The older sibling’s lessons

Our sourcing product has a simpler config framework from a year earlier, and its verification doc is a catalogue of the mistakes the purchasing design was built to avoid.

The sourcing admin endpoint once wrote the uncoerced string into the degraded branch. bool("false") is True in Python. A flag written as the string “false” turned on and the endpoint answered 200 OK. The fix lifted the coercion to module level and called it from both branches.

Its cache is a per-process dict with a 30-second TTL. Under multiple uvicorn workers each process has its own copy, so a write becomes visible across workers only as their TTLs expire. The doc records why pub/sub was rejected: there is no IPC mechanism in the deployment to carry an invalidation, and Postgres LISTEN/NOTIFY would need a dedicated long-lived connection per worker. Eventually consistent, documented as such, and bounded by the refresh interval. Five of its eight agent flags are startup-cached on the kernel and need a restart, which was investigated and confirmed rather than hidden.

The contrast is the point. The purchasing design took the hit of one query per read so that it never had to explain a staleness window.

What I’d tell you to do

  • Give stored config and effective config different function names, and make the names literal. Merge defaults only on the consumer path.
  • When the same process step is forgotten more than twice, turn it into a test. Run reachability checks in a subprocess so earlier imports cannot mask the failure.
  • Decide what an absent flag means per flag kind, and register the default. Fail-closed points in opposite directions for kill switches and compliance checks.
  • Make unknown names loud. A silent False is indistinguishable from an intentional disable.
  • Write down the consistency model of your config reads. “Immediate” and “within one poll interval” are different products.