For months, our purchase requisition app had a compliance panel with twelve rules, green and red ticks, a “rerun checks” button, and a progress bar that moved. Buyers trusted it. None of it was computed. The statuses were hardcoded in a seed file, the rerun slept for a second and a half, and the frontend picked one to three warnings at random on a timer. This is the story of replacing that panel with real checks, what it cost, and the design rules that came out of it.
What the fake looked like
The backend seeded twelve compliance rules per case from a JSON file. Every case showed the same split: ten passed, two failed. The two failures were always the same rules, with the same failure reasons, with dates that came from the seed file rather than from the quote or the contract attached to the requisition. The system never opened those files.
The “rerun all” handler did this:
for r in rules:
r.status = "running"
await asyncio.sleep(1.5) # UI sees "running" briefly over SSE
# increment an in-process counter, then flip the seeded results
Click it once, you saw the seeded result. Click it again, everything passed. Restart the app and the counter reset. On the requestor side, the intake wizard ran its own set of checks on setTimeout, each waiting roughly 800 ms in sequence, then called a helper that shuffled the eligible checks with Math.random() and flagged one to three of them as warnings. Our own analysis doc called it “setTimeout theater,” which was fair.
The dangerous part was not that it was fake. Demos are allowed to be fake. The dangerous part was that fake rows and real rows rendered identically, in the same panel, with the same tick marks. A buyer had no way to tell which statuses meant something.
What the real version does
We had a validation engine (about 13.7K lines, vendored from an upstream team) that reads the requisition and its attachments and runs a set of checks: supplier match, item description match, quantity and unit of measure, price and currency, payment terms, category, after-the-fact detection, totals. The work was wiring it in behind the same API contract the frontend already consumed.
We traced one requisition through both paths: a small workwear order, four lines, five attachments, under $3,000. Before: twelve rules, ten passed, two failed, every time. After: fourteen rules (we added Item Description Match and Total Amount Calculation), eight passed, six failed, and each failure named the actual gap. The totals check found a $59.60 variance between the requisition and the quote. The description match found four lines with no matching quote line. The buying-channel check flagged a catalog bypass.
Each failed rule now carries a proposed_correct_value and instructions_for_requestor generated from the engine’s corrective actions, so the buyer can send the requestor a concrete fix instead of a tick mark. Progress streams over server-sent events as rule_update and summary messages, and the full result JSON is persisted on the job row so the comparison views have something real to read.
The bill for honesty, from our own before/after table:
| Fake | Real | |
|---|---|---|
| Cost per rerun | $0 | about $0.10-$0.30 in Bedrock tokens |
| Latency per rerun | 1.5 s | about 30-90 s |
| Same PR, same statuses? | always | only if the files say so |
| Survives restart? | no | yes, DB-backed |
Thirty to ninety seconds is a long time for a progress bar. We kept the streaming so the buyer sees rules flip one by one instead of staring at a spinner.
Deterministic first, LLM when confidence is low
The engine is not “ask Claude whether the PR is compliant.” Most checks are deterministic: a Decimal math engine for totals, embeddings for line matching, lookups for supplier and category. The LLM is called for edge-case reasoning, and only when a deterministic result is shaky. The gate is a small function:
def should_invoke_edge_case_reasoning(check_id, status, confidence, ctx):
if status == "NEEDS_REVIEW": return True
if confidence < 0.70: return True
if status == "WARNING" and confidence < 0.85: return True
# per-check grey bands, e.g.:
# check_04: 0.70 <= description similarity < 0.85
# check_06: 0.5% < price variance < 2.0%
# check_11: exactly one after-the-fact signal
...
A clean pass or a clean fail never spends a token. The grey bands are where a human would also hesitate, so that is where the model gets to reason. Its prompt ends with one instruction that sets the tone for the whole system: “Be conservative: when in doubt, flag for human review.”
Five outcomes, and only one of them blocks
The second thing we got wrong early was the shape of a result. A check used to produce pass or fail with a confidence. Then we needed “could not evaluate,” and then “does not apply to this tenant,” and we tried to squeeze both into the same bucket. It broke twice. Once, a check that was switched off for a tenant rendered as “needs more info” forever. Once, an unknown extension state was treated as open when it was closed.
The fix was an enum with five values and a ruling that they stay distinct:
class CheckOutcome(Enum):
PASS = "pass"
FAIL = "fail"
WARNING = "warning"
INCONCLUSIVE = "inconclusive" # could not be evaluated; someone can still supply the input
NOT_APPLICABLE = "not_applicable" # applicability gate ran; this check does not apply
INCONCLUSIVE is an open state. NOT_APPLICABLE is a closed one. Collapsing them, as the module docstring puts it, “means a closed state renders as an open one forever,” which “is exactly the shape that already broke twice in this build.” An enum value is opt-out: a consumer branching on outcome has to handle it or fail loudly. A details key would have been silently ignored.
The runner applies exactly one rule on top: only FAIL blocks submission. Warnings inform. Inconclusive waits. Not-applicable is non-blocking for free.
Refuse the fallback
With a feature flag controlling the real engine, the obvious code is “if the engine is off, fall back to the mock.” We did that for the manual rerun button, because a demo environment without Bedrock access still needs to show something. We refused to do it for the automated path that runs when a requisition arrives from the ticketing system. That handler skips validation and records the skip in its ledger, with a comment that says why: falling through to the mock “would write fabricated ComplianceRule rows a buyer cannot tell apart from real ones.”
That sentence is the rule. If you cannot run the real check, say nothing. Never write a fake answer where a real one is expected.
Measuring the AI after the fact
Once the statuses were real, we could ask whether they were right. The rule row stores two values: ai_status, the model’s verdict frozen at evaluation time, and status, the reviewer’s final decision. Every accuracy metric compares the two. Overrides need a justification of at least twenty characters and a known category. A check whose override rate goes above 15% (the default; tenants can change it) shows up as “Needs Attention,” with a twelve-week trend behind it. Every case where a reviewer contradicts a real AI verdict is tagged for the retraining dataset.
Separately, any check with confidence below 0.70 is a hard route to human review. Random sampling for the rest uses a deterministic hash of case, rule, and tenant, so a sampled case stays sampled.
What I’d tell you to do
- Fake results are fine in a demo. Fake results rendered next to real ones, in the same shape, are not. Label them or remove them before the first real user sees the screen.
- Put the LLM behind a confidence gate. Clean passes and clean fails should cost zero tokens.
- Give “could not evaluate” and “does not apply” different enum values. You will need both, and merging them will bite you twice.
- Decide one blocking rule and write it down. Ours is “only FAIL blocks.”
- When the real path is unavailable, refuse. A skipped check with a reason is honest. A fabricated pass is a lie with a tick mark.