A supplier match drawer in our sourcing platform showed four criteria with moderate scores next to a low overall match. Users asked the obvious question: if these four look fine, why is the total bad? We had no answer, because the three criteria that explained it were not on the screen, and they carried 45% of the weight. That was the visible symptom. The real problem was that the headline number and the breakdown had been produced independently, and nothing in the system forced them to agree.
This post is about the discipline we ended up with: the model gives judgement, the code does arithmetic, and every number a user sees is traceable to one of the two.
How the score used to be made
Our supplier discovery agent proposes suppliers for a sourcing project, then scores how well each one fits the project charter. Three sources feed it: the client’s own supplier history, a wider reference network searched with vector embeddings, and web discovery through Amazon Nova.
When I audited the scoring in July, I found that the numbers on screen were not what they claimed to be.
Web suppliers were scored by list position. The code was literally this:
web_score = max(0.30, 0.85 - idx * 0.03)
The first web result got 85%. The second got 82%. No comparison to the project was involved. A supplier we had never evaluated showed a confident 85% because it happened to come back first.
The real scorer was mostly ignored. A charter-aware LLM scorer existed, with seven weighted criteria. It was blended into the final number at only 40%: 0.6 * positional + 0.4 * llm. And the save path fell back to the raw positional score whenever the blended score was missing, which was the path that produced the 85%.
The breakdown was fabricated on the frontend. The drawer that showed “Capability: 100% Match” and “Regional fit: Partial” decided those labels from whether the field was populated on the supplier object and from the single overall score. It never looked at the charter. The narrative text was a score-bucketed template.
So we had a positional number dressed as a relevance score, a real score diluted to a minority share, and an explanation that explained nothing.
The two-number problem
Once the LLM scorer was promoted, we hit a second, subtler version of the same issue, which we logged as Bug 42a.
The prompt asked the model for two things in one JSON object: an overall score from 0 to 100, and a percentage for each of seven criteria. The prompt asked them to “roughly agree.” That is a request, not a constraint. The model could return an overall score of 0 alongside seven criteria at 10% each. Each number was individually plausible. Together they were nonsense, and the headline could never be reconstructed from the breakdown.
On top of that, only four of the seven criteria were shown in the drawer. The hidden three (performance, compliance, financial) carried 45% of the weight. A supplier could tank on the hidden factors while the visible rows looked moderate, which is exactly what the user saw.
The fix: judgement from the model, arithmetic from the code
The change was small in code and large in principle.
The model is now asked only for per-criterion judgement: a 0 to 100 percentage for each of seven criteria, plus a short reason. It is never asked for the overall score. The overall score is computed:
_CRITERIA_WEIGHTS = {
"capability": 0.30, "regional_fit": 0.15, "performance": 0.15,
"compliance": 0.10, "financial": 0.10,
"reputation": 0.10, "strategic_fit": 0.10,
}
def compute_weighted_criteria_score(criteria_percentages):
# weighted sum over the criteria the LLM actually returned;
# missing criteria are excluded and the weights renormalised
The docstring says what this buys us. The headline “is always the literal weighted sum of what the drawer shows, reproducible and auditable, not ‘the AI said so.’”
Two rules follow from this.
Never invent a missing criterion. If the model omits a criterion or returns something unparseable, that criterion is None. It is excluded from the sum and the remaining weights are renormalised. We do not fill the gap with a formula-derived number, because that would reintroduce the fabrication we just removed.
Show every point of weight. All seven criteria now have a visible row. If the score is low, the row that made it low is on the screen.
The clamp, and why it is a compromise
There is one more guardrail, and I want to be honest that it is a judgement call.
Sometimes a per-criterion percentage from the model is wildly inconsistent with the overall score. The example in the code is a supplier at 9% overall with a 95% “regional fit.” Both numbers are real model output. Shown together, they look absurd.
We clamp. The criterion is capped to a ceiling tied to the score band, and a warning is logged. The docstring lays out the alternatives we rejected: leaving it as-is lets the UI show something absurd; replacing it with a formula-derived number brings back the fabricated-number problem for that criterion. Clamping keeps the model’s relative judgement (a criterion clamped from 95 to 40 still reads as the strongest of its group) while guaranteeing the UI never contradicts the headline. The log entry keeps it auditable.
I would not call this elegant. I would call it less misleading than either alternative, and visible.
The search that was not searching
While validating this end to end I found the deepest version of the same disease, one layer down.
The “internal” supplier search was supposed to be vector similarity. It was actually full-text search against a vector table that no longer existed. The query embedding was computed, accepted, and silently discarded. Every “semantic” result was keyword luck.
Fixing it meant wiring real HNSW cosine search against the current table, which is embedded with the same Titan v2 model. The verification numbers from the commit: internal match type went from 9 real matches out of 49 to 100 of 100, zero duplicate suppliers, zero tier-versus-score mismatches.
The same pass fixed three related number lies:
- The supplier tier (A, B, C) was computed once from a provisional retrieval score and never recomputed after the LLM blend. A supplier could show Tier A next to a 35% match.
- The client-spend table has many rows per supplier. A join multiplied one matching supplier into N duplicate results.
- The similarity
thresholdparameter was a no-op everywhere. Wiring it up exposed defaults calibrated for a no-op, so it was recalibrated to 0.2 from observed cosine ranges.
Every one of these is a number on a screen that nobody could trace to a computation.
What I’d tell you to do
- Split every displayed number into judgement or arithmetic. The model owns judgement. Code owns arithmetic. Nothing is both.
- Never ask a model for a total and its parts in the same call. Ask for the parts. Compute the total.
- Show every input to a computed score. A hidden 45% is a bug report waiting to happen.
- When the model omits something, leave it empty. Filling it in is fabrication with extra steps.
- Audit the fallbacks. The positional score, the frontend template and the dead threshold all lived in fallback paths that ran more often than the real thing.