All writing

Shipping LLM systems · 8 of 9

Testing LLM systems without a cassette recorder

2,804 tests and zero recorded LLM calls: what mocking buys, what SQLite hid, and how to evaluate the model itself.

Fortan Pireva 6 min read llmtestingquality

Our sourcing service has 2,804 collected tests, 2,768 passing, and 80.15% line coverage. It makes zero recorded LLM calls. There is no VCR, no cassette directory, no replay of saved completions. Every model interaction in the suite is a mock. I want to be precise about what that buys and what it does not, because the honest answer is more useful than the coverage number.

The shim that made the suite fast

The ORM models use Postgres-only column types: JSONB, pgvector Vector, ARRAY, UUID, INET. The production DDL has to stay exactly that. The test suite had to run on in-memory SQLite, hermetically, with no external database.

The compatibility module uses two mechanisms, chosen per type by what actually breaks. For types whose value binding already works once the column exists, a @compiles(..., "sqlite") hook overrides only the DDL rendering: UUID to CHAR(32), INET to VARCHAR, JSONB to JSON. For ARRAY and Vector, which hold Python lists that cannot bind under their native types, a metadata patch swaps the whole type for a JSON variant using with_variant, which affects only the SQLite dialect. The Postgres path is untouched.

Two details in the conftest earned their comments. In-memory SQLite is per connection, so with the default pool every session would see a fresh empty database and fail with “no such table”. The engine uses StaticPool to keep one connection alive for the whole run. And there is a guard that refuses to start if the test database URL equals the application database URL, because the suite drops and recreates tables.

Coverage went from 27.79% to 80.15% on this foundation. The low end of the file table is where the model calls live: section extraction at 14%, the RAG service at 27%, the agent registry at 31%. The doc itself contradicts its own threshold, saying 75% in the summary and “threshold set to 27%” in the history table. I leave that in because it is the kind of drift this post is about.

What SQLite hid

The purchasing build, a parallel team on the same monorepo, learned the cost of the shim in one incident. A package was merged without its database migration. Because the test path builds the schema from the ORM, the suite stayed green. Real Postgres would have failed for every caller. The onboarding doc for new engineers records it as one of “the five rules people break in their first week”, right next to a second one: three commit() calls inside tests took the suite from 11 to 120 failures in one day, because the session fixture isolates only by rollback and a commit persists rows into the shared in-memory engine permanently.

The remedy was not to abandon SQLite. It was to run the full suite on Postgres as a separate bar and name the backend every time a number is quoted. The first full Postgres run came back at 4 failed / 3,037 passed. All four were classified: one documented S3 end-to-end test that fails on both backends, one Postgres-only bug with a fix waiting to merge, and two tests that executed a SQLite-only PRAGMA in setup and so had never been Postgres-capable. The baseline file carries a warning I now quote to people: “Different trees AND different backends. These two numbers are not a control for each other.”

The same file explains why a stale baseline is dangerous rather than merely wrong. The team had been checking “no new failures” against a list of eleven known failures from two weeks earlier. The suite had grown by about 60%. A machine diffing a fresh run against that list would see ten failures that no longer existed and have no way to tell an improvement from a broken collection, and no reason to notice a new failure hiding among ten it expected. The bare branch was observed ranging from 11 to 33 failures on an unchanged commit, which led to a two-sample rule before anyone was allowed to attribute a failure to a change.

Evaluating the model without mocking it

Mocks test the plumbing. They cannot tell you whether the model’s output is any good. For that we have an eval framework behind its own kill switch, because runs spend real tokens.

The LLM-as-judge for RFP overview narratives is where I learned the most. It runs a 1-to-5 rubric prompt several times and aggregates. It runs at temperature 0.3, and the docstring says why: at temperature 0 the calibration probe collapsed variance to zero, so the dashboard chip for “high judge variance, review” never fired. The signal we wanted was the disagreement between runs, and determinism erased it. Scores are normalized with (s-1)/4 so they sit on the same scale as the other metrics.

The RFP scorer’s docstring does something I wish every eval harness did. It splits its eleven metrics into two labelled groups. Four are “REAL SIGNAL” and reflect actual agent capability; the pricing-model match had a phase-one baseline of 0.333. Seven are “SMOKE / STRUCTURAL CHECKS” with high baselines because the underlying mechanism is simple. One of the seven is a pass-through of the agent’s self-reported confidence, annotated: phase one showed this is “UNCORRELATED with correctness (high confidence on wrong picks)”. The closing note says a pass on the seven smoke checks “does not certify agent quality”. Anyone reading a green dashboard gets told, in the code, which green means something.

Letting a model write tests, carefully

The end-to-end suite has a generator that takes one acceptance criterion from a work item plus the set of locators confirmed to exist on the live page, and drafts a runnable Playwright test. It uses Claude Sonnet 4.5 on Bedrock. Every locator in the output is mechanically restricted to the confirmed set, the same check the page-object generator uses.

The module’s own docstring draws the line that matters. A page object only claims that an element exists and is reachable, which is provable against the live DOM. A generated test additionally drafts assertion logic, a claim about intended behaviour that no amount of crawling can verify. “The test’s premise can still be wrong even when every locator in it is real.” So the pipeline runs generated tests as report-only, with continueOnError, and promoting a criterion to blocking after it has passed unattended several times is a human decision the generator never makes.

The same pipeline has a plugin whose first line reads: “CI guardrails against a run that reports green while testing nothing.” Local runs skip cleanly when credentials are absent, which is deliberate. In CI credentials are always supposed to be present, so the run is invoked with --fail-on-credential-skip --min-tests 1. A green run with zero executed tests is a failure.

What the numbers mean

Put together, here is what I believe the suite proves and does not prove.

It proves the plumbing: request parsing, persistence, state transitions, formatting, error paths, and that the system degrades the way we designed when a model call fails. It proves it fast, hermetically, on every commit.

It does not prove that a model call returns something good. That is the eval framework’s job, run deliberately, with real tokens, and read through the real-signal versus smoke-check split. It does not prove the schema matches production unless a Postgres run says so, with its backend named. And a generated end-to-end test proves its locators exist and nothing more until a human has watched it pass.

What I’d tell you to do

  • Run on SQLite for speed, but keep a Postgres bar and name the backend next to every number you quote.
  • Retire a known-failures baseline the moment the suite’s size changes materially. A stale list hides new failures among old ones.
  • If your LLM judge runs at temperature 0, you have deleted the variance signal you wanted. Pick a small non-zero temperature and run it more than once.
  • Label eval metrics as real-signal or structural in the code, next to their baselines. Do not let a self-reported confidence score count as evidence.
  • Treat model-generated tests as report-only until a human promotes them, and make a zero-test green run fail.