All writing

Shipping LLM systems · 6 of 9

Shipping an agentic rewrite behind flags that keep prompts byte-identical

How to rewrite a working assistant as a tool-calling agent behind flags that leave prompts byte-identical, using shadow mode.

Fortan Pireva 6 min read llmmigrationtesting

We had a working supplier discovery assistant. It routed user messages with regular expressions, ran a hybrid search, and built result cards in Python. It worked, and people used it. We wanted to rewrite it as a tool-calling agent on Semantic Kernel and Bedrock, where the model decides when to search, when to add a supplier, when to refine. The question was how to ship that rewrite to production without changing what any user saw until we chose to.

The answer we landed on has three parts: characterization tests that pin today’s behaviour, feature flags that default off with a provable “byte-identical when off” property, and a shadow mode that runs the new path next to the old one without ever serving it.

Phase 0: pin what exists before you touch it

Before any new code, we wrote characterization tests. Not tests of what the system should do. Tests of what it does today, including the parts we didn’t like.

The intent router is a good example. It parses “add more suppliers” as an add-supplier action with the supplier name “more suppliers”. That is wrong. It is also what production does, and the test asserts exactly that:

("add more suppliers",
 {"type": "add_supplier", "supplier_name": "more suppliers"}),

If a later change fixes that behaviour, the test fails and someone has to decide, on purpose, that the fix is wanted. Nothing drifts by accident.

The same suite snapshots the context block the agent injects into the prompt, verbatim:

=== SEARCH RESULTS (use ONLY these, never invent suppliers) ===
...
=== END SEARCH RESULTS ===

It also pins the shape of the actions[] payloads the frontend consumes, and it asserts that every agentic flag defaults to False. The plan document counts 61 characterization cases across intent corpus, context strings and action shapes.

That last assertion is the cheapest and most important test in the suite. If someone flips a default in a config refactor, CI goes red before a deploy does.

Flags off means byte-identical, and that is an invariant to test

The feature flags live in settings: ENABLE_AGENTIC_DISCOVERY, ENABLE_AGENTIC_DISCOVERY_SHADOW, ENABLE_AGENTIC_INTAKE, ENABLE_AGENTIC_NEGOTIATION. All default off.

The property we wanted is stronger than “the feature is hidden”. It is that with the flag off, the prompts and tool configuration sent to the model are byte-identical to what they were before the rewrite existed. If the bytes are identical, the model’s behaviour is unchanged by construction. No argument about “probably fine”.

This has a consequence for architecture. The agentic plugins decide kernel plugin membership. They are registered once at startup, inside a function whose docstring says in capitals that a restart is required. Writing the flag through the admin config endpoint updates the database and the settings object, but it does not register or unregister a plugin in a running process.

Making those flags live would be a redesign, not a per-call check, and the docstring lists why:

  1. Each agent reads its flag in __init__ to set its tool scope, so scope is fixed at construction.
  2. There is one kernel shared by every request. Adding and removing plugins per request would race.
  3. The alternative, register everything and hide plugins per request, depends on “hidden” being byte-identical to “never registered”. That equivalence is exactly what we have not yet proven; the tests for it are in the failing baseline.

The last line of that docstring: “Do not describe these five as runtime-live without a test that proves it.” I like that a comment can carry a rule like that.

Shadow mode: run the new path, serve the old one

Characterization tests prove the old path is unchanged. They say nothing about whether the new path is good. For that we needed real traffic without real consequences.

Shadow mode does this. When ENABLE_AGENTIC_DISCOVERY_SHADOW is on, the procedural path handles the request and the user gets its result as before. In the background, the agentic path runs the same turn. Write tools simulate success and record what they would have done. The conversation save is skipped, so shadow turns never appear in a real conversation. A telemetry row named chat_agentic_shadow_compare captures served-versus-shadow actions, suppliers and latency.

The mechanism underneath is small. Agent instances are singletons in the registry, so a shadow flag on the instance would leak between concurrent turns. The state lives in contextvars instead:

_shadow_mode: ContextVar[bool] = ContextVar("agentic_shadow_mode", default=False)

@contextmanager
def shadow_mode():
    token = _shadow_mode.set(True)
    try:
        yield
    finally:
        _shadow_mode.reset(token)

Contextvars are isolated per asyncio task and propagate into everything the turn awaits, including the kernel invoke and the tool executions. A write tool checks in_shadow_mode() and records intent instead of acting. The per-turn card-action buffer uses the same pattern for the same reason.

Guardrails belong inside the tools

In the procedural version, guardrails lived in the orchestration code: deduplicate against suppliers already on the project, drop recommendations below a minimum score. When the model is deciding what to call, orchestration code is no longer in the loop.

So the guardrails moved into the tools. The discovery plugin’s docstring puts it this way: the tool-calling loop “physically cannot violate them”. A discover tool filters out anything below MIN_RECOMMENDATION_SCORE before the model ever sees it. An add tool deduplicates automatically. The model can be as creative as it likes about which tool to call; it cannot produce a result that breaks the rule, because the rule is on the other side of the tool boundary.

The loop itself is bounded. BEDROCK_MAX_TOOL_ITERATIONS defaults to 3, down from 12, with a comment saying “to prevent runaway loops”. Each tool invocation gets its own trace span, and a tool.repeated attribute is set when the same tool is called more than once in a turn. Tool results handed back to the model are truncated to 4,000 characters.

Two bugs the flags-off design caught

Two incidents are worth recording because both happened with the flags on in a non-production environment, and both would have been invisible without the rest of the design.

First, duplicate tool names. The agentic intake plugin defined a get_project_charter tool. So did the existing intake plugin. With the agentic intake flag on, both landed in the Bedrock tool configuration. Bedrock rejects duplicate tool names, so every agent that didn’t scope its tools, meaning RFP, orchestration and the command-center helper, failed with a validation exception. The fix removes the duplicate, dedupes names when building the tool config with first definition wins and a warning on skip, and adds a regression test asserting the intake plugins share no tool names.

Second, a cache key poisoned by a positional argument. The flag reader has the signature _get_flag(key, project_id=None, fallback=False). All five startup call sites passed False positionally, so it bound to project_id. Every lookup cached under (KEY, False), while cache invalidation pops (KEY, None). Those entries could never be invalidated. The resolved value happened to be correct, because if project_id: is falsy for False, which is why nobody noticed. The fix passes fallback= by keyword and adds a test that drives the real registration path and asserts the cache keys.

Neither bug was about the model. Both were about plumbing that only exists because of the model.

Where it stands

As of the last plan update, all the prototypes are implemented with flags off. Shadow mode is wired. The live-Bedrock shadow comparison, the step that tells us whether the agentic path is actually better, has not been run. I consider that the honest state: the safety net is built and tested; the thing it protects has not yet earned its way through it.

What I’d tell you to do

  • Write characterization tests first, including the bugs. A snapshot of today’s quirks is the only thing that lets you prove you haven’t changed them.
  • Make “flag off” mean byte-identical bytes to the model, and assert flag defaults in CI.
  • Use shadow mode to collect evidence before you serve a new path. Simulate writes, skip persistence, log the comparison.
  • Keep turn state in contextvars when your agents are singletons.
  • Put guardrails inside tools, not around the loop. The loop is the model’s now.