Our case assistant could send email to a requestor. Early on, the way it decided to send was a phrase list. If the user’s message looked like a “yes,” a Python helper injected a line into the prompt: “[CRITICAL INSTRUCTION]: you MUST call send_email_to_requestor NOW.” It worked in demos. It was also not a control. A prompt injection, a hallucinated phrase match, or a replayed “yes” could all fire a real send. This post is about what we replaced it with: a small state machine that every external write in the service now has to pass through, and the caveats we wrote into the code so nobody oversells it.
The problem with “the model decides”
An agent with tools will call a tool when it believes it should. For reads, that is fine. For writes to systems outside your database, it is not. Creating a ticket in ServiceNow, inserting a purchase requisition, sending an email: none of these can be un-done by rolling back a transaction. The model’s belief that the user said yes is not evidence that the user said yes.
Our old heuristic made it worse by pretending to be a safeguard. The comment that now sits where it used to live says it plainly: “coercing a tool call by text pattern-matching is not a security boundary.”
The state machine
Every external write starts as a proposal:
proposed --confirm(token)--> confirmed --execute()--> executed
proposed --cancel()-------> cancelled
confirmed --cancel()-------> cancelled
proposed|confirmed --(TTL elapsed, checked lazily)--> expired
propose() records what the write would do, hashes the parameters, and returns a confirm token exactly once. The database stores only sha256(token), the same pattern we use for public case links. The raw token is never persisted.
confirm() requires the proposal to be in proposed. A second confirm hits InvalidStateTransitionError. It does not silently succeed again, because a replayed confirm is exactly the thing we were trying to stop.
execute() is the opposite: deliberately idempotent. If the proposal is already executed, it returns the stored result without calling the executor a second time. That is what makes a retried HTTP call, or a duplicated tool invocation from the agent, safe.
Expiry is checked lazily on the next touch. The default window is 900 seconds. A caller may ask for a shorter window; the service takes min(requested, configured), so nobody can extend it from the outside.
Tamper detection
The parameters are hashed at propose time with a canonical serialization (sorted keys, fixed separators). When the caller re-supplies parameters at execute, the hash is compared. A mismatch raises ParamsMismatchError. The docstring gives the scenario it exists for: “propose a $10 PO amendment, confirm, then execute with a $10,000 payload instead.”
One executor goes further. The path that inserts a requisition into the client’s purchasing system rebuilds its payload at execute time from the stored proposal, so any drift between what was proposed and what would be written fails closed. The reason is in one line: an inserted PR cannot be un-inserted.
Two executes at once
Check-then-act is the classic race. Two concurrent execute() calls could both read confirmed, both run the executor, and send twice. We hold a SELECT ... FOR UPDATE on the proposal row for the duration of the state check, the executor call, and the result write, inside one transaction. On Postgres, the second caller blocks, re-reads, finds executed, and returns the cached result.
Two honest notes live next to that design. First, on SQLite the lock is a documented no-op (SQLAlchemy drops with_for_update() silently), so the guard degrades to the state check alone; our unit tests prove the sequential-retry case on both dialects and say in their docstring what the SQLite path does not prove. Second, holding a row lock across a network call is a choice with a cost, and every executor has to carry its own timeout so a hung vendor call does not hold the lock forever.
The gate is enforced at runtime, not by convention
A state machine that writers can bypass is documentation. So the service has a second piece: require_gated_write(). It is the first line of every ServiceNow write (create_case, push_pr_intake, push_compliance_results, submit_case_feedback). It raises UngatedExternalWriteError unless it is called from inside the dynamic scope of ActionGate.execute_external_write(), and it sends an admin alert when it fires. The connector factory returns gated wrappers instead of the raw HTTP client, so a new call site cannot forget.
The scope is tracked with a contextvars.ContextVar, which gives per-task isolation for tasks that are awaited directly. The first version used a bare boolean, and an adversarial reviewer on another machine reproduced an escape: asyncio.create_task() copies the current context at spawn time, so a task spawned inside an open gate keeps _gate_open=True for its whole lifetime, long after the scope that spawned it has exited. No executor spawned tasks at the time. The reviewer’s note says it anyway: “the guard is advertised as fail-closed runtime enforcement, so the escape needed fixing rather than just documenting.” The fix was a per-open identity token in the ContextVar plus a module-level set of currently valid tokens; leaving the scope removes the token from the shared set, so a copied context holding a stale token no longer passes.
The same-turn rule, and what it does not prove
On the agent path, the model still decides when to call confirm_send_email, the same way it decides to call any tool. So the server records the user-turn id at propose time and refuses a confirm attempted in the same turn:
class SameTurnConfirmError(ActionProposalError):
def __init__(self):
super().__init__(
"This action cannot be confirmed in the same turn it was proposed in, "
"ask the user to confirm, then try again once they've responded.",
status_code=409, code="proposal_same_turn_confirm",
)
That forces at least one turn boundary between draft and send. It is real, server-enforced, and the model cannot spoof it. The model docstring then says what it is not: “It is NOT a guarantee that a human clicked anything.” The plugin docstring is blunter: “Treat this as turn-separated, model-mediated confirmation, not a security boundary.” The old phrase list did not disappear either. It moved into the system prompt, where it is less reviewable than Python was, and the code says so.
I would rather ship those sentences than a false sense of safety. The next step, a per-tenant kill switch for agent-initiated sends, is noted as a TODO where it belongs.
Small decisions that add up
- An unknown or cross-tenant token returns a 404
proposal_not_found. The code does not explain why; my reading is that a 403 would confirm the token exists. - Any actor in the same tenant may confirm. Actor mismatch is recorded for alerting rather than refused.
- The proposal’s idempotency key is forwarded to ServiceNow on the envelope, so a retry at our end de-duplicates at the vendor too. Keys of expired or cancelled proposals are retired, so one timeout does not block that subject forever.
- Notifications are the contrast case. They are Jinja2-templated and never generated, so they do not need a gate. LLM-written text, like the email draft, only reaches the outside world through one.
The gate has 18 tests and the proposal service 47. The numbers matter less than what they cover: replay, double-confirm, idempotent execute, params mismatch, the create_task escape.
What I’d tell you to do
- Separate proposing a write from executing it, and make the model only ever propose.
- Hash the parameters at proposal time. Compare at execution. Fail closed on mismatch.
- Enforce the gate at runtime at the write site, not by code review. A factory that only hands out gated clients helps.
- Make execute idempotent and confirm non-replayable. They are different operations with opposite rules.
- Write the limits into the code next to the mechanism. “Not proof a human clicked” is worth more than a security claim you cannot back.