- Human-in-the-loop means an agent pauses before a sensitive action and continues only after a person approves, rejects or edits it.
- Implementing the pause as a blocking call (a thread or request waiting on a reply) fails on deploys, timeouts and restarts, and holds compute while nobody is looking.
- A durable wait persists the run's state, records an approval request, and resumes from the log when a signal arrives. The waiting run holds no compute.
- Good approvals show the exact action and its arguments, have a timeout with a defined default, and record who decided what, for audit.
Most agent demos implement approval with something like input("Approve? [y/n]"), or a web request that stays open until someone clicks a button. Both work for a minute. Real approvals take hours: the manager is in a meeting, it is the weekend, the request sits in a queue. Anything that waits in memory for that long will not be there when the answer comes.
01What goes wrong with a blocking human-in-the-loop approval?
A blocking approval keeps the run alive in a process while waiting. In production that process will not survive:
- Deploys and restarts kill it, and the pending approval is lost or, worse, re-requested.
- Timeouts on load balancers and HTTP clients end the request long before a person replies.
- Compute is held for the whole wait: a worker slot, a connection, sometimes a model session.
- Nobody can see what is pending, because the state only exists in memory.
02The decision
03What does a durable wait look like?
The workflow asks to wait for a named signal with a timeout. Under the hood the runtime writes the wait into the run's log, releases the worker, and later replays the run up to that point when the signal arrives.
let approved = if amount_cents > APPROVAL_THRESHOLD_CENTS {
let a: Approval =
ctx.wait_signal("manager_approval", Duration::from_secs(2 * 24 * 3600)).await?;
a.approved
} else {
true
};That snippet is from the refund agent in KEEL, our durable execution runtime. The two-day timeout is real time; the run consumes nothing while it waits. You can build the same pattern with any durable workflow engine, or with a state table and a resume handler.
04What should the approval request contain?
The person approving is the last check before something irreversible happens. Give them what they need to decide in seconds:
- The exact action and arguments. “Refund $640.00 to card ending 4242 for order ORD-1002”, not “The agent wants to issue a refund”.
- Why. The model's short reasoning and the evidence it used: the ticket, the order record.
- The choices. Approve, reject, or edit the arguments. Editing is often the most useful option.
- The deadline and the default. What happens if nobody answers. For money, the safe default is usually reject and notify.
05What happens after the answer?
- Approved: the run resumes and executes the action with the same idempotency key it would have used anyway.
- Rejected: the run resumes with the rejection as input, so the agent can explain it to the customer or take a different path.
- Edited: the run resumes with the edited arguments, and the log records both the proposal and the edit.
- Timed out: the run resumes with the default, and the timeout is recorded as a decision in its own right.
Every outcome lands in the log with who decided and when. That record is what an audit, or a curious customer, will ask for.
A good approval step holds the decision, not the server.
06Where it sits in the stack
The signal and the resume belong to L4 actions, next to effect classes and idempotency keys. The request itself is L5 product and workflow: it shows up where the approver already works, in a chat thread, an inbox or a queue, not in a new dashboard nobody opens. See put the model inside the workflow.