KEEL: agent runs that survive a crash and never refund twice
“The agent crashed after calling Stripe. Did the customer get refunded zero times, once, or twice?”
- LAYERS
- L0 Infrastructure · L2 Reasoning · L4 Actions
- BUILT WITH
- durable execution · replay · fork-and-diff
- STATUS
- In build. The core runtime works; recordings are being re-made on real models.
01Problem
An agent run is long, non-deterministic, and touches the real world. Most agent frameworks keep the loop in process memory, so when the process dies mid-run the work is either lost or repeated. Repeating a refund, an email, or a deploy is an incident, not a retry.
Generic durable-execution engines fix the crash, but they treat a model call as an opaque black box. They can’t tell you which prompt and model produced a decision, and they can’t re-run a real production run on a new model to see whether it would have done better.
02System concept
Don’t make the model deterministic; make the run deterministic by recording. Every source of non-determinism goes through one boundary (ctx) and lands in an append-only, hash-chained log before it is used: model output, tool results, the clock, randomness.
A run is then a pure fold over its log. Replay re-reads results instead of recomputing them. A fork is a new branch that shares the log up to step N. A side effect is safe because its intent is durable before it happens.
03Interaction
Run a refund agent on a real ticket, then kill its worker process at the worst moment: after Stripe has moved the money, before the agent has written it down. A fresh worker picks the run up, replays three recorded steps with zero model calls, retries the refund with the same idempotency key, and Stripe hands back the existing refund.
Then do it 20 times at random kill points: 20 runs, 20 refunds. Finally, fork the run at its first model call onto a cheaper model with refunds simulated, and compare the two branches side by side: the steps taken, the reply, and tokens, cost and time.
04Architecture
- L4 Actions. Every tool declares an effect class (pure, idempotent, compensable, irreversible) and a tier that says what the API lets us guarantee: A idempotency keys, B reserve-then-confirm, C queryable, D blind. The coordinator runs intent → prepare → commit, and always reconciles before retrying an ambiguous result. Human approvals are durable signals; a run can wait days holding zero compute.
- L2 Reasoning. Every model call is fingerprinted (model, messages, tools, parameters). Replay fails loudly with a diff if code or prompt changed, and fork-and-diff turns real traffic into an evaluation set.
- L0 Infrastructure. Stateless workers claim branches with fenced leases on Postgres. A worker that loses its lease cannot write. Each step carries tokens, cost and latency, so economics come straight out of the log.
05Technical decisions
- Postgres first, FoundationDB later. The storage trait is the contract, not the database. One Postgres primary carries a long way, and fencing is a row lock in the same transaction as the append. Rejected: starting on a distributed store before there is a throughput problem.
- Step identity is name + counter, not position.
triage#2survives small code edits that would shift every positional index. Rejected: positional replay, as used by most generic engines. - Four honest tiers instead of one “exactly-once” claim. Exactly-once delivery is impossible over a network, so each adapter states what it can guarantee, and tier D runs park as “in doubt” instead of guessing. Rejected: marketing exactly-once for every tool.
- Deterministic simulation from week one. The whole system runs per seed with injected crashes, zombie workers and lost responses, then five invariants are checked. 100,000 seeds, 0 violations; any failing seed reproduces exactly. Rejected: chaos testing as a late hardening task.
06Prototype / demo
The demo replays runs that were recorded from the real engine: real worker processes, real kill -9, Postgres, and a Stripe-compatible payment API. Every number on screen comes from those logs, and keel replay reproduces any of them offline.
Run it yourself in one command: cargo run -p agensphere-keel-cli -- dev.
Before launch the recordings will be re-made with Azure AI Foundry models and Stripe test mode. Until then the demo labels them as scripted-model / Stripe-API-fake runs.
RUN THE DEMO ↗07Production path
- Now (in build). The core runtime is done: crash recovery, effect tiers, timers and signals, fork and diff, the simulation suite and the chaos harness.
- Next. A blob store and snapshots for long histories, a per-step overhead benchmark (target < 5 ms), and recordings on real models and Stripe test mode.
- SDK launch. Python and TypeScript bindings over the one Rust core, a Pydantic AI integration, and a web console with a replay scrubber.
- Hardening. Versioned replay for code deployed mid-run, multi-tenancy and encryption, and a 1M-seed nightly run.
- 1.0. Three design partners run production traffic for 30 days with no duplicate effects and no lost runs.