AGENSPHERE/ JOURNAL
← JOURNAL
J-001TEARDOWN4 MIN READ

KEEL: make the run deterministic, not the model

A refund agent crashed after Stripe moved the money. KEEL is the durable execution runtime we are building so the answer to “how many refunds?” is always one, and so any run can be replayed or forked onto another model.

IN SHORT
  • Durable execution for agents means every source of non-determinism (model output, tool results, clock, randomness) is written to an append-only log before it is used.
  • With that log, a crashed run resumes on a fresh worker with zero repeated model calls, and a side effect like a refund is retried with the same idempotency key instead of run twice.
  • Every tool declares an effect class and an honest tier (A to D) that states what the external API lets the runtime guarantee. There is no blanket exactly-once claim.
  • The same log makes replay, fork and diff possible: re-run a real run offline, or fork it at step N onto a cheaper model and compare cost and outcome.

KEEL is PW-01 on our Artifacts shelf: a durable execution runtime for AI agents, written in Rust, with Postgres underneath. It is in build. This entry covers the problem it solves, how it works, one decision we made on purpose, and what is still missing. Every number here comes from recorded runs of the real engine, not from client work.

01What problem does KEEL solve?

An agent run is long, non-deterministic and touches the real world. A support agent reads a ticket, asks a model what to do, looks up an order, and issues a refund. Most agent frameworks keep that loop in process memory. If the process dies halfway, the work is lost or repeated.

Repeating is the expensive failure. Kill the worker after Stripe has moved the money but before the agent has written that down, restart naively, and the customer gets refunded twice. A repeated refund, email or deploy is an incident, not a retry.

Generic durable-execution engines already fix the crash part. What they do not do is understand model calls. To them a model call is an opaque function. They cannot tell you which prompt and model produced a decision, and they cannot take a real production run and re-run it on a new model to see whether it would have done better.

02How does durable execution work in KEEL?

The idea fits in one line: do not make the model deterministic, make the run deterministic by recording.

Every source of non-determinism passes through one boundary, ctx, and lands in an append-only, hash-chained event log before the workflow uses it. That covers model output, tool results, the clock and randomness. A run is then a pure fold over its log. Replay re-reads results instead of recomputing them.

Here is the refund step from the example agent in the repo. The workflow never calls Stripe directly. It asks ctx to run the effect, and ctx handles the log, the idempotency key and the retry rules.

examples/refund-agent/src/lib.rs
let args = RefundArgs { payment_intent: o.payment_intent.clone(), amount_cents, reason };
match ctx.effect("refund", &deps.refund, args).await {
    Ok(r) => {
        let v = serde_json::to_value(&r)?;
        refund = Some(r);
        v
    }
    Err(Error::EffectRejected { msg, .. }) => json!({ "error": msg }),
    Err(e) => return Err(e),
}

When a worker dies mid-refund, a fresh worker claims the run, replays the three recorded steps with zero model calls and zero tokens, finds the open refund intent, and retries it with the same idempotency key. Stripe returns the existing refund. The payments panel still shows one refund.

The fork-and-diff view: the run is forked at triage#1 onto another model with refunds simulated. Stripe still shows one refund, because the fork moved no money.
FIG. 1Forked at triage#1 onto an alternate model with effects simulated. The fork moves no money; the diff compares steps, reply, tokens, cost and time.

03What does each layer own?

KEEL touches three layers of the stack.

L4 Actions. Every tool declares an effect class and a tier. The coordinator runs intent, then prepare, then commit, and always reconciles before it retries an ambiguous result. Human approvals are durable signals: a run waiting two days for a manager holds no compute.

L2 Reasoning. Every model call is fingerprinted by model, messages, tools and parameters. If you change the prompt or the code and replay an old run, it fails loudly with a diff instead of silently diverging. Fork-and-diff turns real traffic into an evaluation set.

L0 Infrastructure. Stateless workers claim branches with fenced leases on Postgres. A worker that loses its lease cannot write, because the fence check sits in the same transaction as the append. Each step records tokens, cost and latency, so the economics of a run come straight out of its log.

04Why four tiers instead of “exactly-once”?

DECISION RECORD · DR-01LAYER L4
CONTEXTAgent tools call external APIs over a network. A timeout after sending a request leaves the outcome unknown, and some APIs offer no way to check.
OPTIONSClaim exactly-once for every tool · At-least-once with retries everywhere · Per-adapter tiers that state what the API can actually guarantee
CHOSENFour tiers. A: native idempotency keys. B: reserve, then confirm. C: queryable, reconcile before retry. D: blind, run at most once.
REJECTEDA blanket exactly-once claim is not achievable over a network, and at-least-once double-charges people. Tier D runs park as “in doubt” and escalate to a human instead of guessing.

The tiers live in the type system, so an adapter cannot be written without choosing one:

crates/keel-core/src/effect.rs
pub enum Tier {
    /// Native idempotency keys: retry with the same key until a definitive answer.
    A,
    /// Two-phase: prepare a pending resource tagged with the key, then commit or abort.
    B,
    /// Queryable: `reconcile(key)` before any retry.
    C,
    /// Blind: execute at most once; ambiguity becomes `EffectInDoubt`.
    D,
}

Exactly-once is not a property of a framework. It is a property of the API on the other end, and the runtime should say which one you have.

05How do we know it holds?

The whole system runs under deterministic simulation from the first week. Each seed injects crashes, zombie workers and lost responses, then checks five invariants, including “no duplicate effects” and “no lost runs”. The current suite has run 100,000 seeds with zero violations, and any failing seed reproduces exactly.

The demo shows the same thing at human scale: 20 runs with random kill points, 19 workers killed, 20 refunds. Those runs were recorded with KEEL's scripted model and a local Stripe-compatible fake. They will be re-recorded on Azure AI Foundry models and Stripe test mode before launch.

06What would we change, and what is next?

The honest limits today: long histories need a blob store and snapshots, and we have not yet published a per-step overhead benchmark (the target is under 5 ms). Code deployed in the middle of a running workflow needs versioned replay, which is not built yet.

The path from here is SDKs for Python and TypeScript over the one Rust core, a web console with a replay scrubber, then hardening: multi-tenancy, encryption and a one-million-seed nightly run. We will call it 1.0 when three design partners run production traffic for 30 days with no duplicate effects and no lost runs.

Run it yourself with one command: cargo run -p agensphere-keel-cli -- dev. The source is at github.com/Agensphere/keel, and the recorded demo is at keel.agensphere.com.

Building something like this?DESCRIBE A SYSTEM →