AGENSPHERE/ JOURNAL
← JOURNAL
J-002DECISION3 MIN READ

Every agent tool needs an effect class before it ships

If you cannot say what happens when a tool call runs twice, the tool is not ready to give to a model. Four questions, asked before an AI agent tool ships, decide how the runtime is allowed to retry it.

IN SHORT
  • An effect class describes what happens if a tool call is executed again: pure, idempotent, compensable or irreversible.
  • Agents retry more than normal code does, because models re-plan, workers crash and networks time out. A tool without a declared class will eventually run twice.
  • Declare the class in code next to the tool, and let the runtime pick the retry policy from it. Irreversible tools get a human approval or an at-most-once rule.
  • A timeout after sending a request is not a failure. It is an unknown outcome, and it must be reconciled before any retry.

When a model gets a tool, it gets the power to change something outside your system. Most teams describe a tool by its inputs and outputs. That is enough for the model to call it. It is not enough for the runtime to call it safely a second time.

01What is an effect class?

An effect class is one word that answers: what happens if this call runs again with the same arguments?

ClassRunning it twiceExample
PureNothing changesLook up an order, search documents
IdempotentSame result, if you pass the same keyCreate a refund with an idempotency key
CompensableA second effect, but it can be undoneReserve inventory, create a draft
IrreversibleA second effect that cannot be undoneSend an email, wire money, deploy

The class belongs to the tool, not to the call. You decide it once, in code, before the tool is exposed to a model.

02Why do agents need this more than normal code?

A normal request handler runs a function once, top to bottom. An agent does not. Three things make repeats routine:

  1. The model re-plans. It sees a tool result it does not like and calls the tool again, sometimes with the same arguments.
  2. Workers crash. Long runs outlive deploys, autoscaling and out-of-memory kills. Whatever was in flight gets retried by someone.
  3. Networks time out after sending. The request reached the provider, the response did not reach you. You do not know if it happened.

Every one of those turns into a duplicate side effect unless the runtime knows which tools are safe to repeat.

03The decision

DECISION RECORD · DR-02LAYER L4
CONTEXTAgent tools were registered with a name, a description and a JSON schema. Retry behaviour was decided ad hoc inside each tool, and two of them had no protection against running twice.
OPTIONSLeave retries to each tool's own code · Never retry any tool automatically · Declare an effect class per tool and derive the retry policy from it
CHOSENEvery tool declares its effect class at registration. The runtime refuses to register a tool without one, and picks retries, keys and approvals from the class.
REJECTEDPer-tool retry code drifts and nobody reviews it. Never retrying makes pure lookups fragile for no reason. A declared class is one line, reviewable in a PR, and enforceable by the runtime.

In practice the registration looks like this:

agent/tools.py
@tool(effect="pure")
def lookup_order(order_id: str) -> Order: ...

@tool(effect="idempotent", key=lambda a: f"refund:{a.order_id}:{a.amount_cents}")
def issue_refund(order_id: str, amount_cents: int) -> Refund: ...

@tool(effect="irreversible", approval="manager", over=lambda a: a.amount_cents > 50_000)
def send_customer_email(to: str, body: str) -> None: ...

04How does the class drive the retry policy?

  • Pure: retry freely on any error.
  • Idempotent: retry with the same key until the provider gives a definitive answer. Generate the key from the intent (order and amount), not from a random UUID per attempt, or the key protects nothing.
  • Compensable: retry only after checking whether the first attempt landed. If it did, keep it. If the run is abandoned, run the compensation.
  • Irreversible: run at most once. If the outcome is unknown, stop and ask a human. Large or sensitive calls wait for an approval signal before they run at all.

05What about a timeout?

A timeout after the request was sent is the case teams get wrong most often. It is tempting to treat it as a failure and retry. It is not a failure. It is an unknown.

The rule is simple: on an ambiguous result, reconcile before you retry. Ask the provider whether an effect with this key exists. If the provider cannot answer and the tool is irreversible, park the run as “in doubt” and escalate.

A tool that has never been asked “what if this runs twice?” has already decided the answer, and it is usually the wrong one.

06A checklist before a tool ships

  • The effect class is declared in code, next to the tool.
  • Idempotent tools derive their key from the intent, and the provider actually honours it.
  • Irreversible tools have an approval rule or an at-most-once guard.
  • Ambiguous errors are separated from definitive ones in the tool's error type.
  • There is a test that runs the tool twice with the same arguments and checks there is one effect.

This is the same rule our durable execution runtime, KEEL, enforces in its type system. You do not need a runtime to start: a decorator, a review rule and one test per tool get most of the value.

Building something like this?DESCRIBE A SYSTEM →