AGENSPHERE/ JOURNAL
← JOURNAL
J-004NOTE3 MIN READ

Treat a prompt change like a code change: ship it with an eval

A prompt is code that runs on someone else's computer and changes behaviour without a compiler error. LLM evaluation gives it a test set, a score and a review gate, so model upgrades stop being surprises.

IN SHORT
  • An eval is a fixed set of inputs with a way to score the outputs, run every time a prompt, model or tool definition changes.
  • Start with 30 to 100 real cases, mostly from production logs and past failures, and score with deterministic checks first. Use a model as judge only for what code cannot check.
  • Run the eval in CI and block merges that drop the score, exactly as a failing unit test would.
  • Version the prompt, the model name and the parameters together, and log all three with every call, so any output can be traced to what produced it.

A prompt edit looks like a copy change. It is a behaviour change. It can fix the ticket in front of you and quietly break five cases you are not looking at. A model upgrade does the same thing to every prompt at once.

The fix is the same one software already uses: a test suite that runs on every change. For model-backed features that suite is called an eval.

01What is an LLM evaluation, concretely?

An eval is three things:

  1. Cases: a fixed list of inputs, each with what a good output must satisfy.
  2. Scorers: functions that turn an output into a pass, a fail or a number.
  3. A threshold: the score below which the change does not ship.

It is not a benchmark leaderboard and it does not need a platform. A folder of JSON files and a script is a real eval.

evals/test_triage.py
CASES = load_jsonl("evals/triage_cases.jsonl")   # real tickets, labelled

def score(case, out) -> dict:
    return {
        "valid_json": is_valid(out, TriageResult),
        "right_action": out.action == case["expected_action"],
        "no_refund_over_total": out.refund_cents <= case["order_total_cents"],
        "cites_order_id": case["order_id"] in out.reply,
    }

def test_triage_prompt():
    results = [score(c, run_triage(c["ticket"])) for c in CASES]
    assert rate(results, "valid_json") == 1.0
    assert rate(results, "no_refund_over_total") == 1.0     # a hard rule, never trade it
    assert rate(results, "right_action") >= 0.92

02Where do the cases come from?

Not from your imagination. The cases that matter are the ones that already happened.

  • Production logs. Sample real inputs across the categories you care about, with personal data removed.
  • Every bug report. When a model output is wrong in production, its input becomes a case before the fix is written. This is the single habit that compounds most.
  • Edge cases on purpose. Empty input, very long input, another language, an adversarial instruction inside a document.

Thirty good cases beat a thousand synthetic ones. Grow the set as failures arrive.

03How should outputs be scored?

Use the cheapest scorer that is correct.

Deterministic checks first. Is it valid JSON for the schema? Does it call the right tool? Is the number within the allowed range? Does it contain the required ID? These are exact, fast and free.

Reference comparisons second. For extraction or classification, compare to a labelled answer.

A model as judge last. For tone, helpfulness or faithfulness to a source, a second model with a written rubric can score outputs. Keep the rubric specific, and check the judge against a sample you labelled by hand before you trust it.

Separate hard rules from quality scores. A refund above the order total is never acceptable at any accuracy level, so it gets its own assertion at 100%.

04When does it run?

On every change that can alter behaviour: the prompt text, the model name or version, temperature and other parameters, tool definitions, retrieval settings. Put it in CI and block the merge on a drop, the way a failing unit test blocks it.

If a prompt can change in production without a test running, you do not have a prompt. You have an untested deploy.

05Version everything that shaped the output

Store the prompt in the repo, not in a dashboard. Log the prompt version, the model identifier and the parameters with every call. When someone asks why the system said something last Tuesday, the answer should be a commit hash, not a guess.

This also makes model upgrades boring, which is the goal. Run the eval on the new model, read the diff of failing cases, adjust the prompt if needed, and ship with a number instead of a feeling.

06Where this goes next

Logs plus an eval give you one more tool: replay. Take real recorded runs and re-run them against a new prompt or model, then compare. That is the fork-and-diff idea behind KEEL, and it turns production traffic into the best eval set you will ever have.

Building something like this?DESCRIBE A SYSTEM →