·8 min read

Replaying an agent should not call the model

A runnable example of agent replay, request matching, and the difference between reproducing a failure and testing a new prompt.

Imagine a release agent that proposes publishing after CI fails. You save the conversation, rerun the task, and watch it pass. The next attempt looks fine. Nobody has changed the code.

The convenient explanation is that the model had a moment. It worked the second time. Ship it.

But the second run may never have encountered the condition that broke the first one. CI now passes. A retrieved document changed. The model chose another tool. You have a successful new run and an unexplained old failure.

For reproducing the code around a model, I want replay to mean something fairly strict. Run the recorded version of that code against the observations it received, including the model's responses. Stop if the program requests something the recording cannot supply. A fresh model call belongs in a separate experiment.

"Replay" already describes several operations with different guarantees. Looking at a saved conversation, resuming a checkpoint, and executing code against recorded inputs answer different questions. A button with the same name does not make them interchangeable.

A bug small enough to keep

Here is a deliberately constructed example. An agent asks a tool for CI status on main. The surrounding code decides whether to propose publication. No actual model, CI service, or deployment system is involved.

The gate contains an ordinary Python bug:

def buggy_gate(result):
    return "publish-intent" if result["status"] else "stop"

The string "failed" is truthy. So is "passed". Both produce the same publication intent.

Our authored recording contains a model response requesting get_ci(ref="main"), followed by a tool result with status="failed". Another fixture simulates the world later, when that same branch has passing CI on a different commit. The old code proposes publication in both cases. Only the first case exposes the mistake.

Download the complete Python example. It uses the standard library, makes no network calls, and has no executor that could publish anything. Save it and run python3 replay_example.py.

Its first three results are:

historical replay: publish-intent on failed CI (bug reproduced)
changed world: publish-intent on passed CI (bug hidden)
code regression test: stop on failed CI (fix handles old input)

The regression test changes only the gate to require result["status"] == "passed". It keeps the old failed observation. That is a controlled code change against an old input. The historical replay keeps the original buggy gate too.

In this example, the program receives an authored model response. Freezing its response lets us investigate what the rest of the program did with that response. We have learned nothing about how often a real model would request this tool or whether a revised prompt would improve its judgment.

The output should tell you which of those questions the test can answer.

The request has to match too

Returning a stored response is the easy part. Returning it to the right request is where a green test can stop meaning anything.

Suppose you change the prompt to inspect a release branch. Your test intercepts a call to the same model endpoint and supplies the old response, which still asks for main. The test finishes. You conclude that the new prompt works. The new prompt never reached a model.

The example therefore checks the next request before supplying each recorded response:

expected_kind, expected_request, response = self.events[self.cursor]
if (kind, request) != (expected_kind, expected_request):
    raise ReplayDivergence(f"request changed at event {self.cursor}")

It checks the model request as well as the tool name and arguments. It consumes events in order and refuses an incomplete or exhausted recording. Changing the prompt stops at event zero; changing the tool arguments stops at event one. There is no fallback that quietly calls a live service when a fixture is missing.

HTTP testing tools already make this choice explicit. VCR.py's configuration makes request matching configurable. Its documented defaults include the method and URL components but omit the request body. For a model API, the body can contain the entire experiment. A match on the endpoint alone can therefore accept a different prompt. Its none record mode refuses new requests within the intercepted traffic. Local files and other uncaptured I/O still need separate controls.

A real recording needs a deliberate matching policy. You may want to ignore a transport request ID. You probably do not want to ignore a changed system prompt or tool argument. Every field you discard becomes a claim that it cannot affect the behavior you are testing.

An unmatched request is useful evidence. It tells you where the old recording stopped describing the new execution. Treating that disagreement as permission to improvise a response destroys the experiment.

Other people already solved parts of this

This is an old debugging problem with a new external dependency. In his 2005 discussion of event sourcing, Martin Fowler identifies both sides of the problem. Historical queries need historical answers. Replaying events must also avoid sending the same external updates again. His proposed gateway can remember query responses and suppress outgoing effects during replay.

The 2016 rr paper describes a more demanding implementation at the boundary between user-space programs and the kernel. It records inputs and sources of nondeterminism, then applies their recorded effects during execution. It does not need the filesystem to spontaneously recreate yesterday. Its fidelity depends on its recording boundary and supported execution conditions, including its approach to scheduling threads.

An agent fixture is a much smaller promise. It can hold the observations crossing model and tool interfaces steady while ordinary application code runs. It does not inherit rr's instruction-level guarantees because someone put "time travel" in the product name.

Current orchestration systems make the distinctions concrete. Temporal requires deterministic Workflow code, with nondeterministic work such as model calls and database queries in Activities. During replay, Workflow commands are checked against history. A successful history replay says something about that workflow's compatibility with recorded events; it does not demonstrate that a model would generate the same response today.

LangGraph's time-travel documentation describes a different operation. Results before the chosen checkpoint are saved, while nodes after it execute again. That includes model calls and API requests, which may return different results. This is useful for trying another continuation. It is not the frozen-input experiment above unless you add controls around those calls.

Read what the system executes again before deciding what a passing run proves.

But the model was the bug

This is the strongest objection, and it is correct. If the failure was a bad model choice, replaying that choice cannot show that a new model would do better. You can reproduce how the program accepted it, where validation failed, or which action followed. You cannot evaluate new reasoning by substituting old reasoning.

For a prompt or model change, preserve the old prefix, mark the point where the experiment changes, and call the candidate model there. Keep the relevant tool world controlled if you want to isolate that change. Run enough trials for the reliability claim you intend to make. A single successful continuation tells you that one continuation succeeded.

For a code change, the frozen recording can become a regression case, as it does in the example. For a changed database or service, choose explicitly whether to restore old observations or test current behavior. Save those results separately. Otherwise a code fix, a different model answer, and a repaired external dependency can all arrive under the same "passed" label.

There is no need to build a general replay platform to fix a one-line condition. A captured response and a unit test may be plenty. Use a recording when the sequence and surrounding state matter enough that a single fixture loses the explanation.

An HTTP recording misses a local file read. A serialized list of tool results may miss the order in which concurrent tasks changed shared state. A redacted value can change a branch condition. Keep the original executable and dependency versions if the question is what the original program did, and document whatever the recorder leaves live. Missing inputs are a limit on the claim, not an invitation to declare the run deterministic anyway.

Keep replay isolated from production because the intended action may be exactly the bug you are reproducing. Record the proposed write, or route it through an inert replacement. Debugging a duplicate email by sending another email would be an unusually literal interpretation of reproduction.

I want a failed agent run to arrive with a small executable case: the code version, the requests it made, the observations it received, and a way to run it without touching the real account. Then I can change one thing and see what follows.

Before celebrating a successful rerun, make the old failure happen on purpose.