
Reviewing Agent Task History to Catch Missed or Duplicate Runs
October 5, 202610 min read2,263 words
Text: Priya Nambiar
Seven log fields expose whether your agent ran twice or skipped entirely.
Production agents run without anyone watching the screen in real time. When something breaks, there's no recording to rewind and no human memory to lean on. The task log is what's left, and it's either good enough to explain what happened or it isn't. Standard server logs confirm that an API call happened, but they can't explain why the model chose that action over another, which tools it called along the way, or what state it left behind when it finished. The miniorange audit trail guide states that these logs show the call, not the reasoning path behind it.
The deeper issue is structural. Traditional software follows a fixed call sequence, so a log of that sequence tells you almost everything. An AI agent is different: it combines application logic, external tools, retrieval systems, and a probabilistic model making choices at each step. Even when the input looks the same, the path it takes can shift from run to run. A log built to capture a fixed sequence will miss that dynamism entirely, recording the shell of an event without the decision that produced it. Operators often assume their logging setup tells them what their agent did. In practice, it tells them an agent did something, on some input, with some result, and leaves the rest to guesswork.
How missed runs and duplicate runs happen in production
Two failure modes account for most of the damage agents do in production: missed runs and duplicate runs. They are mechanically distinct, and conflating them leads to misreading the log.
Duplicate runs come from retries that don't know they're retries. A trigger fires more than once: an email arrives, gets processed, and the "unread" flag doesn't clear fast enough before the next polling cycle sees the same message and processes it again. Or a tool call succeeds, but the agent crashes before saving that fact, so a resumed workflow retries the call. Without an idempotency key attached to that call, the result is a duplicate payment, a duplicate ticket, or a duplicate deploy. The framework's response timeout was shorter than Jira's actual response time under load, so the agent retried and Jira created a second ticket. The agent's own dedup pass tried to merge the two, but that merge call timed out too, so it retried and produced a third ticket. One support email turned into three Jira tickets, and at each step the agent looked like it was behaving reasonably.
Missed runs come from the opposite problem: memory that isn't actually durable. An agent without persistent memory can repeat work it already did or skip work it wrongly believes is finished, because session memory tracks conversation, not execution. Saving chat history tells you what the agent discussed. It doesn't tell you which shell command ran, which email went out, which approval was granted, or whether a retry would step on a side effect that already happened.
Both failure modes trace back to the same infrastructure reality. Webhooks, queues, and every major delivery substrate the agent stack sits on are built so they deliver at least once, but not exactly once. Stripe retries webhooks for three days with exponential backoff. That's the correct design for a payments system that can't afford to drop an event, but most agent frameworks don't enforce idempotency constraints on top of it by default. The agent inherits a delivery guarantee built for reliability and turns it into a duplication risk, because nothing downstream is checking whether an action already happened before doing it again.
What a task log must contain to catch either failure mode
A task log that can't answer whether a run completed, which tools it invoked, and what each one returned is a record of activity without a way to audit it.
Seven fields turn a log into something a reviewer can use. A run_id is the correlation key you need for an entire execution chain. Without it, individual tool-call records float free of each other, and you can't group them into one coherent run. A user_id records who triggered the run, which matters for compliance and for attributing cost. The model and its parameters, meaning which model, what temperature, what max tokens, need to be logged because two runs against different model configurations aren't equivalent even when the prompt text is identical. Token usage, broken into input, output, and total, supports cost tracking and anomaly detection: a run burning an unusual number of tokens is often a retry loop in disguise. Latency in milliseconds establishes a performance baseline and feeds SLA monitoring, and a latency spike is frequently the leading indicator of a timeout-driven retry about to happen. The list of tool invocations, with every tool called, every parameter passed, and every value returned, is the one field that can directly expose a duplicate tool execution; nothing else in the log does that job. And a status field, covering success, failure, timeout, or partial completion, closes the loop, since partial completion is the status most likely to sit just upstream of a missed or duplicated follow-on run.
None of these fields are exotic. They let someone reconstruct, after the fact, what an agent did and why, beyond simply recording that something happened.
Reading the log to spot a duplicate run
A duplicate run leaves a fingerprint, and once a reviewer knows the shape of it, it's easy to spot: two run records with different run_ids but nearly identical trigger inputs, overlapping time windows, and the same tool invocations sitting side by side.
The primary signal is two run_id entries that point back to the same source event: the same task_id, the same user_id, the same triggering input hash. Add to that a pattern where both records show tool invocations with identical parameters executed within a short window of each other, and the case gets stronger. Status matters here too: if the first record shows "success" and the second shows anything at all, a side effect may have fired twice, and if both show "success," it almost certainly did. A latency anomaly on the first record, especially one consistent with a timeout, rounds out the picture, since it tells the reviewer the tool call likely completed on the server side even though the agent never got the acknowledgment before its own timeout fired.
The Jira cascade case lays this out cleanly. The log would show two tool invocations to the ticket-creation endpoint, identical parameters, status "success" on both. That pairing is the only evidence a retry happened at all, and the 2026 incident record traces it to a framework response timeout of 8 seconds against a Jira instance that was regularly taking 12 seconds to respond under load.
A filtering habit makes this kind of search fast. The Nylas CLI audit guide shows the pattern: nylas audit logs show --source claude-code --status error. Filtering by status error on a task known to have succeeded surfaces the retries that the error triggered, which is often the fastest way to find the duplicate before it's reported by someone downstream.
One distinction keeps this from turning into false alarms: legitimate parallel runs have different task_ids and different trigger inputs. Duplicates share a trigger input and differ only in the run_id generated at retry time. A reviewer checking for shared trigger inputs, not just shared timing, won't mistake normal concurrent work for a duplication bug.
Reading the log to spot a missed run
A missed run is defined by what's absent from the log, not by what's sitting in it, which makes it harder to catch than a duplicate without a deliberate review habit built around looking for gaps instead of entries.
Sequence checking is the most direct pattern. If runs are expected on a schedule or in response to a steady trigger stream, sorting by timestamp and scanning for gaps wider than the expected interval reveals a trigger that got consumed somewhere without ever producing a run record. Input-coverage checking works from the other direction: cross-reference the trigger source, an email inbox, a queue, a webhook log, against the trigger input hashes that actually appear in the task log. Any trigger event with no corresponding run_id is a missed run, full stop on the method, not on the consequence.
A third pattern catches a subtler case. A run that ends in a terminal error status, with no follow-on run and no human escalation logged afterward, is a candidate missed run: the agent stopped, and nothing ever picked the work back up.
A gap in these patterns is worth keeping in mind. Input and argument hashes preserve what was requested, but they say nothing about whether the target was still in the same state by the time the tool actually ran. A run can look complete in the log, with every expected field filled in, and still have run against stale state, so you get a miss in practice even though no run_id is technically absent.
The structural weaknesses in logs that make both failure modes harder to catch
Even a log with every one of the seven fields in place can still hide failures if the underlying audit architecture has gaps in how it generates and verifies those records.
One common gap is what amounts to a self-signed chain. If the same runtime component that mints run_ids also controls the step counter tracking them, a crashed or compromised runtime can drop a run and start a new sequence without leaving a hole you can detect. So if a sound implementation enforces a monotonic sequence scoped to the signing key across every run, any gap in step numbers becomes visible on its own, and no one has to notice a missing event by chance.
The second gap, and the more consequential one, involves who's doing the recording. If only the agent signs what it claims to have executed, there's no independent witness to what the tool on the other end actually did. A discrepancy between the agent's own record and the tool's own record can reveal a covert duplicate or a covert miss at that boundary, and that discrepancy can only exist if the tool side is countersigning. Logging more fields doesn't fix this. The fix is structural: an independent record on both sides of the call, not a richer record kept by only one side of it.
The check-before-act review pattern as a routine quality gate
The most effective use of task history is a regular check, before and after each run, that catches missed and duplicate runs before they compound into something larger.
Applied before a run: query the task log for any existing run_id tied to the same trigger input. If one exists and already completed successfully, the new run is a duplicate candidate and should be gated before it executes, not after. Applied after a run: confirm every expected tool invocation shows up in the record and that the status isn't "partial completion," since a partial completion with no follow-on run is the exact signature of a missed run sitting undiscovered. The overhead of running this check is one read against the task log. The alternative is finding out about a duplicate payment or a missed customer action days later, after it's already done damage.
Durable execution frameworks build this pattern into the infrastructure itself. Completed operations get recorded in a journal, and on recovery the framework replays that journal and skips whatever's already done. Log review in that setup becomes a way to verify what the framework already enforced, not a substitute for the enforcement.
Claude Code's checkpoint behavior shows what this looks like at the tool level. Checkpoints capture file state before each user prompt, and if you run /rewind, it restores an earlier checkpoint instead of re-running everything from the start, so the window of work that needs to be idempotent shrinks. The Nylas CLI audit system demonstrates the filtering half of the same discipline: time-window filters (--since, --until) combined with source and status filters let an operator reconstruct what an agent did in a given session and confirm nothing was duplicated or dropped. Neither tool replaces the habit of checking. Both make the habit fast enough that it actually happens.
What harness and runtime choice means for log quality
The richness of a task log is set largely by the harness and runtime an agent runs on, not by how carefully an operator wishes it were logging. A framework that treats structured, queryable logs as a first-class output gives a reviewer run_ids, tool invocation records, and status fields ready to query. But if a framework treats logging as an afterthought, you end up stitching together partial evidence from whatever free-text output happened to get captured.
This is why the seven fields and the review patterns described above matter as evaluation criteria, not just as practices to bolt onto an existing setup. Before adopting a harness for production use, the right question is whether it generates run_ids automatically, whether it records tool invocations with full parameters, and whether it exposes a query interface for time-window and status filtering of the kind the Nylas CLI example shows. A durable execution model that journals completed work, the way the frameworks described earlier do, closes much of the duplicate-run risk at the infrastructure layer before any manual review is needed. If you don't want to build journaling and countersigning infrastructure yourself, managed hosting built around these patterns closes the same gap. The review patterns in this piece work regardless of which harness sits underneath them. They work better, and faster, the less manual reconstruction they require.


