Prompt InsightsOpen Prompt Builder

Agent Reliability // Claude

Silent Agent Failures: 8 Failure Modes and 10 Prompts to Catch Them (2026)

Updated · 10 prompts

A silent agent failure is an AI agent run that ends in a success message while the work is wrong, incomplete, or has changed something it should not have. Nothing throws, the exit code is zero, the final answer reads well, and the only evidence is buried in the trace. This page lists the eight failure modes behind most of them, the checks that catch each one, and ten prompts for reviewing agent transcripts that you can run in Claude or any other capable model.

The reason these matter more than crashes: a crash gets fixed the same day because someone sees it. A silent failure gets copied into a report, merged into a codebase, or sent to a customer, and is found weeks later, if at all. In one field test we covered, 11 unguided agents run against a real product finished 3 tasks, and the more worrying number in runs like that is not the eight that stalled but how many of the three would have passed a careful review.

The eight silent failure modes

  1. Swallowed tool errors. A tool returns an error, a timeout, or a permission denial, and the agent carries on as if it had succeeded, often filling the gap with what the result probably would have said.
  2. Empty results read as answers. A search with a wrong filter, a misspelled name, or the wrong date range returns nothing, and the agent concludes that the thing does not exist.
  3. Unverified completion claims. The final report says tests pass, the file was created, or the record was updated, and no step in the trace checked it.
  4. Fabricated observations. The agent describes the contents of a page, file, or API response it never opened, usually because an earlier step failed and the model filled in plausible content.
  5. Instruction loss. A constraint from the system prompt stops being followed partway through a long run, most often after the context is summarized or compacted and the constraint does not survive the summary.
  6. Scope drift and quiet substitution. The agent solves an easier neighbouring task: it mocks the API instead of calling it, analyses a sample instead of the full dataset, or hardcodes the value a test expects.
  7. Lossy handoffs. In a multi-agent system a subagent hits a problem and its summary to the orchestrator leaves the caveat out, so every later decision is built on a cleaner story than the data supports. Our news desk's piece on multi-agent failure modes traces the same mechanism with fabricated statistics: one agent invents a number and the next treats it as a confident input.
  8. Side effects inside a successful run. The task completes, and somewhere in step 40 the agent also deleted a directory, overwrote a config, or sent a message nobody asked for.

Why output checks cannot see them

Every mode above produces a final answer that looks fine on its own. That is the definition of silent. So the checks have to read the trace: the sequence of tool calls, their arguments, and their raw results. Tracing is now standard in most agent observability tools, but a trace nobody reads catches nothing. What works in practice is three layers, from cheapest to most expensive.

Deterministic trace checks on every run. Plain code over the trace log: an error result not followed by a retry or an acknowledgement, a zero-result search followed by a negative claim, a claim of passing tests with no test command in the log, a step count more than about twice the median for the task type, a write outside the allowed paths. These cost nothing to run and catch the most common modes, 1, 2, 3, and 8, outright. The trace assertions prompt on this page writes them for your log format.

Model review on a sample. A second model reads the full transcript with one narrow question at a time: does each claim have a supporting tool result, were errors handled, were the original constraints still being followed at the end. Narrow questions matter. "Was this run good?" gets a yes most of the time. "Quote the tool result that supports each claim" does not.

Canary tasks with known answers. A fixed set of tasks, rerun on every change to the prompt, model, or tools, where the correct behaviour is known in advance, including tasks where the right answer is "the tool failed" or "this does not exist". Canaries catch the regressions that sampling would take weeks to surface.

Turning the prompts into a weekly routine

Run the deterministic checks on everything. Once a week, pull a random sample of about 20 transcripts per agent plus every flagged run, and run the claims-versus-evidence review on each. Any failure you confirm becomes two things: a new deterministic check if the pattern is mechanical, and a new canary task if it is not. After a month the checks catch most of what the reviews used to, and the review time goes to the failures nobody has seen yet.

For coding agents specifically, the Claude prompts for coding pack has a brief template that sets success commands and scope before a run starts and a diff audit that checks the agent's summary against what it changed. When the question is whether a change in your agent's success rate is real or noise, the significance and triage prompts in the Claude prompts for data analysis pack apply directly, and the Claude prompts for product management pack helps turn a failure review into a spec the team will act on. When a review format works for your team, save it as a template in Prompt Builder so every reviewer asks the same questions.

How to use these prompts

  1. 01

    Pick the prompt that matches your task

    Each prompt below targets one job. Choose the one closest to what you need instead of asking for everything at once.

  2. 02

    Replace every bracketed placeholder

    Swap [like this] for your real context: the product, the audience, the document, the constraints. Context is what separates a usable draft from generic output.

  3. 03

    Run it, then push back

    Read the first output critically and ask for a revision: tighter, more specific, a different angle, or with the weak assumptions named.

  4. 04

    Save the version that worked

    Once a prompt produces the output you want, keep it as a reusable template so the next run starts from your best version, not a blank box.

01

Claims versus evidence transcript review

The one review to run on every sampled transcript

You are reviewing the full transcript of an AI agent run, including every tool call and tool result. <task>[paste the original task]</task> <transcript>[paste transcript]</transcript> List every factual claim the agent makes in its final answer. For each claim, cite the step number of the tool result that supports it, or mark it UNSUPPORTED if no tool result in the transcript contains that information. Then give a verdict: PASS if every claim is supported, FAIL otherwise, with the single most serious unsupported claim quoted.

02

Swallowed error scan

Find tool errors the agent walked past

Scan this agent transcript for every tool result that contains an error, a non-zero exit code, a timeout, a permission denial, a rate limit message, or an empty body where content was expected: [paste transcript]. For each one, report the step, the error text, and what the agent did next: retried, changed approach, reported it, or continued as if the call had succeeded. Flag every case in the last category and explain how it could change the final answer.

03

Empty result audit

Catch 'no results' being turned into 'none exist'

In this transcript: [paste], find every search, query, list, or lookup call that returned zero results or an empty list. For each, check whether the call's arguments were plausible (right filter, right date range, right spelling, right scope) and whether the agent later stated that something does not exist, has not happened, or is not present because of that empty result. List each such conclusion and the alternative explanation the agent did not rule out.

04

Completion claim verifier

Check 'done' against what actually ran

The agent's final report says: [paste report]. The command and tool log is: [paste log]. For every statement of completion (tests pass, file created, email sent, record updated, deployed), find the command or tool call that proves it and quote its result. Mark each statement VERIFIED, CONTRADICTED (the log shows the opposite), or UNVERIFIED (no step checks it). Do not accept the agent's own later statements as evidence.

05

Instruction retention check

Find constraints that were dropped mid-run

Here are the constraints the agent was given at the start: [paste system prompt and task constraints]. Here is the transcript: [paste]. Number each constraint. For each, find the last step where the agent's behaviour shows it was still following it, and any step where it acted against it. Pay particular attention to the steps after any context summary or compaction event. Report the constraints that were followed throughout, the ones violated, and the step where each violation began.

06

Handoff loss check

Compare what a subagent found with what it reported

A subagent was asked: [paste subtask]. Its raw tool results were: [paste]. The summary it passed back to the orchestrator was: [paste summary]. List every caveat, error, partial result, or uncertainty present in the raw results that does not appear in the summary. For each, say whether the orchestrator's later decisions would plausibly have changed had it known.

07

Scope drift check

Did it solve the task you asked, or an easier one?

Original task: [paste]. Delivered result: [paste final output and a summary of the actions taken]. Restate the original task as a list of explicit requirements. For each requirement, say whether the result meets it fully, partially, or not at all. Then say whether the agent quietly substituted a simpler or adjacent task (a mock instead of a real call, a subset of the data, a hardcoded value, a skipped step) and quote where that happened.

08

Side effect inventory

List everything the run changed in the world

From this transcript: [paste], list every action that changed state outside the conversation: files written or deleted, commands that modify a system, database writes, messages or emails sent, API calls with POST, PUT, PATCH, or DELETE semantics, and anything installed. For each give the step, the target, and whether the task required it. Flag anything destructive or irreversible, and anything outside this allowed scope: [paste scope].

09

Write deterministic trace assertions

Turn review findings into checks that run on every trace

Our agent traces are stored as [format, e.g. JSONL with fields step, type, tool_name, args, result, exit_code]. Here are three example traces: [paste]. Write [language] functions that flag a trace when: a tool result contains an error and the next step is not a retry or an acknowledgement; a zero-result search is followed by a negative claim in the final answer; the final answer claims tests passed but no test command appears; the step count exceeds [2x] the median for this task type; or any write touches a path outside [allowed paths]. Include tests for each function using the example traces.

10

Build a canary task set

Detect regressions before users do

Our agent does [describe what it does and the tools it has]. Design 12 canary tasks with known correct answers that we can run on every prompt, model, or tool change. Include at least 3 tasks where a tool is expected to fail and the correct behaviour is to report it, 2 where the correct answer is that the information does not exist, 2 that require a constraint from the system prompt late in a long run, and 2 where the easy shortcut produces a wrong answer. For each give the task, the expected outcome, and the exact check that marks it passed.

READY TO ASCEND

Get the full pack by email

Curated prompt packs and prompt-engineering insight. No noise.

Questions

What is a silent agent failure?

A silent agent failure is an AI agent run that finishes without any error and reports success, while the work is wrong, incomplete, or has done something it should not have. Typical causes are a tool error the agent ignored, an empty search result read as proof that something does not exist, or a claim that tests passed when they were never run. They are dangerous because every signal a dashboard normally watches, exceptions, exit codes, and a well-formed final answer, looks healthy.

Why do normal evals miss silent failures?

Most evals score the final answer, and a silent failure usually produces a plausible final answer. The evidence is in the trace: the tool result that said permission denied, the search that returned nothing, the test command that never ran. Catching these needs checks over the steps, not only over the output, which is why the prompts on this page all take the full transcript as input.

How many agent transcripts should I review?

Run deterministic trace checks on every run, since they cost almost nothing. For model-based or human review, a practical starting point is a random sample of around 20 transcripts a week per agent, plus every run that a trace check flagged and every run above a risk threshold you set, such as any run that wrote to production or sent a message. Increase the sample after any change to the prompt, the model, or the tools.

Can an LLM judge catch silent failures reliably?

It catches many of them if you give it the full trace and a narrow question, such as matching each claim to a tool result, rather than asking whether the run was good. It is weaker at spotting what is missing than what is wrong. Before trusting a judge, have a person label about 50 transcripts, run the judge on the same 50, and look at the disagreements. Use a different model or at least a fresh context from the one that ran the agent.

How do I prevent silent failures in the agent's own prompt?

Three instructions do the most. Report every tool error verbatim in the final answer. Never state that something is done, exists, or does not exist without citing the tool result that shows it. When a check fails, stop and report instead of working around it. These do not make an agent reliable on their own, but they turn many silent failures into visible ones, which is what the review prompts need.

Related prompt packs