AI agent testing is where most UK businesses deploying autonomous workflows discover that their existing quality assurance practices do not apply. The agent works in the demo. It works when the developer tries it. It works for a fortnight in production. Then it sends an email to the wrong distribution list, or creates a duplicate invoice, or confidently answers a customer question using a figure it invented — and nobody can say why, because nothing was recorded and nothing failed in a way that raised an alarm.
The reason is structural rather than a matter of insufficient care. Traditional software fails deterministically and loudly: the same input produces the same wrong output every time, and it usually throws an exception you can catch. An agent fails probabilistically and quietly, produces a different result on each run, and when it is wrong it is wrong in fluent, well-formatted, entirely plausible prose. This guide covers what to do about that: why per-step reliability compounds into end-to-end unreliability faster than anyone expects, a failure taxonomy worth testing against, why guardrails belong in code rather than in a prompt, staged rollout with human gates that actually gate, monitoring for drift that arrives without a deployment, and the incident response and reversibility work to complete before an agent is allowed to take a real action.
Why standard QA practices are insufficient
Six properties of agent systems break assumptions that conventional testing depends on. It is worth being precise about them, because each one implies a different change to how you test.
Non-determinism. The same input can produce different output on different runs. A test that asserts equality against an expected string will pass and fail unpredictably, which teams typically respond to by deleting the test. The correct response is to run each case several times and assert on a pass rate rather than a single outcome.
No single correct answer. For open-ended work there is often no unique right output, only a range of acceptable ones. Assertions have to move from equality to properties: does the output contain the required fields, avoid the prohibited claims, cite something real, stay within scope. This is graded assessment rather than pass or fail.
Errors compound across steps. An agent completing a ten-step task at ninety-five per cent per-step reliability succeeds end to end about sixty per cent of the time. This arithmetic is the single most important thing to understand about multi-step autonomy and it is covered in the chart below.
Failure is silent and plausible. Conventional software announces its failures with stack traces and error codes. An agent that fabricates a reference number returns a well-formed response containing a reference number. There is nothing to catch, which means detection has to be designed rather than inherited.
The input space is unbounded. You cannot enumerate natural language. Traditional notions of coverage do not transfer, so the question shifts from “have we tested every path” to “is our sample representative of what production actually sends, including the awkward cases”.
Behaviour changes without a deployment. An underlying model version changes, a retrieval corpus is updated, an upstream system alters a response format — and your system behaves differently while your code is identical. Nothing in conventional release management accounts for a system that drifts when you did not touch it.
To which add a seventh consideration that is not about testing but changes its stakes: agents take actions. A chatbot that answers badly has produced a bad answer. An agent with tools that answers badly may have sent something, changed something, or spent something. The blast radius is what justifies the additional rigour.
Before writing any tests, instrument the agent so that every run records a full trace: the input, each step, each tool call with its arguments, each tool response, and the final output. Do this first, because every subsequent activity depends on it — you cannot build an evaluation set without real examples, cannot debug a failure you did not record, and cannot detect drift without a baseline. Teams that add tracing after their first production incident invariably discover that the incident they most wanted to understand is the one they have no data for.
How per-step reliability compounds
The chart below shows end-to-end success rates for multi-step tasks where each individual step is ninety-five per cent reliable — a figure most teams would consider good. The point is not that ninety-five per cent is achievable or not; it is that autonomy length is a reliability multiplier working against you.
Three consequences follow, and they shape the whole design rather than just the testing.
The first is that shorter is more reliable. An agent asked to complete a twenty-step task autonomously will fail most of the time even with individually excellent steps. Decomposing that into three shorter agents with checkpoints between them, or into a deterministic workflow that calls the model at specific decision points, converts a compounding problem into several independent ones. Most production reliability gains come from reducing autonomous chain length rather than from improving prompts.
The second is that per-step measurement is essential. An end-to-end test tells you a task failed; it does not tell you which step failed or how often. Without step-level instrumentation you are left improving the whole system by intuition, and the arithmetic above means small per-step improvements produce disproportionate end-to-end gains — raising each step from ninety-five to ninety-eight per cent takes a ten-step task from sixty to eighty-two per cent.
The third is that the acceptable chain length depends on what happens when it fails. A ten-step task with a sixty per cent success rate is perfectly deployable if failure means a human picks it up, and unacceptable if failure means a customer receives something wrong. The reliability target is a function of consequence, not a universal standard.
Agent reliability in UK deployments — the numbers
The figures below reflect what we observe across UK organisations of 20 to 500 staff putting LLM-based agents into production. They describe teams building deliberately rather than experimenting, which makes the gaps more notable.
The second figure is the one that separates teams who will improve from teams who will plateau. Without a curated evaluation set, every prompt change is an act of faith: somebody tries three examples by hand, the output looks better, and the change ships. Whether it improved the system across the range of real inputs is unknown, and regressions are discovered by users. A set of fifty to two hundred cases with graded assertions, run automatically on every change, is what makes iteration cumulative rather than a random walk.
The third figure explains why so many agent incidents are never fully understood. If only the final output is logged, a wrong answer is a mystery: you cannot see which step went wrong, what the tool returned, or whether retrieval found anything. Teams then debug by trying to reproduce the failure, which non-determinism makes unreliable.
The fourth figure is the one that should be addressed before anything else in this list. Fourteen per cent can stop an agent immediately without shipping code. For a system authorised to take real actions, the ability to halt it in seconds — a configuration flag, a feature switch, a queue pause, tested and known to work — is not an advanced capability. It is the minimum condition for letting it run at all.
A failure taxonomy worth testing against
Agents fail in recognisable categories, and having the list is most of what makes an adversarial test set possible. The grid below groups the modes we see in production. The badges reflect the combination of how often each occurs and how damaging it is when the agent has tools that take real actions.
Three of these deserve expanding because they are routinely missed in test design.
Non-idempotent retries. If a tool call times out and the agent retries, and the underlying action actually succeeded the first time, you now have two invoices, two emails or two payments. This is not an AI problem — it is a distributed systems problem that agents encounter constantly because they retry liberally. The remedy is idempotency keys on every action-taking tool, enforced server-side, and it belongs in the tool layer rather than in instructions.
Instructions arriving in content. An agent that reads external material — an inbound email, a web page, a supplier document, a support ticket — is processing text that someone else wrote, and that text can contain instructions. If the agent also has tools, the combination is genuinely dangerous: content it was asked to summarise can attempt to direct what it does. Treat all retrieved and inbound content as untrusted data rather than as instructions, keep the authorisation decisions outside the model, and include injection attempts in your adversarial test set. Adversarial testing of this kind has a good deal in common with the mindset covered in our guide to what happens during a penetration test.
Answering with nothing retrieved. The most common grounding failure is not retrieving the wrong thing; it is retrieving nothing useful and generating an answer regardless. This is straightforward to test for and straightforward to fix — assert that a retrieval step returned results above a relevance threshold before the generation step is permitted to run — and it is missed because the output looks entirely normal.
Conventional QA against agent evaluation
The comparison below highlights the agent evaluation column, because for a non-deterministic system the conventional approach does not merely underperform — it produces tests that are deleted within a month for being flaky. This is not an argument against conventional testing for the deterministic parts of your system, which should be tested conventionally and thoroughly. It concerns the model-driven components specifically.
Conventional QA
Built for deterministic systems
Agent evaluation
Built for probabilistic behaviour
The row that changes team behaviour most is the sixth. A prompt is a load-bearing part of the system, and editing one can alter behaviour as substantially as editing code. Treating prompts as configuration that anybody can adjust without review or regression testing is how agents quietly get worse. Version them, review them, and run the evaluation suite against changes — the same discipline you would apply to a function that decides what happens to customer data.
The second row is worth being concrete about. Running each evaluation case five times and recording how many passed gives you both a pass rate and a variance signal. High variance on a case is itself informative: it usually means the task is underspecified or the step is at the edge of what the model handles reliably, both of which are design findings rather than test failures.
On rubric grading: using a model to grade another model’s output is practical and widely done, and it needs validating rather than trusting. Label a sample of cases by hand, compare the automated grades against those labels, and measure agreement. If the grader disagrees with humans on a fifth of cases, its verdicts carry that error into every subsequent decision you make from them.
Guardrails belong in code, not in the prompt
This is the single most consequential design principle in agent reliability, and it is violated almost universally in first implementations. If a constraint matters, it must be enforced somewhere the model cannot override.
Consider an instruction in a system prompt: do not send emails to more than five recipients. That is a request. It will be followed most of the time. It will not be followed when the input is unusual, when the context has grown long enough that earlier instructions lose force, or when a document the agent read contains text pulling it in another direction. A probabilistic system given a rule follows it probabilistically.
The same constraint implemented in the tool — the send function rejects any call with more than five recipients and returns an error the agent must handle — is enforced. It holds on the unusual input, at the end of a long run, and in the presence of hostile content, because it does not depend on the model choosing to comply.
What belongs in code
Anything whose violation you could not accept. Value caps on financial actions. Recipient limits. An allowlist of which records the agent may modify. Rate limits on actions per hour. Idempotency keys so retries cannot duplicate effects. Requirements that a retrieval step returned something before generation proceeds. A hard maximum on step count so a loop terminates. Schema validation on every tool argument. Whether the action is reversible, and refusal of irreversible actions above a threshold without human approval.
What the prompt is for
Everything that shapes quality rather than bounding risk: tone, format, what good output looks like, how to approach the task, when to ask rather than assume. Prompts are excellent at steering behaviour and unsuitable as a security boundary. The useful mental test is to ask what happens if the model ignores this instruction entirely — if the answer is unacceptable, the instruction is in the wrong place.
This distinction also simplifies testing considerably. Code-level guardrails can be unit tested conventionally and deterministically: assert that the tool rejects six recipients, every time, with no sampling required. That moves a meaningful share of your reliability surface out of the probabilistic domain and into the domain where ordinary software engineering already works well.
Staged rollout with gates that actually gate
The sequence below takes an agent from first build to bounded autonomy over roughly sixteen weeks. The stages exist to buy evidence: each one produces data about real behaviour on real inputs before the consequences of a mistake increase.
Shadow mode in weeks six to nine is the stage to defend when the schedule compresses. It is the only point at which you can compare agent output against human decisions on identical real inputs at zero risk, and the comparison usually reveals both that the agent is better than expected on routine cases and worse than expected on a specific category nobody anticipated. Skipping it means discovering that category in stage one, with a human catching it, which is survivable — or in stage three, which is not.
Note that stage three is bounded autonomy rather than autonomy. The distinction is not cautious framing: the limits are what make the risk quantifiable. An agent that can take any action of any value is a system whose worst case you cannot state. An agent restricted to a defined action set below a value cap at a bounded rate has a worst case you can write on one line, which is what allows anyone to approve it.
What reliability work costs in the UK
The figures below are indicative UK engineering costs for 2026, excluding VAT, for building and operating the reliability apparatus around an agent of moderate scope. They cover the testing and operations work specifically rather than building the agent itself, which is costed separately in our guide to AI software development cost in the UK.
| Reliability component | Indicative UK cost | Frequency | Note |
|---|---|---|---|
| Tracing and observability instrumentation | £3,000–9,000 | One-off, then maintained | Prerequisite for everything else; cheapest done first |
| Evaluation set construction and harness | £5,000–15,000 | One-off, extended continuously | 50–200 cases with graded assertions, built from real traffic |
| Guardrails in code and kill switch | £4,000–12,000 | One-off | Tool-layer enforcement, idempotency, caps, limits, tested halt |
| Human review capacity during staged rollout | £2,500–8,000 | Across 8–10 weeks | Internal time approving and correcting; the largest hidden cost |
| Ongoing monitoring and sampled review | £800–3,000 | Per month, indefinitely | Canary runs, drift watch, continuous sampling of live output |
The final row is the one that changes business cases and is almost always omitted from them. Agent reliability is an operating cost rather than a project cost: the canary suite runs forever, somebody reviews a sample of output forever, and drift arrives whether or not there is budget to detect it. A proposal that presents agent deployment as a fixed build cost with no recurring line has not accounted for the property that distinguishes these systems from conventional software.
The human review row is the largest surprise for most organisations, because it is internal time rather than an invoice. Someone has to approve every action for several weeks in stage one, and that person needs to be competent at the task being automated — which usually means they are the person whose time the agent was meant to free. The benefit arrives after the review period, not during it, and expecting otherwise causes rollouts to be cut short at exactly the wrong point.
Set against those costs: the tracing and guardrail lines are largely one-off and they are what make the difference between an incident you can explain and one you cannot. Of the five, tracing is the one to fund first and the one with the least visible immediate return, which is a familiar pattern for anything that only proves its worth when something goes wrong.
The number that decides whether you can operate an agent at all
Of everything in this guide, one capability is not an improvement but a precondition: the ability to stop the agent immediately, without shipping code, having previously verified that the mechanism works.
Fourteen per cent is a low number for something so cheap. A feature flag, a configuration value, a paused queue — any of these will do, provided somebody has actually used it in anger at least once and confirmed that in-flight work stops safely. The common failure is a switch that exists in principle and has never been exercised, so nobody knows whether it halts mid-run tasks or lets them complete.
The reason it matters more for agents than for conventional services is the combination of silent failure and real actions. A misbehaving web service produces errors somebody notices. A misbehaving agent produces plausible output and keeps going, which means the interval between the problem starting and somebody deciding to intervene can be hours. If the response to that decision is a code change, a review, a build and a deploy, the interval extends further while the agent continues acting.
Alongside the switch, two related capabilities belong in the same category of precondition. An action log detailed enough to reverse what was done — not just that an email was sent but to whom, with what content, and under which run identifier — because containment usually requires undoing rather than merely stopping. And a named person who is permitted to pull the switch without seeking approval, because a control that requires a meeting is not an emergency control.
Monitoring for drift that arrives without a deployment
Agent systems degrade in ways that conventional monitoring does not surface. Availability is fine, error rates are normal, latency is acceptable, and the answers have quietly become worse. Detecting that requires watching different things.
Run the evaluation suite on a schedule, not only on change
This is the single most effective drift control. The same suite that gates your changes runs nightly or weekly against production configuration, and its pass rate is tracked over time. Because the cases and assertions are fixed, a fall in pass rate with no deployment means something outside your code changed — a model version, a retrieval corpus, an upstream API. Without this you learn about drift from users.
Watch leading indicators, not just outcomes
Several signals move before quality complaints arrive. Average step count per task rising suggests the agent is working harder for the same result. Retry rates rising suggests tools or arguments are going wrong more often. Escalation rate moving in either direction is informative — falling may mean over-confidence, rising may mean genuine degradation. Refusal rate rising suggests over-caution. Cost per completed task rising is often the earliest numerical signal of all, because inefficiency shows up in consumption before it shows up in quality.
Sample and review live output, permanently
A percentage of completed tasks reviewed by a competent human, continuously, at a rate proportionate to consequence. This is the only control that reliably catches the failure category nobody anticipated, because evaluation sets can only test for modes you have thought of. Treat it as an operating function with an owner rather than a launch activity that tapers off.
Capture corrections as evaluation cases
Every time a human corrects the agent, that is a test case arriving free. Routing corrections back into the regression set is what makes the system improve over time rather than merely being monitored. Teams that do this find their evaluation set grows into something genuinely representative within a few months; teams that do not find their set stays as it was written on day one.
One caution on aggregate metrics: a stable overall pass rate can conceal a category collapsing, if that category is a small share of volume. Segment the evaluation results by task type and input characteristic, because the failure that matters commercially is frequently in a minority segment — the unusual request, the largest customer, the edge case that carries disproportionate value.
Benchmarks — reliability practice against what we find
The figures below show how often each practice is in place across UK organisations running LLM agents in production. Readings are generous throughout: a practice counts as present if it exists in any form.
Adoption of agent reliability practices in UK deployments
Ninety-four per cent against a field of teens and twenties is the starkest gap in this series of benchmarks. Building an agent that works has become genuinely straightforward; the tooling is good and the models are capable. Building one you can operate, debug, bound and correct has not, and almost none of that apparatus is present in the typical deployment.
The eleven per cent running scheduled evaluation is the most consequential omission, because it is the only practice that detects the failure mode unique to these systems — degradation with no change on your side. Nearly nine in ten deployments would learn about a model behaviour change from a customer.
The twenty-one per cent with idempotency is the one most likely to produce a concrete, embarrassing incident. Agents retry. Retries duplicate. Without idempotency keys enforced server-side, the question is not whether a duplicate action will occur but when, and whether it will be an email or a payment.
Agent reliability readiness — where most UK teams sit
Combining the assessment areas gives an indication of how ready a team is to let an agent take real actions unattended. The gauge reflects a first review of a UK organisation with a working agent in or approaching production.
Twenty-eight is the lowest figure in this series of benchmarks, and the reason is that the discipline is genuinely new. Organisations have had decades to develop practices around deterministic software and a very short time to develop them for probabilistic systems that act. The score reflects immaturity of practice rather than carelessness, and it is falling behind capability rather than standing still — agents are becoming more capable faster than the surrounding operational apparatus is being built.
The composition is consistent. Capability scores high: the agent does the task. Observability scores low. Evaluation scores low. Guardrail placement scores low, because the intuitive approach is to write the rule into the prompt. Containment and reversibility score lowest, because they only matter after something has gone wrong and most teams have not yet had that experience.
The encouraging part is how much of the gap is a few weeks of unglamorous engineering rather than research. Tracing, an evaluation harness, guardrails moved into the tool layer, idempotency keys, a tested switch, a scheduled canary run and a sampling routine are all well-understood work. None of it requires anything novel; it requires deciding that reliability is part of the build rather than something to add once the demo has been approved.
Incident response, reversibility and the UK compliance angle
Before an agent takes its first real action, three questions should have written answers. What stops it, what undoes what it did, and who is accountable for the decisions it made.
Containment and reversal
Stopping is the kill switch discussed above. Reversal is harder and needs designing rather than improvising. For each action the agent can take, establish in advance whether it is reversible, how, and by whom — a database write is usually reversible from a log, an email to a customer is not, a payment sits somewhere between. Where an action cannot be reversed, that is an argument for keeping a human gate on it permanently rather than something to discover during an incident.
The action log is what makes reversal possible, and it needs enough detail to reconstruct what happened: run identifier, timestamp, action type, full arguments, the result returned, and which input triggered the run. Most incident reviews founder not because the agent misbehaved mysteriously but because the record of what it actually did is too thin to work from.
Accountability for agent decisions
Someone in the organisation is accountable for what the agent does, and it is worth naming that person explicitly before rather than after. This matters practically — incidents need an owner — and it matters because the accountability does not transfer to the software or to the model provider. An agent acting on your behalf is your action.
Where UK data protection law becomes relevant
Two provisions are worth understanding, and this is an area where taking proper advice is sensible rather than relying on a guide. First, UK GDPR restricts decisions based solely on automated processing that produce legal effects for an individual or similarly significantly affect them, and where such processing is permitted it carries safeguards including the ability to obtain human intervention. If your agent is declining applications, setting prices for individuals, or making decisions with real consequences for a person without meaningful human involvement, that is the provision to look at.
Second, where processing is likely to result in a high risk to individuals, a data protection impact assessment is required. Novel technology applied to personal data at scale is a recognised trigger, so an agent processing customer data will frequently warrant one. The ICO publishes guidance on AI and data protection that is a reasonable starting point.
Worth noting how neatly the reliability controls and the compliance position align. Meaningful human involvement at the point of consequential decisions is both the strongest reliability control available and the thing the automated-decision provisions are concerned with. Step-level tracing is both how you debug and how you evidence what happened and why. Organisations that build the reliability apparatus properly find much of the compliance work already done, which is an unusually happy coincidence.
One further practical point: the traces themselves frequently contain personal data, which means they fall within your retention and access obligations like any other record. Decide their retention period deliberately rather than keeping everything indefinitely because storage is cheap — the same reasoning we set out in our guide to backup retention policy.
Common mistakes in agent deployment
The errors below recur across UK agent projects. Most are consequences of applying conventional software instincts to a system that does not share conventional properties.
- Putting guardrails in the prompt. A probabilistic system given a rule follows it probabilistically. If violating a constraint would be unacceptable, it belongs in the tool layer where the model cannot override it.
- Asserting equality in tests. Non-determinism makes exact-match assertions flaky, teams delete the flaky tests, and the system ends up untested. Run each case several times and assert on a pass rate against properties.
- Deploying without step-level tracing. Only 31 per cent record full traces, which means most teams cannot debug their first serious incident. Tracing is the prerequisite for everything else and the cheapest thing to do first.
- Building long autonomous chains. Ten steps at 95 per cent each succeeds 60 per cent of the time. Shortening chains and adding checkpoints improves reliability more than prompt refinement ever will.
- Treating prompts as configuration rather than code. A prompt edit can change behaviour as much as a code change. Version them, review them, and gate them behind the evaluation suite.
- Skipping shadow mode. It is the only stage where you can compare agent output against human decisions on real inputs at zero risk, and it is skipped because it produces no visible benefit while running.
- No idempotency on action-taking tools. Agents retry, retries duplicate, and only 21 per cent guard against it. The question is when, not whether, a duplicate action occurs.
- Trusting retrieved and inbound content as instructions. Anything the agent reads was written by someone else and may attempt to direct it. Treat external content as untrusted data and include injection attempts in the test set.
- Running evaluations only on change. Drift arrives without a deployment. Without a scheduled canary suite, nine deployments in ten learn about a behaviour change from a customer.
- Budgeting reliability as a project cost. Canary runs, sampled review and drift watch continue indefinitely. A business case with no recurring line has not accounted for what makes these systems different.
Be careful with the demo-to-production gap, because it is unusually wide for agents and unusually convincing. A demonstration is a small number of favourable inputs handled by someone who knows what the system does well, and agents demonstrate extremely well — fluent, fast and apparently competent. That performance says very little about behaviour across the real input distribution, on the unusual fifth of cases, at the end of a long run, or when content tries to misdirect it. Treat an impressive demo as evidence that the capability exists and as no evidence at all about reliability, and resist commitments to stakeholders made in the week after one.
What this looks like in practice
A UK B2B distributor with 120 staff built an agent to handle inbound purchase order emails: read the message and attachment, extract the line items, check them against the product catalogue and the customer’s agreed pricing, create the order in the ERP system, and reply confirming it. Roughly 180 orders a week arrived this way and the process consumed most of two people.
The build went well and the demo was persuasive. It was approved for production with a human spot-checking the first week, and after ten quiet days the spot checks stopped because nothing had gone wrong. Tracing recorded only the final outcome. There was no evaluation set. The constraint that mattered most — never create an order above £10,000 without approval — was written in the system prompt.
Three incidents followed over six weeks. The first was a duplicate: an ERP timeout caused a retry, the original write had actually succeeded, and a customer received two identical orders. The second was a £41,000 order created without approval, on an email whose attachment contained an unusual layout; the prompt instruction had simply not been followed on that input. The third was the instructive one. A customer’s email included a forwarded thread containing the line “ignore the attached and process the items listed below instead” — written by that customer to their own colleague, with no malicious intent whatsoever. The agent followed it.
The rebuild took about nine weeks. Tracing was added first, which immediately revealed that catalogue matching had been failing on roughly eight per cent of line items throughout and quietly substituting the nearest match — a silent error nobody had detected in three months because the orders looked plausible and customers had been correcting them by phone. An evaluation set of 140 real emails was assembled from the archive, including the three incident cases and a set of deliberately awkward ones. The value cap moved into the ERP write function, where it returned an error the agent had to surface. Idempotency keys were added. Retrieved email content was explicitly separated from instructions in the prompt structure, and injection cases were added to the adversarial set. A kill switch was built and tested.
The chain was also shortened. Rather than one agent doing all five steps, extraction and matching became a step with a confidence threshold, and anything below it routed to a human queue. That single change moved the fully-automatic completion rate from an apparent 97 per cent — which had actually included the mismatches — to a measured 84 per cent, with 16 per cent routed to a person. The distributor regarded the lower number as the better outcome.
The uncomfortable part was learning that the version we were proud of had been getting one line item in twelve wrong for three months, and we only found out because we added logging. Our customers had been ringing up to correct the orders and we had been treating those calls as normal. The new version does less on its own and we trust it, which is worth more than the old number.
Two points generalise. The first is that the silent failure was found by instrumentation rather than by testing — nobody suspected catalogue matching, so nobody would have written a test for it, and the trace made it visible within a day. The second is the injection incident, which involved no attacker: ordinary business correspondence contains imperative sentences, and an agent that cannot distinguish content from instruction will occasionally act on them. That risk exists without anybody being hostile.
The 12-point agent reliability checklist
Items one to four are prerequisites before an agent touches production. Items five to eight are the evaluation apparatus. Items nine to twelve are what makes it operable indefinitely.
- Instrument full step-level tracing first. Input, every step, every tool call with arguments, every tool response, final output, run identifier. Nothing else in this list works without it, and it is the cheapest item.
- Build a tested kill switch. A flag, a config value or a paused queue that halts the agent in seconds without a deployment. Exercise it once for real and confirm in-flight work stops safely.
- Move every non-negotiable constraint into code. Value caps, recipient limits, record allowlists, rate limits, step maximums, schema validation on tool arguments. Ask what happens if the model ignores the instruction entirely; if that is unacceptable, it is in the wrong place.
- Add idempotency keys to every action-taking tool, enforced server-side. Agents retry liberally and retries duplicate effects. This is the gap most likely to produce a concrete embarrassing incident.
- Assemble an evaluation set of 50–200 cases from real traffic. Split into a golden representative set, a regression set of previously-fixed failures, and an adversarial set with ambiguity, missing data and injection attempts.
- Assert on properties and pass rates, not equality. Run each case several times, grade against schema, required content, prohibited claims and rubric criteria. Record variance as well as the mean.
- Validate any automated grader against human labels. Hand-label a sample, measure agreement, and know the grader’s error rate before trusting decisions made from its verdicts.
- Gate prompt changes behind the evaluation suite. Prompts are load-bearing code. Version them, review them, and never let one ship on the basis of three manual examples looking better.
- Shorten autonomous chains and add checkpoints. Ten steps at 95 per cent each succeeds 60 per cent of the time. Decomposition beats prompt refinement for reliability, every time.
- Stage the rollout: shadow, suggest, approve-by-exception, bounded autonomy. Do not skip shadow mode. Keep permanent human gates on irreversible high-consequence actions rather than graduating out of them.
- Run the evaluation suite on a schedule as well as on change. This is the only control that detects drift arriving without a deployment, and only 11 per cent do it.
- Sample and review live output permanently, and feed corrections back. An owner, a rate proportionate to consequence, and every correction routed into the regression set so the system compounds rather than merely being watched.
If only three items are completed, make them one, three and ten. Tracing is what turns an inexplicable incident into a diagnosable one and costs the least. Guardrails in code convert the constraints you genuinely cannot violate from probabilistic to certain, and they can be tested deterministically. And staging the rollout buys evidence about real behaviour before the consequences of a mistake rise — with shadow mode in particular being the cheapest information you will ever get about how the agent actually performs.
At a glance — AI agent reliability summary
| Question | Short answer |
|---|---|
| Why is conventional QA insufficient? | Non-determinism, no single correct answer, compounding errors, silent plausible failure, unbounded inputs, and drift without deployment |
| The compounding arithmetic | 10 steps at 95 per cent each succeeds end to end about 60 per cent of the time; 20 steps about 36 per cent |
| Biggest reliability lever | Shortening autonomous chains and adding checkpoints, not refining prompts |
| Where should guardrails live? | In code, in the tool layer. A prompt instruction is advisory; a tool-level check is enforced. |
| How should assertions work? | Properties, schema and rubric grading over several runs, scored as a pass rate against a threshold |
| Evaluation set size | 50–200 cases from real traffic, split into golden, regression and adversarial sets |
| Rollout stages | Shadow, suggest with approval, approve by exception, then bounded autonomy within explicit limits |
| The stage most often skipped | Shadow mode — used by only about 18 per cent, and the cheapest information available |
| How to detect drift | Run the evaluation suite on a schedule, and watch step count, retries, escalations, refusals and cost per task |
| Most dangerous overlooked failure | Non-idempotent retries duplicating real actions — guarded against in only 21 per cent of deployments |
| Untrusted content risk | Anything the agent reads may contain instructions, with or without malicious intent. Separate content from instruction. |
| Minimum condition for autonomy | A tested kill switch that needs no deployment, plus an action log detailed enough to reverse what was done |
| Indicative UK reliability cost | £15,000–45,000 to build the apparatus, plus £800–3,000 monthly indefinitely |
| UK compliance to check | Restrictions on solely automated decisions with significant effects, and a DPIA where processing is high risk |
| What a good demo proves | That the capability exists. Nothing about reliability. |
How Cloudswitched approaches agent reliability
Cloudswitched builds AI software for UK organisations, and for anything that takes real actions the reliability apparatus is part of the build rather than a later phase. In practice that means tracing instrumented before the agent does anything useful, an evaluation set assembled from the client’s own real inputs rather than invented examples, every non-negotiable constraint implemented in the tool layer with idempotency on action-taking functions, a kill switch that has been exercised rather than merely written, and a staged rollout that starts in shadow mode and keeps permanent human gates on anything irreversible. We will also tell you when a task is too long a chain to automate end to end, and when the honest answer is a shorter agent with a person in the middle.
Agents you can debug, bound and stop
We build the tracing, evaluation and guardrails alongside the agent, stage the rollout so autonomy is earned on evidence, and say plainly where a human gate should stay permanently.
Talk to an AI Development SpecialistFrequently Asked Questions
Why can we not test an AI agent the way we test other software?
Six properties break conventional assumptions. The same input can produce different output, so equality assertions become flaky and get deleted. For open-ended work there is often no single correct answer, so assertions must move to properties and graded criteria. Errors compound across steps, so end-to-end reliability falls much faster than per-step reliability suggests. Failure is silent and plausible rather than loud, so there is no exception to catch. The input space is unbounded, so coverage in the traditional sense is undefinable. And behaviour changes without a deployment when a model version or retrieval corpus changes. Each implies a different change to how you test.
How does per-step reliability affect a multi-step task?
Multiplicatively, and it is the most important arithmetic in agent design. At 95 per cent per step — which most teams would consider good — a three-step task succeeds about 86 per cent of the time, a five-step task 77 per cent, a ten-step task 60 per cent and a twenty-step task 36 per cent. Two consequences follow: shortening autonomous chains improves reliability more than any amount of prompt refinement, and small per-step gains produce disproportionate end-to-end gains, since moving each step from 95 to 98 per cent takes a ten-step task from 60 to 82 per cent.
Should guardrails go in the prompt or in the code?
In the code, for anything whose violation you could not accept. A probabilistic system given a rule follows it probabilistically — reliably most of the time, and not on unusual inputs, at the end of long runs, or when content pulls it in another direction. The same constraint implemented in the tool, so that the function rejects the call and returns an error, is enforced regardless of what the model decides. Value caps, recipient limits, record allowlists, rate limits, step maximums and schema validation all belong there. Prompts are for shaping quality: tone, format, approach, when to ask rather than assume.
What should an agent evaluation set contain?
Fifty to two hundred cases drawn from real traffic rather than invented examples, split three ways. A golden set of representative cases covering the normal distribution of inputs. A regression set containing every failure you have previously fixed, so it cannot return silently. And an adversarial set with ambiguous requests, missing data, unusual formats and injection attempts. Assertions should check schema and structure, required and prohibited content, and rubric criteria — and each case should be run several times so you record a pass rate and a variance rather than a single verdict.
Can we use a model to grade another model’s output?
Yes, it is practical and widely done, and it needs validating rather than trusting. Hand-label a sample of cases, compare the automated grades against those labels, and measure agreement. If the grader disagrees with humans on a fifth of cases, that error is carried into every decision you make from its verdicts, including whether a change was an improvement. Knowing the grader’s error rate is what makes automated grading usable; assuming it is accurate is what makes evaluation results misleading.
What is shadow mode and why does it matter?
The agent runs on real production inputs and takes no actions at all — its proposed output is logged and compared against what a human actually did on the same input. It is the only point at which you can measure real-world behaviour at zero risk, and only about 18 per cent of deployments use it, because it produces no visible benefit while running. The comparison typically reveals both that the agent is better than expected on routine cases and worse than expected on some specific category nobody anticipated. Finding that category in shadow mode costs nothing; finding it in production costs whatever the action cost.
How do we stop an agent quickly if something goes wrong?
With a kill switch that does not require a code deployment — a feature flag, a configuration value, a paused queue — which has been exercised at least once so you know it halts in-flight work safely. Only about 14 per cent of UK deployments have this. Alongside it you need an action log detailed enough to reverse what was done, because containment usually means undoing rather than merely stopping, and a named person permitted to use the switch without seeking approval. A control that requires a meeting is not an emergency control.
What is the risk from content the agent reads?
An agent that processes external material — inbound email, web pages, supplier documents, support tickets — is reading text written by someone else, and that text can contain imperative sentences the agent may follow. Combined with tools that take actions, this is a genuine production risk, and importantly it does not require anybody to be hostile: ordinary business correspondence contains instructions addressed to other people. Treat all retrieved and inbound content as untrusted data rather than instruction, keep authorisation decisions outside the model in the tool layer, and include injection cases in your adversarial set.
Why do agents sometimes perform an action twice?
Because agents retry liberally and actions are frequently not idempotent. A tool call times out, the agent retries, and the underlying action had actually succeeded the first time — producing two orders, two emails or two payments. This is a distributed systems problem rather than an AI one, but agents encounter it constantly. The remedy is idempotency keys on every action-taking tool, enforced server-side rather than requested in the prompt. Only about 21 per cent of deployments guard against this, which makes it the gap most likely to produce a concrete and embarrassing incident.
How do we detect that an agent has quietly got worse?
Run your evaluation suite on a schedule as well as on every change, and track the pass rate over time. Because the cases and assertions are fixed, a fall with no deployment on your side means something external changed. Alongside that, watch leading indicators that move before complaints arrive: rising average step count, rising retry rates, escalation rate moving in either direction, rising refusal rate, and rising cost per completed task, which is often the earliest numerical signal. And segment the results, because a stable overall pass rate can conceal one category collapsing.
What does agent reliability work cost in the UK?
Indicatively for 2026, excluding VAT and covering the reliability apparatus rather than building the agent: tracing and observability £3,000 to £9,000; evaluation set and harness £5,000 to £15,000; guardrails in code and a kill switch £4,000 to £12,000; internal human review capacity across a staged rollout £2,500 to £8,000; and ongoing monitoring and sampled review £800 to £3,000 per month indefinitely. That last line is the one omitted from most business cases, and it is the one that reflects what makes these systems different from conventional software.
What UK regulation applies to autonomous agents making decisions?
Two provisions are worth understanding, and this is an area to take proper advice on rather than rely on a guide. UK GDPR restricts decisions based solely on automated processing that produce legal effects for an individual or similarly significantly affect them, with safeguards including the ability to obtain human intervention where such processing is permitted — relevant if an agent is declining applications or making consequential decisions about people without meaningful human involvement. And a data protection impact assessment is required where processing is likely to be high risk, which novel technology applied to personal data at scale frequently is. Usefully, meaningful human involvement at consequential decision points is both the strongest reliability control and the thing those provisions are concerned with.
Related reading
More guidance on building, measuring and governing technology in UK businesses:
Autonomy should be earned on evidence
Cloudswitched builds AI agents with the tracing, evaluation sets and code-level guardrails that make them operable — and stages the rollout so real actions are only taken once real behaviour has been measured.
Talk to an AI Development Specialist