Is Your AI Agent Saving Money or Creating More Work?
Measure AI agent ROI after review, exceptions, maintenance, failures, recovery, provider cost, and the value of useful new capability.
AI agent ROI is the useful value an agent creates minus the human attention, operating cost, mistakes, and recovery work required to keep it useful. If the calculation counts only the task the agent completed, it is not ROI. It is a demo metric.
An agent can generate a report in five minutes and still make the business slower. Someone may have to prepare special inputs, inspect every statement, resolve exceptions, rerun failed jobs, repair duplicated records, update instructions, and explain the output to the next person. The visible task shrinks while the surrounding system grows.
The honest question is not, “Can the agent do this?” It is: After the whole workflow is counted, does the organization produce more valuable work with less total burden?
Start with a baseline the agent cannot rewrite
Measure the current workflow before comparing it with an AI-assisted version. A useful baseline includes more than active task time:
- how often the work occurs;
- active human time by role;
- elapsed time between start and usable completion;
- rework and error frequency;
- the ordinary exceptions that interrupt the process;
- the cost and time required to recover from mistakes;
- the quality or business consequence of the finished output; and
- work that is currently skipped because the team lacks capacity.
Use real cases, not the clean version in a procedure manual. If the operator searches email, asks a colleague for missing context, fixes a spreadsheet, and checks the final packet twice, that is the baseline.
This is also why the first project should be a bounded workflow with observable inputs and outputs. The guide on how to pick a first AI project gives a fuller selection scorecard. Without a stable unit of work, the team cannot tell whether the new system changed anything.
Count contribution and burden separately
A practical equation is:
> Net productive effect = useful contribution − total operating burden
Useful contribution can include:
- active labor capacity reclaimed;
- shorter elapsed time when speed changes a real outcome;
- more correct, complete, or consistent output;
- avoided rework, missed handoffs, or preventable loss;
- additional throughput the organization can actually use; and
- a valuable new capability that was previously uneconomic.
Total operating burden can include:
- preparing or cleaning inputs for the agent;
- reviewing its output;
- handling exceptions and uncertainty;
- maintaining prompts, instructions, tools, and integrations;
- building and rerunning evaluations;
- monitoring the production system;
- provider and infrastructure expense;
- training people and changing the workflow;
- investigating and recovering from incidents; and
- correcting downstream mistakes.
Keep unlike units separate until the organization has a defensible conversion. Minutes, dollars, error rates, customer response time, and additional coverage do not become one number merely because they fit in a spreadsheet.
A fictional example: the six-hour packet
Consider a fictional internal workflow that assembles ten review packets each week. It takes six human hours without an agent.
In the first AI-assisted month, the agent prepares the packets in 45 minutes of unattended runtime. Calling that an 87.5% time saving would be wrong. The weekly human burden is:
| Work around the agent | Weekly human time |
|---|---|
| Prepare and check inputs | 0.5 hour |
| Review ten packets | 2.0 hours |
| Resolve ordinary exceptions | 1.0 hour |
| Maintain instructions and tests | 0.5 hour |
| Average incident recovery | 0.5 hour |
| Total human burden | 4.5 hours |
The net capacity reclaimed is 1.5 hours, before provider cost and any change in quality are counted. That may still be worthwhile. It is simply a different decision from “the agent did six hours of work in 45 minutes.”
Now suppose the agent also catches missing evidence that previously caused one packet per month to return for rework. That avoided rework belongs on the contribution side. If it produces more packets than anyone needs, the extra output does not.
The example is fictional and illustrates the accounting method; it is not a Richardson Applied AI client result.
Measure actual work, because intuition can be wrong
Research does not support one universal productivity number for AI.
An NBER field study of 5,179 customer-support agents reported a 14% average increase in issues resolved per hour after a conversational assistant was introduced. The effect was much larger for novice and lower-skilled workers and minimal for the most experienced workers. The tool, work, users, and outcome measure all mattered.
A different result appeared in METR's early-2025 randomized study of 16 experienced open-source developers completing 246 real tasks. With AI tools available, they took 19% longer, even though they believed the tools had made them faster. METR explicitly warned against generalizing that result beyond the studied setting.
The measurement problem has become harder as tools and work patterns change. In a February 2026 update, METR said its newer developer experiment could not produce a reliable current speedup estimate because participation and task selection changed, the pay rate changed, and concurrent agents made time reporting difficult. Its later explanation of task substitution and uplift separates speed on the old task list from the value of a new task mix made possible by AI.
The practical conclusion is narrower and more useful: measure the people, workflow, tool configuration, and result that actually exist in your organization. Do not borrow a percentage from a different deployment and call it a business case.
Reliability has an economic weight
Average success rate can hide expensive failures. A workflow that is correct nine times out of ten may still be negative if the tenth case creates a long recovery or an external consequence.
Track at least four reliability measures:
1. Useful completion rate: Did the run produce the complete output the next person needed? 2. Exception rate: How often did a person have to intervene, and for how long? 3. Recovery burden: How much human time and operating disruption followed a failed or ambiguous run? 4. Consequence severity: What happened when the system was wrong, late, incomplete, or duplicated?
An expected recovery burden can be estimated as incident frequency multiplied by average recovery time. Expected financial consequence can be estimated separately when the organization has enough evidence. Rare, severe events should not be disguised by a strong average.
The NIST AI Risk Management Framework calls for performance to be demonstrated in conditions similar to deployment, production behavior to be monitored, and mechanisms for override, incident response, recovery, change management, and decommissioning to be documented. Those controls are not free. They are part of the system being evaluated.
Monitoring should reduce uncertainty, not create a second job
Monitoring has its own ROI problem. Collect too little and the team learns about failures from a customer or a broken downstream record. Collect everything without a decision rule and people inherit another dashboard they do not use.
NIST's 2026 report on challenges in monitoring deployed AI systems separates functionality, operations, human factors, security, compliance, and broader impacts. It also notes that monitoring methods and terminology remain fragmented and that user burden is an open challenge.
For a bounded business workflow, start with signals tied to decisions:
- Did the job start and finish?
- Did it process the expected items exactly once?
- Did the output pass the workflow's acceptance checks?
- Which exceptions required a person?
- How long did review and recovery take?
- Did the provider, model, data, or instructions change?
- Is the net productive effect improving, flat, or declining?
The production AI agent checklist covers the operational controls in more detail. A log without ownership and response behavior is evidence storage, not monitoring.
Give new capability its own line
Some agents are valuable because they do work that was not performed before. They may review every item instead of a sample, assemble information before a decision, or maintain follow-through the team could not support manually.
Do not force that value into “hours saved” when no old labor was replaced. Name the capability and its observable business consequence:
| New capability | Evidence to track |
|---|---|
| Review every item in a queue | Coverage rate, useful findings, false alarms, review burden |
| Prepare work before a specialist begins | Ready-on-arrival rate, specialist cycle time, missing context |
| Detect missed handoffs | Confirmed misses caught, false positives, recovery avoided |
| Maintain faster follow-through | Response time, completed handoffs, unwanted or duplicate actions |
The value counts only when the organization uses the output and can observe a consequence. Generating more material is not automatically more capability.
This distinction is consistent with the GAO AI accountability framework, which separates performance against program objectives from monitoring reliability and relevance over time. The business objective comes first; the agent's activity is evidence only when it advances that objective.
Run a net-productivity review
Use one scorecard for the baseline and one for the AI-assisted workflow:
| Category | Record each review period |
|---|---|
| Useful output | Complete outputs accepted by the next user |
| Active human time | Input preparation, direct work, review, and exceptions |
| Elapsed time | Start to usable completion, including waiting and retries |
| Quality | Completeness, accuracy, consistency, and rework |
| Reliability | Successful runs, partial runs, duplicates, and safe stops |
| Recovery | Incidents, human recovery time, and downstream consequence |
| Maintenance | Instruction, integration, evaluation, and monitoring work |
| Direct cost | Model, software, infrastructure, and vendor expense |
| New capability | Coverage or useful work that did not exist before |
| Business consequence | Capacity, response, quality, risk, or decision improvement |
Review enough cycles to include normal exceptions. Two weeks may be adequate for a frequent, stable internal task. A monthly or seasonal process may need several occurrences. Compare equivalent periods where possible and record any change in task mix, staffing, model, provider, or instructions.
Then choose one disposition:
- Expand when the net effect is positive and the evidence supports a larger boundary.
- Keep bounded when value is real but errors or operating costs justify the current scope.
- Repair when a specific measurable burden is suppressing an otherwise useful workflow.
- Return part to manual work when a human performs one stage more economically.
- Retire when the burden remains larger than the useful contribution.
The time already spent building the agent is not a reason to keep it. It is evidence about the next decision.
Measure the system, not the model trick
AI agent ROI belongs to the whole operating system: the workflow definition, evidence, instructions, tools, evaluations, human boundaries, monitoring, recovery path, and the people who use the result.
That is why an AI project should be built alongside the organization's durable operating knowledge, not as an isolated prompt. Building the company brain and the first workflow together makes the source material, ownership, evaluation, and maintenance burden visible from the beginning.
If the agent produces more useful work after every required human and technical cost is counted, it is a capability. If it creates a stream of review, confusion, and recovery larger than its contribution, it is an experiment that has delivered its lesson.
The point is not to make every agent prove an impressive percentage. The point is to know what the system is actually doing to the business.
Common questions
What should an AI agent ROI calculation include?
Include the value of useful output, time or capacity reclaimed, avoided rework, and valuable new capability. Subtract review, exception handling, maintenance, evaluation, monitoring, provider costs, incident recovery, and the expected consequence of errors.
How long should you measure an AI agent before deciding whether it works?
Measure enough real cycles to include ordinary variation and exceptions. A two-to-four-week bounded pilot is often more informative than a demo, but seasonal or low-frequency workflows need a longer evidence window.
How do you value a capability the company did not have before?
Define the business consequence first, such as greater coverage, faster response, or fewer missed items. Track that outcome separately until the organization has a defensible financial conversion for it.
When should an AI agent be retired?
Retire or redesign it when the full operating burden persistently exceeds its useful contribution, when errors cannot be contained economically, or when a simpler workflow can produce the same result more reliably.