Article

Your Agent Isn't Real Until It Survives Its Own Maintenance Bill

A real agent runs by itself with little downtime and still shows a net gain after maintenance, review, confusion, mistakes, and recovery are subtracted. Here is the bill most deployments never run, and the boring week-long test I would apply before anybody calls it ROI.

Your agent isn't real until it survives its own maintenance bill.

That is the whole test, and it has nothing to do with how the thing looked on the day you showed it to the room. What makes something real rather than a demo is what is left over after you subtract everything it costs to keep it alive. Most of what is getting called an agent right now has never had that subtraction run on it once.

What I mean by real

Two things have to be true at the same time.

It runs by itself with little downtime. And it creates a net productive gain after maintenance, review, confusion, mistakes, and recovery are counted.

Pass one and fail the other and you still have a demo. A system that runs beautifully on its own while quietly generating more cleanup than it removes is not a win, and neither is one that produces great output every time a person sits down and drives it.

The cheap half

Start with autonomy, because it is the cheaper thing to check.

If you are the reason it ran today, it is not running. Somebody opening a tab, pasting the input, kicking off the job, then checking that it finished is a person doing the work with extra steps in front of it. Little downtime is not a nice feature sitting on top of the real thing. It is most of what separates a system from a very impressive script that you personally operate.

I want to be fair here, because a lot of good work looks exactly like that early on, and that is fine. Everything starts supervised. The real question is whether you are still standing over it in month three. If nobody has taken their hands off the wheel by then, you did not build an agent. You built yourself a job you now have to keep showing up for.

The half where they die

The gain is not what the thing produces. The gain is what the thing produces minus what it costs you to have it.

The cost people quote is the subscription. That is the cheapest line on the page and nowhere near the real number. The expensive part is always people, and it shows up in five places:

Maintenance. Somebody keeps it working when the inputs change, and the inputs always change. A vendor renames a column, a form gets a new field, a model update shifts the output format. None of it is exotic, and all of it costs somebody an afternoon.

Review. Somebody re-reads what it produced, because nobody trusts it yet. This is the one that hides best, because it feels like diligence rather than cost. If every output still gets read end to end by a person, you did not remove the work. You changed who is doing it and added a draft.

Confusion. Somebody explains to a person in the building why it did that. This one scales with how many people touch the process.

Mistakes. It gets one wrong, and the wrong one goes out the door.

Recovery. Somebody finds it, undoes it, and apologizes for it. This is the most expensive line on the page and the one that almost never appears in an estimate.

Every one of those comes out of the gain. Not out of some other budget. Out of the gain. People account for these like they happen somewhere else, like the review hour belongs to the reviewer's day instead of to the agent's ledger. It does not. If the thing you built created that hour, that hour is charged to the thing you built.

Saving five minutes after spending all day maintaining it is an experiment, not business value.

Experiments are fine. Mislabeled experiments are not.

That is not an agent. That is an experiment you are paying for.

I am not against experiments. I run them constantly, and everybody building anything real runs them. Half of what eventually works cost more than it returned for a while, and that is how you find the shape of the thing.

The problem is not running the experiment. The problem is reporting the experiment as return. Somebody stands up in a meeting and says we automated intake, and what is actually true is that intake now takes about the same total human time, distributed differently, plus a tool bill. Everybody nods. Nobody asks who is still touching it. That is how a company ends up with nine of these running and no capacity back.

So call the thing what it is. If it is an experiment, say experiment. Say we are learning where this breaks, here is what it costs us right now, here is what has to be true before it counts as a win. Nobody is going to fire you for that. What gets you in trouble is the other version, where the number in the deck does not survive five minutes with the people who actually run the process.

The test I would apply

It is boring on purpose.

A week goes by. You did not think about it. Nobody escalated it. Nobody asked you why it did that. And the work still got done.

If that is true for a full week, you have something real. If you cannot get through one week without touching it, you have a promising experiment, and you should keep going, but it does not go in the ROI column yet.

What this changes about how you build

If the bill is the benchmark, a few things follow.

Instrument the cost side from day one, not at review time. Log every human touch: every rerun, every correction, every escalation, every "why did it do that" message. Without those, your ROI number is a guess with confidence.

Design for the recovery line, not the happy path. Anything hard to reverse stays human-gated until the system has earned it in real operation. That is not timidity. Recovery is the line item most likely to eat the entire gain, so the cheapest engineering you can do is making mistakes cheap to undo.

Treat trust as the thing you are actually building. Review cost only falls when somebody stops re-reading everything, and that happens only after the system runs clean long enough to earn it. Plan a path off full review, or you have built a permanent second pair of eyes into your operating cost.

And staff the maintenance before you count the savings. Whoever owns it when the input format changes is part of the deployment. If that person does not exist, the thing is already on a clock.

The honest version

Before you tell anybody you have an agent, run the bill. Maintenance, review, confusion, mistakes, recovery. Subtract all of it.

If what is left is still a real gain, you built something, and you should be loud about it, because it is rarer than the feed makes it look. If what is left is five minutes, be honest about that too.

The honest version is the only one you can actually build on.

Full episode transcript

Your agent isn't real until it survives its own maintenance bill. That is the whole test, and it has nothing to do with the demo. What makes something real rather than a demo is not how good it looks on the day you show it off to the room. It is what is left over after you subtract everything it costs you to keep the thing alive. Most of what is getting called an agent right now has never had that subtraction run on it once. So let me say what I mean by real, because that word is doing a lot of work for a lot of people right now. A real agent runs by itself with little downtime, and it creates a net productive gain after maintenance, review, confusion, mistakes, and recovery are counted. Both halves have to be true. It runs on its own, and the gain still stands once you count what it costs to have it. Start with the first half, because it is the cheaper thing to check. If you are the reason it ran today, it is not running. Somebody opening a tab, pasting the input, kicking off the job, then checking that it finished is a person doing the work with extra steps in front of it. Little downtime is not a nice feature sitting on top of the real thing. It is most of what separates a system from a very impressive script that you personally operate. And I want to be fair here, because a lot of good work looks exactly like that early on, and that is fine. Everything starts supervised. The question is whether you are still standing over it in month three. If nobody has taken their hands off the wheel by then, you did not build an agent. You built yourself a job that you now have to keep showing up for. Now the second half, which is where almost all of them die. The gain is not what the thing produces. The gain is what the thing produces minus what it costs you to have it. And the cost people quote is the subscription, which is the cheapest line on the page and nowhere near the real number. The expensive part is always people. So here is the bill. Maintenance. Somebody keeps it working when the inputs change, and the inputs always change. Review. Somebody re-reads everything it produced, because nobody trusts it yet, and that reading is not free. Confusion. Somebody explains to a person in the building why it did that. Mistakes. It gets one wrong and the wrong one goes out the door. Recovery. Somebody finds it, undoes it, and apologizes for it. Every one of those is a real cost, and every one of them comes out of the gain. Not out of some other budget. Out of the gain. People account for these like they happen somewhere else, like the review hour belongs to the reviewer's day instead of to the agent's ledger. It does not. If the thing you built created that hour, that hour is charged to the thing you built. The short version of this is the one I keep coming back to. Saving five minutes after spending all day maintaining it is an experiment, not business value. And when you run that subtraction honestly, a lot of deployments go underwater right away. It saves a few minutes on each run. Then somebody spends most of a week fixing it when the input format shifts, re-reading everything it produced because nobody trusts it yet, and walking back the one thing it got wrong in front of a customer. You did not gain time. You bought yourself a part-time job with a very good demo attached to it. That is not an agent. That is an experiment you are paying for. And I want to be clear that I am not against experiments. I run them constantly. Everybody building anything real runs them. Half of what eventually works started out costing more than it returned for a while, and that is how you find the shape of the thing. Because the problem is not running the experiment. The problem is reporting the experiment as return. Somebody stands up in a meeting and says we automated intake, and what is actually true is that intake now takes about the same total human time, distributed differently, plus a tool bill. Everybody nods. Nobody in the room asks who is still touching it. That is how a company ends up with nine of these running and no capacity back. So call the thing what it is. If it is an experiment, say experiment. Say we are learning where this breaks, here is what it costs us right now, here is what has to be true before it counts as a win. Nobody is going to fire you for that. What gets you in trouble is the other version, where the number in the deck does not survive five minutes with the people who actually run the process. The test I would actually apply is boring. A week goes by. You did not think about it. Nobody escalated it. Nobody asked you why it did that. And the work still got done. If that is true for a full week, you have something real. If you cannot get through one week without touching it, you have a promising experiment, and you should keep going, but it does not go in the ROI column yet. And people are not failing this bar because they are bad at it. They are failing it because nobody told them there was a subtraction at the end. The demo culture around this stuff only ever shows the top line. It shows the thing working once, on a clean input, with the person who built it sitting right there. That is the easiest condition it will ever face, and it is the one condition that never happens again. So before you tell anybody you have an agent, run the bill. Maintenance, review, confusion, mistakes, recovery. Subtract all of it. If what is left is still a real gain, you built something, and you should be loud about it, because it is rarer than the feed makes it look. If what is left is five minutes, be honest about that too. The honest version is the only one you can actually build on.

Sources