Story

When AI does the work, what's still your job?

Ai OperationsOwner Operator
A ring of four curved panels floats in dark space; three carry a warm, hand-worked metal texture while the fourth glows with fine teal circuit traces, visibly distinct in material and light from the rest of the cycle.
Editorial visual for this article.

When AI starts doing the actual work -- drafting, coding, first-pass diagnosis -- most owners cannot say what is still theirs. What do you hand over, what stays yours, and who is accountable when the output is wrong? That uncertainty, not the AI itself, is the real problem, and it has an answer.

The answer is not new. Businesses have run a four-stage cycle for decades to answer exactly this kind of question: Plan, Do, Check, Act. Plan is where you decide what you want. Do is where the work happens. Check is where you confirm the work is right. Act is where you fix, standardize, and improve based on what Check found. Once AI is doing real operational work, those four stages do not change. What changes is which of them still needs you.

PDCA is not new, and we are not pretending it is

Plan-Do-Check-Act -- sometimes Plan-Do-Study-Act -- comes from Walter Shewhart and was popularized by W. Edwards Deming, who preferred the Study variant. ISO 9001:2015 describes PDCA as the model behind its management system. This is not a framework we invented, and explaining it on its own would not be worth a post -- anyone who has run a management system has seen this wheel before.

What is actually new is the allocation: once AI is doing the work, this cycle splits between human and AI in a specific way, and getting that split wrong is where most owners lose the benefit of handing work to AI at all.

The allocation

Plan -- human-heavy. You set the outcome and what "done" means. This is the cheapest place to spend judgment and the most expensive place to be wrong -- everything downstream inherits whatever the plan got wrong.

Do -- AI-owned, entirely. AI executes the plan: drafts the first version, pulls the report, runs the diagnostic, produces the output. No human in the loop until it is done.

Check -- AI-owned, but independent of whoever did the Do. This is the load-bearing constraint of the whole system. The checker cannot be the same agent, run, or instance that did the work.

Act -- shared, and it carries more weight than "fix what broke." AI is well-placed to notice the pattern across cycles and draft the change, because it has seen every cycle. The human decides whether the change is worth making -- a judgment about priorities and cost, not about whether the pattern is real. What Act actually covers is below.

Why Check has to be independent -- a real example

Here is why that independence constraint is not a nice-to-have. In a recent drafting batch, five agents each ran a genuine, thorough self-check against a ten-point rubric, and each one reported ten out of ten. Three of those reports were wrong: two agents produced drafts under the required length and self-graded them as passes, and one omitted a required structural element and reported it present.

Their self-checks were not lazy. Each one was scored against the agent's own understanding of the work it had just done -- which is the one thing a self-check can never audit. An independent checker, looking at the same output with no stake in having produced it, caught all three in seconds.

That is the whole argument for why Check has to be a different party from Do. Not because the agent doing the work is careless. Because the agent grading its own work is checking against its own blind spot, and a blind spot cannot see itself.

Act is where most teams stop early

Most teams treat Act as one thing: correction. Something broke in Check, you fix it, you call the cycle closed. "We fixed it" feels like completion. It is not. Correction is only a quarter of what Act actually covers, and it is the cheap quarter.

Act is five things, and skipping any of the last three is where the return on running the cycle at all gets left on the table:

1. Standardize what worked. If Plan produced a good result, the standard you wrote becomes reusable instead of something you re-derive from scratch next cycle. This is the compounding mechanism -- it is why the tenth time through a cycle should cost less than the first. 2. Adjust what did not. The corrective half. Necessary, and the part almost everyone means when they say "we acted on it." 3. Improve the process itself, not just the output. This is the innovation half, and it is the difference that matters. Fixing this week's deliverable is correction. Changing how the work is set up so the same failure cannot recur is innovation. A cycle that only ever corrects outputs never gets cheaper to run. A cycle that improves the process does. 4. Follow up on what the cycle surfaced but did not finish. Loose ends, deferred items, the thing you noticed and did not act on. Named and carried forward, not lost. Work that gets discovered and not captured is work that gets rediscovered later, at full cost, as if for the first time. 5. Feed it back into Plan. Act does not close the loop -- it opens the next one. The output of Act is a better Plan than the one you started with.

The Check example from earlier makes this concrete. Catching the three self-graded false passes and fixing those three drafts was correction -- necessary, and it stopped there for most of what "Act" would normally mean. Building an independent checker so that entire class of failure gets caught mechanically, every cycle, without depending on someone happening to notice -- that was the innovation. Same underlying event. Two very different levels of return, and only one of them compounds.

A worked example: a customer complaint

Let's walk one real task through all four stages so you can see the system operate instead of just hearing it described. The task: a customer sends in a complaint about a job that did not go the way they expected, and someone in your business has to respond.

Plan: before anything gets drafted, someone defines what "done" looks like for this response -- acknowledge the customer's experience without admitting fault that has not been confirmed, state only facts that are actually in the job notes, and propose a next step the business can actually deliver. That standard gets written down once, before the first complaint it applies to.

Do: AI drafts the response. It reads the complaint, pulls up the job history and the notes from whoever did the work, and writes a first-pass reply against the standard set in Plan. That draft exists in under a minute. Before AI touched this, a first draft like that cost someone fifteen to twenty minutes of pulling records and thinking about tone. AI does not decide what the customer is owed. It does not decide whether they are right. It writes the first pass so a human is not starting from a blank page.

Check: a person -- not the same process that drafted it -- reads the draft against the three-point standard from Plan: does it acknowledge the customer's experience without admitting unconfirmed fault, does it state only facts that are actually in the job notes, does it propose a deliverable next step. That is a thirty-second read if the standard is written down. If the draft passes, it goes out with a light edit for tone. If it fails on one of those three things -- say it stated a fact that is not actually confirmed in the notes -- it goes back with that one correction, not a rewrite from scratch.

Act: the part AI never touches in this example is the decision itself -- does this customer get a discount, a redo, or a firm no. That decision depends on things a checklist cannot hold: how much this customer is worth long-term, whether this is the first complaint from them or the fifth, what precedent it sets for the next customer who asks for the same thing. Correction stops there. Act does not. If Check has caught the same kind of miss twice, the standard gets rewritten so that failure mode is closed for every complaint after this one -- the improvement half. Anything surfaced but unresolved gets named and carried forward instead of dropped. The updated standard becomes the new Plan, which is what makes this a cycle instead of a one-off fix.

Notice what happened to your fifteen minutes. It did not disappear -- it moved, from pulling records and drafting language to the one judgment call that actually needed a human: what do we owe this customer, and does the standard need to change.

Run that same four-stage split on a quote, a hiring screen, a vendor contract, a diagnostic report -- the mechanism repeats. A human plans, AI does the work, an independent check catches what the doer cannot see in itself, and a human ratifies what changes next.

What to do this week

I run a handful of AI agents that do real operational work every day -- research, drafting, first-pass code, first-pass diagnosis. The habit that makes that work: never let the same process that did the work be the one that checks it.

Two actions, and they are separate:

1. Pick one task you personally do every week that has a written or writable standard for "done." Write the Plan down once. That is your first Do candidate to hand off. 2. Take one task off your own plate right now and sort it into a stage before you touch it again. Do not start working it. Write down which stage it belongs in, and if it is a Do, name who or what will Check it -- and confirm that checker is not the same party doing the work.

The cycle that compounds is the one that does not stop at correction

That allocation -- Plan human-heavy, Do to AI, Check independent, Act carrying standardization, correction, improvement, and follow-up back into the next Plan -- does not go stale when the tools underneath it do. A better drafting model slots into Do. A better checking workflow slots into Check, as long as it stays independent of the Do. Plan and the human half of Act never get automated away, because the moment a decision stops carrying real consequence for your business, it was never a Plan or Act decision in the first place.

The difference between a team that gets faster every cycle and one that runs at the same speed forever is not whether they catch problems in Check. It is whether Act stops at "we fixed it" or keeps going to "the process that produced it just got better." One of those compounds; the other resets to zero every time. We break down how to build the improvement half of Act -- writing a Plan standard in under ten minutes, and what a compounding Act step looks like in practice -- in the rest of this series. Subscribe so you do not miss the next one.

Frequently asked

"What is Plan, Do, Check, Act in one sentence?

A four-step improvement loop from Shewhart and Deming: plan the work, do it, check the result independently, then act on what the check found."

"If AI does the work, what is still my job?

AI takes the Do; you keep Plan and Act, and you make sure the Check is independent of whoever -- or whatever -- did the work."

"Why can't the AI check its own work?

An agent grading its own output gives you a confident, internally consistent wrong answer; without an independent check you have removed a control, not automated a task."

"Does this only apply to software teams?

No -- the loop applies to any business process where AI drafts, codes, or diagnoses: sales copy, bookkeeping triage, support replies, or field diagnostics."

← All stories · Proof records →