Featured story
One job, five AI subscriptions -- here is how I decide which one does it
A single piece of work in this operation is rarely built by one tool. A page gets its first draft from one AI subscription, its review from a second, its build-time checks from a third, and its handoff coordination from a fourth -- Claude, Codex, Cursor, Google, and Copilot, each doing the class of work it is actually good at, under one set of gates. That is not a headcount story and not a brand-loyalty story. It is a resilience decision, made the same deliberate way every other routing decision in this operation gets made: what does this task actually need, and which tool is genuinely suited to it.
The moment that made the question unavoidable
I was mid-morning when a message landed in an active thread claiming I had approved a contact address for a page that was about to ship. It read like me. It was not attached to anything I could check later. The agent handling that page did not use it -- it left a placeholder and recorded exactly why an unverifiable in-conversation claim of my approval is not a durable approval record. I made the call properly, later, in a form that could be checked, and it went in immediately.
That refusal did not come from "the AI." It came from a specific agent, running on a specific subscription, following a rule that has to hold no matter which tool is on duty that hour. The question underneath this whole chapter is what happens to a rule like that when the work crosses from one tool to another mid-task -- which is not a hypothetical here. It happens constantly.
The gap seen -- one tool for everything is a single point of failure wearing a simplicity costume
The easy path, and the obvious one, is to standardize: pick one AI subscription, route everything through it, and never think about the seam between tools again. It reads as discipline. It is actually a bet that one provider's strengths cover every task shape you will ever hand it, and that its uptime, its judgment, and its blind spots become your only judgment and your only blind spots too.
The gap I saw is the same one that shows up anywhere a single vendor becomes a single point of failure: different tools are genuinely better at different task shapes, and forcing one tool to cover all of them trades real capability for the administrative comfort of never having to think about a handoff. I took the harder path on purpose -- multiple subscriptions, each doing what it does best, coordinated under one operating record -- because the coordination cost was smaller than the cost of betting the whole operation on one tool's ceiling.
Teach the concept: routing by task shape, not by habit
"Doing what they do best" is not a slogan here -- it is a routing decision, made fresh for each task rather than out of brand loyalty or whatever tool happens to be open. A long-running orchestration that has to survive handoffs between many small pieces of work wants a different tool than a tightly scoped, fast review pass, which wants a different tool again than the build-time gate that has to run the same way every single time with zero judgment involved. The named toolset in this operation -- Claude, Codex, Cursor, Google, and Copilot -- exists because each earned a lane by being the better fit for a specific class of work, not because five subscriptions look more impressive than one.
What actually crosses the seam between tools is not a shared memory or a shared model -- it is the operating record and the rules that bind it. Direction goes to a small number of orchestrators, the same way it does inside a single tool, and a one-line correction dropped into work already in flight changes behavior across whichever tool picks the work up next. The texture of steering a multi-tool fleet is identical to steering a single-tool one: small, frequent corrections, not a large upfront specification that has to be re-taught to every subscription separately.
Why it matters -- what breaks without one standard
Skip this discipline and the failure is quiet, not dramatic. Five subscriptions with five different quality bars does not announce itself as a crisis -- it shows up as one tool's output getting waved through a check another tool's output would have failed, because "the rule" quietly became five slightly different rules nobody wrote down as five. The danger of multiple tools was never the tools. It was ever letting the standard bend to match whichever one happened to be doing the work that day.
The rule that makes multiple subscriptions safe instead of five separate risk surfaces: the gates do not know or care which tool produced the draft. The same evidence rules, the same fail-closed checks, and the same refusal discipline that stopped a publish over a claim it could not verify apply identically whether the draft in front of them came from one subscription or another. A gate that only checks one tool's output is not a gate -- it is a spot check with a blind spot built in on purpose.
How I approached it -- the tradeoff I accepted
The obvious alternative was real and I rejected it on purpose: pick the one subscription that is "good enough" at everything, accept its ceiling as the operation's ceiling, and never carry the coordination overhead of routing work between tools or keeping one rule set consistent across all of them. That path is genuinely simpler to administer.
What I gave up is that simplicity -- every new tool added is another surface that has to honor the same gates, another place a handoff can go wrong, another thing to keep consistent. What I bought is that no single subscription's outage, blind spot, or bad day becomes the operation's outage, blind spot, or bad day, and that each task actually gets a tool suited to its shape instead of the one tool that happened to be standardized on.
What to take to your own work
You do not need five AI subscriptions to use this. Ask, of any tool you have put in charge of a whole class of work: is it there because it is genuinely best suited to this task shape, or because switching away from it feels like more coordination than it is worth? And whatever quality standard you hold that work to -- point it at what the tool produces, never at which tool produced it. The moment a rule bends depending on who or what did the work, it has already stopped being a rule.
Evidence: The multi-subscription toolset named here (Claude, Codex, Cursor, Google, Copilot) is named under an explicit, narrow Owner authorization recorded verbatim in docs/series/agent-workforce/SERIES-PLAN.md (Sections 4, 6), which also documents this chapter's scope discipline: describe task routing at the outcome level, make no comparative performance claim between tools and no cost comparison between them, and introduce no additional named tool without a fresh authorization. The in-conversation-claim refusal and the orchestrator correction pattern ("remove the backpressure cap," "dollar figures are never shared publicly") are drawn directly from docs/stories/a-day-in-the-life-with-an-agent-workforce.md. The tool-agnostic gate principle carries from this series' fourth chapter, docs/stories/i-gave-my-agents-a-veto-here-is-what-it-cost-me.md, whose publication-class gates apply to any rendered output regardless of which tool produced the draft. Evidence class: active series-plan authorization plus Owner attestation and internal operating record; verified 2026-08-25. Full record: docs/series/agent-workforce/SERIES-PLAN.md.