Story
Just In Time AI
Why Can a Second AI Opinion Still Leave You With the Wrong Decision?

CHALLENGE: Can you prove an AI recommendation is strong enough to authorize the work?
Ask a second AI for a check on a plan, get an answer that agrees with the first one, and you still have no idea whether the work is safe to fund. In our own platform planning, the working plan said 8 features and 24 stories while the design of record already held 9 features and 31 stories -- one feature and seven stories missing. Those seven omitted stories are 22.6 percent of the final 31-story scope, and moving from the stated 24 to the actual 31 is a 29.2 percent increase in story count. Two models reading that same plan could have agreed with each other, sounded certain, and sent us to fund the smaller number.
That is the challenge, and we can meet it without becoming the technical expert in the room. A bounded review is a review whose question, evidence, limits, required answers, and decision authority are fixed in writing before anyone -- person or model -- is asked to answer. Fix those first and a recommendation becomes inspectable. Skip them and we are grading confidence.
The short answer: a second opinion can repeat the same stale premise, because agreement measures overlapping inputs rather than truth. A bounded review makes the question, the evidence revision, the exclusions, the reviewer's scope, the shape of the verdict, the disagreement rule, and the basis for confidence explicit, while final authority stays with the accountable owner.
None of this is exotic. Practices such as architecture decision records and threat modeling share the same two properties a bounded review depends on: evidence someone else can inspect, and authority someone is accountable for.
Turn a high-stakes AI decision into a decision you can defend. Just In Time AI can help you turn a high-stakes AI decision into a bounded review with explicit evidence, authority, and a decision record, then implement the smallest useful change once the decision is clear. Explore AI Systems Setup and Coaching.
WHAT: What makes a bounded review decision defensible?
Three things make it defensible: a packet written before dispatch, a review that is able to disagree with the plan it reads, and a decision recorded by the authority assigned to that decision. Here is the packet we wrote for the platform scope question.
| Part of the packet | What we wrote |
|---|---|
| Decision question | Does the next tranche of platform work match the design of record, and what is the smallest tranche worth authorizing? |
| Intent and end state | A tranche we can fund and finish without re-planning it mid-flight |
| Evidence revision | The design of record as frozen on the review date; later edits do not enter this review |
| Exclusions | Staffing, delivery dates, and commercial terms |
| Reviewer | A different model and persona set from the one that wrote the plan, read-only |
| Risk coverage | Architecture, maintainability, reliability, security; recorded gap: commercial risk, held by the owner |
| Required answers | Observations, challenged premises, alternatives with consequences, evidence limits, recommendation, basis for confidence |
| Disagreement path | Both positions preserved, plus the smallest reversible probe that settles them |
| Owner and deadline | Named decision owner, a stated deadline, and no authorization as the default on silence |
| Reopen condition | Any change to the design of record, the scope, or the authority |
Around that packet we run four stages, in order, whenever a recommendation is large enough to justify one. Field-by-field implementation is outside the scope of this business article.
- Frame and authorize before anyone reviews. The accountable decision owner -- not the reviewer, and not the author of the plan -- fixes the decision question, the evidence revision, the exclusions, the deadline, the authority reserved to the owner, and the answers the review must return. The trigger is a recommendation expensive to reverse, and nothing is dispatched until the packet is written.
- Get a review that is able to disagree. The reviewer is someone or something other than the author, works read-only from the frozen evidence, covers the named risk areas, and is asked directly to find what is wrong with the premise. Withholding the original recommendation matters here: a reviewer handed the first answer may grade it instead of deriving its own.
- Resolve dissent, then record and decide. Competing positions stay on the record with the evidence each rests on, along with the smallest reversible probe that would settle the split. The owner then records the decision with its identifier, status, alternatives and consequences, basis for confidence, authority, exclusions, and reopen criteria; an unanswered decision sits in the owner's queue with its deadline, and silence defaults to no authorization.
- Recheck before the work starts, and check the outcome after. Before execution, confirm the reviewed revision still matches, and stop if the evidence, scope, or authority moved. After the stated check date, compare the result against what was predicted, and when a review catches a stale premise, repair the source or the control that let it into the packet.
What the review produced is a record, not a memo. Here is the worked record in its smallest useful form.
| Part of the record | What it held |
|---|---|
| Observations and challenged premise | The working plan stated 8 features and 24 stories; the design of record held 9 features and 31 stories |
| Alternatives and consequences | Fund the plan as written and absorb one feature and seven stories later; fund the corrected scope and commit more now; fund a probe slice and start slower |
| Recommendation and confidence basis | The corrected scope, with confidence resting on the frozen design rather than on reviewer certainty |
| Decision authority and record | In the recorded case, the technical agent made the technical decision after the review; the reviewer recommended and did not authorize. Business sequencing authority remained reserved to the Owner. |
| Exclusions and reopen condition | No commitment to dates or staffing; reopens on a design change to the affected features |
| Outcome check | A named check date and the observable result expected by then |
What this case proves is bounded. It shows one premise correction and one recorded decision process. It does not show shipped results, a dollar saving, a calibrated confidence score, or a correct number of reviewers.
Why can two AI answers agree on the wrong premise?
Because they can read the same page. The recorded case does not establish that two models were handed the same working plan: it documents one independent reviewer comparing the working assumption with the design of record. In a future review, if two models receive the same stale plan and neither is asked whether it still matches the design, their agreement measures shared inputs.
A second model handed the first recommendation can behave worse than one handed nothing: it may grade the answer in front of it instead of deriving its own, and grading can pull toward agreement. Fluency hides this. A well-organized answer built on a stale number reads exactly like one built on a current number.
The correction is cheap: withhold the recommendation, hand over the frozen evidence, and make disconfirmation the assignment.
How do you tell whether the review challenged the evidence?
Read what the reviewer cited, not how confident it sounded. Compare its decision question and its evidence revision against the first recommendation's. Different question or different revision means the two answers were never comparable.
Then apply the criteria the review is graded on:
- Challenged premises. The review names at least one stated assumption, tests it against the frozen evidence, and says plainly whether the assumption held or was corrected.
- Alternatives with consequences. Each option carries what it costs you and what it forecloses, so the choice can be compared rather than argued.
- Evidence limits. The review states in plain language what the frozen evidence cannot settle, instead of implying full coverage.
- Risk coverage and gaps. The review names which material risk areas had a qualified perspective and which did not. An area with no qualified reviewer means the review is incomplete, not passed.
- Basis for confidence. A confidence claim names the evidence it rests on. A number with no scale and no validation record is self-assessment, not measurement.
Agreement with the first answer satisfies none of these.
Who owns the decision after AI recommends a plan?
The authorized decision owner, every time. Reviewers -- human or model -- recommend. They do not inherit authority over funding, priority, customer commitments, release timing, production execution, or secrets.
That split only holds if silence has a stated path. Leave it unstated and one of two things happens: the decision stalls while everyone waits, or somebody downstream reads the delay as permission and authorizes themselves. So every routed decision carries a deadline, sits in a queue the owner actually reads, and states its default. Ours defaults to no authorization.
It also helps to name what never counts as authorization: two models agreeing, a high confidence score, a passing review, a chat message reading "looks good," or nobody objecting. The decision is recorded and authorized before the build starts, by the person accountable for the outcome.
What happens when the evidence cannot settle the options?
Stop averaging and start testing. Competing positions stay on the record with the evidence each one rests on, because a split smoothed into a middle answer hides the thing worth knowing.
Then name the smallest reversible probe and the observation that would settle the split. Run it, and a result decides instead of a preference. Before any decision you cannot walk back, also run a forward test proportional to the stakes: assume this fails in six months, and say what caused it. That premortem runs before the decision, not after the commitment.
WHY: What does a defensible review protect for your business?
A defensible review protects the decision you are about to fund. When the premise behind a recommendation is stale, the money follows the stale premise: you authorize a tranche that was never the real tranche, and the missing work surfaces later as an unplanned commitment nobody budgeted. The review is what stands between an agreeable answer and a wrong funding decision.
It also protects the time you spend deciding. A decision that was never made inspectably gets made again -- the same scope re-planned next quarter, the same argument re-run in front of leadership, the same executive hours spent re-deriving a conclusion that already exists somewhere. Each repeat consumes tokens and compute on reviews that only grade an old assumption, and each repeat wears on the team that has to rebuild work it thought was settled. Left alone, that pattern reaches your customers as a commitment you cannot keep, and it quietly decays compliance: a record that nobody trusts stops being consulted, and controls that are never consulted stop being controls.
Related reading: The five-field form that made 32 architecture decisions citable shows why the decision record stays useful long after the meeting ends, which is exactly what keeps a settled decision from being re-argued.
Artifacts
- Multi-persona holistic review pattern. The maintained pattern file is a released, maintained review format showing two review gates, four mandatory perspectives, and one shared output structure.
That artifact is narrow on purpose. It does not implement the bounded packet described here, does not prove reviewer independence, a frozen evidence revision, the required answer fields, or the outcome of the case above, and makes no claim about how well it works. Its perspectives review the same pipeline's own output, which is the self-review framing this article corrects.
For the technical procedure that turns this decision boundary into a review packet, read How do I scope an AI review panel that can actually be wrong?.
Bottom Line: What should you require before authorizing the work?
Skip the packet once and you pay for it repeatedly: the same scope correction found again next quarter, leadership hours spent re-deriving a decision already made, review tokens and wall-clock time burned grading a stale premise, team frustration as people rebuild work they believed was settled, the risk that a repeat of the same miss reaches a customer as a commitment you cannot keep, and compliance decay as a record nobody trusts stops being consulted.
Before authorizing consequential technical work, require four things. Freeze the decision question and the evidence revision in writing before dispatch, because that is the earliest point where a stale premise is catchable; the reread before authorization is only a backstop. Require a reviewer other than the author, working read-only from that frozen evidence. Require named alternatives, consequences, evidence limits, and a stated basis for confidence, and treat a missing risk perspective as an incomplete review. Record the decision with its authority, exclusions, outcome check date, and reopen condition, signed by the accountable owner.
Price the exposure with your own numbers, not a borrowed benchmark: seven omitted stories multiplied by the loaded cost per story you already accept. We do not know that rate, so we will not invent the total.
The outcome this protects is simple. I want leaders to authorize the smallest defensible next step, so engineering hours go to new work instead of to scope that was ordered short. Share this with the person who owns the next AI-assisted decision, then record the field that is hardest to make explicit in your environment.
Frequently Asked Questions
Does asking multiple AI models the same question make the answer more reliable?
Only when the second model has a real chance to reach a different conclusion. If both models receive the same stale plan, or the second model sees the first model's answer, agreement can repeat the same mistake. Give the second model the current source evidence, withhold the first recommendation, and ask it to challenge the premise. For decisions that are expensive to reverse, compare the evidence each model cited and keep final approval with the accountable owner.
Who should sign off after AI recommends a technical plan?
The person who controls the budget and answers for the outcome signs off, and that name goes into the packet before the review starts, not after the verdict arrives. Reviewers, human or model, recommend only, and nothing in a review confers authority over funding, priority, customer commitments, release timing, production execution, or access to secrets. Write the owner's name, the response deadline, and the default on silence into the packet today; make that default "no authorization" so a missed deadline stalls the work instead of releasing it.
What should I do when AI reviewers disagree?
Record both positions with the evidence each rests on and refuse to split the difference, because the averaged answer hides the disagreement that mattered. Name the smallest reversible test that would settle it and the single observation that decides the result, then run that test first. Use this rule of thumb: if the probe costs more than reversing the wrong choice would, skip the probe and have the owner decide explicitly, recording the uncertainty that remains and the condition that would reopen it.
When is a formal AI review worth the effort?
Apply three tests: the decision is expensive or slow to reverse, it crosses a team or customer boundary, or it rests on evidence the decision owner cannot personally grade. One test met calls for a one-page packet and a single independent reviewer; two or more calls for broader risk coverage and a written decision record. A reversible choice you can undo in an afternoon, such as a config default nobody depends on yet, does not need a review at all -- just make it and move.
Can a confidence score make an AI recommendation trustworthy?
Treat any confidence number as an uncalibrated self-assessment unless a documented scale and a validation record show what past scores at that level actually predicted. Instead of the score, ask which premise would have to be wrong for the recommendation to fail, and grade the answer on challenged premises, alternatives with consequences, and stated evidence limits. If the reviewer cannot name that premise, the score is decoration and the review is not finished.
Join for free for practical AI operating lessons you can use with your team.
Want to work together or talk directly? Contact me.