From AI ROI Hypotheses to Kill Criteria: Managing an AI Initiative Portfolio
Your AI steering committee has seven green pilots. The models hit their evaluation targets, the demos worked, and every sponsor wants another funding tranche. The combined request is €500,000. Nobody can explain what evidence would make the answer no.
I would not approve the next tranche from a progress report. I would ask each initiative to prove that the evidence agreed before the pilot started has crossed the threshold for another commitment of money, data access, integration capacity, and operational risk.
A pilot is a purchased option to learn. It is not a miniature implementation entitled to survive.
Green is a delivery color, not an investment decision
Section titled “Green is a delivery color, not an investment decision”AI portfolio reviews often collapse four different questions into one:
| Decision dimension | What it asks |
|---|---|
| Technical feasibility | Can the system perform the task well enough under defined conditions? |
| Operational viability | Can it work inside the real workflow with usable data, ownership, support, and controls? |
| Business value | Does the outcome improve enough to justify the total cost and displaced work? |
| Acceptable risk | Can you operate it within your safety, legal, quality, and reversibility limits? |
Passing one dimension does not prove the others.
A demand-forecasting model can reduce validation error while planners ignore its recommendations. A contact-center assistant can generate accurate answers while adding verification time to every call. A coding agent can increase output while review load and incidents rise. Technical progress is useful evidence. It is not the funding decision.
That distinction matters because AI adoption is moving faster than value governance. McKinsey’s 2025 global survey found that 78% of respondents reported AI use in at least one function, while fewer than one in five said their organizations tracked KPIs for generative-AI solutions. More than 80% reported no tangible enterprise-level EBIT impact from generative AI. The issue is not a shortage of pilots. It is weak evidence connecting pilots to outcomes. (McKinsey, The State of AI 2025)
RAND reached a similar conclusion from interviews with 65 experienced AI practitioners. Among the 50 participants from industry, 84% cited at least one leadership-driven cause as a primary reason projects failed, including solving the wrong problem, choosing the wrong metric, and failing to fit the workflow. Thirty of those 50 participants also described persistent data-quality problems. A model review will not fix a decision contract that was never defined. (RAND, The Root Causes of Failure for Artificial Intelligence Projects)
Kurio Nova’s AI value decision framework helps you decide whether an idea deserves a pilot. The AI ROI measurement guide helps you connect spend to outcomes. The missing layer is the continuation decision: what must be true before the initiative receives more resources?
Give every initiative a funding contract
Section titled “Give every initiative a funding contract”Before work starts, write the decision that the evidence will have to support. I call this a funding contract.
It is not a procurement contract, and it is not a promise that the pilot will scale. It is a short agreement between the sponsor, delivery owner, risk owner, and funding authority about what the pilot is buying: evidence for a bounded next decision.
| Contract field | Question to settle before the evidence window begins |
|---|---|
| Outcome hypothesis | Which outcome should change, for whom, by how much, and through what mechanism? |
| Baseline or counterfactual | What happens without the initiative, and how will you observe the difference? |
| Evidence window | How much time, volume, or exposure is needed for a decision-ready signal? |
| Continuation threshold | What minimum result earns the next tranche? |
| Kill criterion | Which result, cost, risk, or missing dependency ends the current hypothesis? |
| Risk guardrails | Which safety, compliance, quality, or reversibility condition cannot be traded for ROI? |
| Authority and tranche | Who decides, and what is the maximum next commitment? |
| Revisit trigger | What change invalidates the decision before the planned review? |
The evidence window should match the signal. A high-volume document workflow may produce a stable result after thousands of transactions. A fraud or revenue use case may need a seasonal comparison. A capability whose value depends on adoption may need enough time for training and behavior change.
Do not copy a universal 30-, 60-, or 90-day rule. Pre-commit to the amount of evidence that makes your decision credible.
DARPA’s Heilmeier Catechism follows the same logic at program level. It asks what problem is being solved, how success will be measured, what the risks and costs are, and what the mid-term and final examinations will prove. The value is not the wording. The value is defining the exam before seeing the grade. (DARPA, The Heilmeier Catechism)
Put three pilots under the same contract
Section titled “Put three pilots under the same contract”Consider three initiatives asking for the next tranche in the same portfolio review.
| Initiative | Evidence against the contract | Decision | Resource consequence |
|---|---|---|---|
invoice-exception-copilot | 12,400 exceptions processed; median handling time falls from 18.0 to 12.5 minutes; cost per resolved case falls from €11.40 to €7.80; adoption reaches 73%; rework remains below the 2% guardrail | Fund | Release €220,000 for two more business units and reserve integration and change capacity |
contact-center-answer-assistant | Average handle time falls from 7.4 to 6.5 minutes, but staffing, routing, and policy changed during the pilot; the result cannot be attributed to the assistant | Pause and reset | Cap further spend at €40,000, rebuild the comparison, and reopen after 20,000 comparable interactions |
demand-forecasting-model | Forecast error improves from 23% to 14%; planner acceptance stays at 18%; inventory days move from 34.2 to 34.0; projected integration cost rises from €180,000 to €410,000 | Stop | Retire the pilot and move data-engineering and business-change capacity to the stronger initiative |
The first pilot earns another tranche because its business and operational evidence crossed the agreed threshold without breaking the quality guardrail.
The second pilot is not a failure. It is an invalid experiment. Continuing at full speed would turn uncertainty into sunk cost, so the right answer is a bounded pause with a named evidence repair and a spending cap.
The third pilot is technically better and economically weaker. That is the decision most steering committees avoid. The model metric improved, but the people who must act on the output do not use it, the inventory outcome did not move, and the cost exceeded the agreed limit. More tuning will not repair a missing operating mechanism.
I’ve seen sponsors respond to this result by proposing a pivot during the review. A pivot can be valid, but it does not inherit the old pilot’s funding rights. It creates a new hypothesis, baseline, threshold, owner, and tranche. Otherwise, “pivot” becomes a polite word for moving the goalposts.
A portfolio review must reallocate something
Section titled “A portfolio review must reallocate something”Project governance asks whether each initiative is on plan. Portfolio governance asks where the next scarce unit of capacity creates the most justified option value.
Money is only one constraint. AI initiatives compete for data engineering, subject-matter experts, security review, integration windows, frontline attention, and executive sponsorship. Keeping a weak pilot alive can block a stronger one even when both budgets look small.
The U.S. General Services Administration’s 10x program provides a useful operating precedent. It funds work in phases and continues only while a project demonstrates impact and value. In 2022, 181 submitted ideas became 25 vetted ideas, and seven received Phase 2 funding. The program also describes “Celebrate No” as part of its investment culture. No further funding is a designed outcome, not an embarrassment. (GSA 10x, Phases and Process)
Your review should therefore close with a resource movement, not a dashboard update. In the example, the stopped forecasting pilot releases data-engineering capacity. The paused contact-center pilot releases most of its integration budget while preserving a bounded learning option. The invoice copilot receives the next tranche because it has earned priority over the alternatives.
If nothing moves, you did not make a portfolio decision.
Keep value approval separate from risk approval
Section titled “Keep value approval separate from risk approval”Kill criteria do not replace responsible-AI controls. They answer a different question.
The NISTNational Institute of Standards and TechnologyA United States federal agency that develops cybersecurity standards, guidelines, and vulnerability data infrastructure. AI Risk Management Framework organizes risk work around governing, mapping, measuring, and managing AI risks across the lifecycle. That is useful for deciding whether an initiative can operate within your constraints. It does not decide whether the initiative is worth another €220,000. (NIST AI Risk Management Framework 1.0)
For initiatives that can take action, AI Agents in the Enterprise: Where They Work, Where They Fail, and How to Decide provides a separate method for setting the autonomy ceiling.
Keep two approvals visible:
| Gate | Possible result |
|---|---|
| Value and portfolio gate | Fund, pause, pivot, stop, or grant a bounded strategic exception |
| Risk gate | Approved within guardrails, approved with controls, restricted, or prohibited |
A safe initiative can still be a bad investment. A valuable initiative can still be too risky to continue. Combining the gates creates two common errors: compliance becomes proof of value, or ROI becomes permission to waive safety.
I prefer an explicit strategic exception when evidence is slow but option value is real. Name the executive owner, cap the money and capacity, set an expiry, and state what evidence the exception is buying. “Strategic” without those fields is not an exception. It is an escape hatch.
Protect the contract from two opposite biases
Section titled “Protect the contract from two opposite biases”Hard thresholds can reduce escalation of commitment, but they can also kill useful work too early.
An experiment with 137 R&D managers found that visual decision aids and outside advice reduced continued funding of a losing course of action in a stage-gate setting. The practical lesson is to make evidence and prior criteria visible, and to include someone who is not defending the original proposal. (Behrens, Go/Stop Decisions in New Product Development)
The opposite error is premature termination. Slow-learning initiatives may look weak before their outcome signal matures. Helga Drummond’s work on decision error warns that stopping can destroy value too. Your objective is not a high kill rate. It is symmetric discipline: continuation and termination should face the same burden of evidence. (Drummond, Decision Error and the Risks of Premature Termination)
Use three protections. Freeze the contract before the evidence window begins. Match the window to signal latency. Require a new contract when the hypothesis changes.
Then record the decision in plain language: what evidence crossed or missed the threshold, what resources move now, who owns the outcome, and what event reopens it.
At your next AI steering review, choose three initiatives asking for more capacity. For each one, write the funding contract on a single page before opening the progress deck. Do not ask whether the pilot is green. Ask what evidence has earned the next tranche, and move the resources accordingly.
Article series
AI Value Decision System
Explore every resource in this series, whatever its format.