Skip to main content
Implementation

The 90-Day AI Pilot Roadmap: From Use Case to Production Decision

A week-by-week AI pilot plan covering workflow selection, data, prototypes, evaluations, integration, supervised use and the final scale-or-stop decision.

By Jayson Hao12 min read
Tactile paper and metal editorial artwork for The 90-Day AI Pilot Roadmap: From Use Case to Production Decision
Editorial field noteITL / № 10

Key takeaways

  • Define the business decision the pilot must support before selecting a model or vendor.
  • Spend the first 15 days on workflow evidence and the next 30 on feasibility and evaluation.
  • Use shadow mode and draft-only assistance before permitting business-system writes.
  • Finish with a production brief, not an open-ended extension of the pilot.

Before day one: write the decision brief

Name the workflow, business owner, users, customer affected, current pain, expected mechanism, permitted data, risk level, budget ceiling and decision date. State the final choice the pilot will support: scale, revise or stop. If stakeholders cannot agree on that choice, they are funding exploration rather than a pilot.

Choose one executive sponsor and one process owner. The sponsor removes organizational barriers; the owner defines acceptable work and changes the process. Add technical, privacy, security and legal reviewers according to risk rather than inviting a large committee to every working session.

Days 1–15: map the work and baseline

Observe the workflow, not only its procedure document. Sample 50 to 100 cases and mark inputs, decisions, tools, handoffs, exceptions, outcomes and time. Separate queue time from hands-on work. Record current volume, first-pass acceptance, error, rework, cost and customer result.

End day 15 with a narrow task definition and exclusion list. “Handle support” becomes “classify delivery-status tickets, retrieve the order and draft a cited response; exclude identity disputes, damaged goods and refunds.” The exclusion list prevents the prototype from quietly expanding.

Days 16–30: prove the riskiest assumption

Identify what could make the use case impossible: missing data, weak retrieval, unreliable classification, an unavailable integration or an unacceptable privacy condition. Build the smallest prototype that tests that assumption on representative cases. Avoid polishing the interface.

Use the most capable reasonable model to establish whether the task can work. Optimize model size, latency and cost after a quality baseline exists. Track every input, retrieved source, output and reviewer decision from the first prototype run.

Days 31–45: build the evaluation harness

Create a fixed set containing normal cases, difficult cases, known historical failures, incomplete inputs and adversarial attempts. Define expected evidence, acceptable result and severity before running the model. Use domain experts to label the cases and resolve disagreements.

Measure component and end-to-end performance. Retrieval may fail even when generation is good; a tool may be selected correctly with a wrong parameter; a correct draft may create no value if review takes longer. Set thresholds for quality, serious error, intervention, latency and cost per accepted outcome.

Days 46–60: integrate with limited authority

Connect identity, approved knowledge and the minimum required systems. Use read-only credentials first. Define typed tools, input validation, timeouts, retries, spend limits, audit logs and stop conditions. Add human approval before external messages or state changes.

Run security, privacy and failure reviews against the actual architecture. Confirm data retention, processing location, vendor training use, deletion, access boundaries and incident ownership. Update the risk classification if integrations expand the consequence of a failure.

Days 61–75: run shadow and supervised work

In shadow mode, process live work without affecting the outcome. Compare the proposed decision with what employees did and review disagreements. Then let a small trained group use drafts or approve actions. Provide a one-click way to report a problem and capture the correction reason.

Watch for automation bias and workarounds. Employees may accept fluent outputs too quickly or avoid the system when it adds friction. Sample accepted and rejected cases, interview users weekly and change the workflow or interface when the same issue repeats.

Days 76–90: measure, decide and prepare production

Compare pilot results with the baseline and any control group. Report accepted outcomes, quality by case type, serious failures, human intervention, cycle time, customer result, recurring cost, implementation cost and forecast payback. Separate observed value from projected annual value.

Scale only with a production brief covering owner, service level, support, monitoring, evaluation cadence, model-change testing, incident response, rollback, budget and employee training. Revise with a defined hypothesis and time box. Stop when value is weak, required data cannot be used or the safety threshold remains unmet.

Four reasons pilots never become products

The workflow had no owner; the prototype used clean demo data rather than real cases; success meant stakeholder excitement rather than a threshold; or the team postponed identity, integration and governance until after the demo. Each creates a pilot that proves a model can generate output but says little about operating value.

OpenAI’s 2026 scaling guidance emphasizes workflow design, governance, ownership and quality before scale. Microsoft’s 2026 workplace research similarly found organizational conditions more strongly associated with reported AI impact than individual effort. The 90-day plan must change the operating process, not only test model intelligence.

About the author

Jayson Hao

Founder of Innovation Trigger Lab and a University of Toronto Computer Science graduate with an AI/ML focus. He designs and ships production RAG systems, AI chatbots, web platforms and mobile products.

View profile

Frequently asked questions

Is 90 days enough for an AI pilot?

It is enough for a narrow workflow with accessible data, a named owner and sufficient case volume. Complex regulated integrations may need longer, but the first 90 days should still end with an evidence-based decision.

What should an AI pilot deliver?

It should deliver a measured baseline, representative evaluation set, tested prototype, operating-cost evidence, risk findings, user feedback and a scale, revise or stop recommendation.

What is the difference between an AI proof of concept and pilot?

A proof of concept tests whether the technology can perform a task. A pilot tests whether the system creates acceptable value, quality and risk in a real operating process with actual users.

Sources and further reading

  1. 1.
  2. 2.
    How enterprises are scaling AIOpenAI, May 11, 2026
  3. 3.
    2026 Work Trend IndexMicrosoft, May 5, 2026