Leadership - 2026-10-09 - 9 min read
Choosing a First AI Pilot Across Three Workflows
Choose a first AI pilot with three synthetic workflows, a reusable assessment table, clear handoffs, and a baseline that includes review and rework.
A first AI pilot should start with a small task that has a clear boundary, a checkable result, and someone responsible for the handoff. Turn “support, sales, and operations can use AI” into specific tasks before choosing one.
Support might need a reply grounded in a return policy. Sales might need action items from meeting notes. Operations might need a list of missing form fields. All three can produce convincing demonstrations. Whether they are ready for a pilot depends on the available sources, the reviewer, and what happens when an output is wrong.
What I often see is that the starting point comes earlier than workflow selection. The owner wants AI adopted but has no clear entry point. Some employees already pay for AI tools themselves. Others wait for the company to provide accounts, and a third group doubts AI will help at all. Each group changes who should test first and which data a pilot can touch.
The method below is a proposal, illustrated with three synthetic scenarios. The people, data, and selection decisions are invented to explain the method. There are no measured business results or promised savings.
Check prerequisites before comparing candidates
A weighted score can hide an unresolved gap. An attractive output is of little use if no one owns its review.
Write down five things for each candidate:
- Task boundary. What is one input, and what should it produce? Does this task recur? “Help support” is too broad. “Draft a return-policy reply from approved sources for a person to review” is a testable scope.
- Data and sources. Which data can the test use, and for what purpose? Who owns the source version and its applicability? Keep real customer data out until its permitted use is confirmed; synthetic data can support preparation.
- Responsibility and handoff. Who checks the output and approves sending or execution? Who takes over when information is missing, the source does not answer, or an exception appears?
- Checkable output. What counts as correct or requiring rework? A support reply must trace to the policy; an action item must trace to the notes; a missing-field report must trace to field rules.
- Reversible action. Can the initial scope stop at a draft, recommendation, or list for confirmation? Sending, making commitments, and writing to a production system require separate human approval.
These are proposed minimum checks, not a universal standard validated across enterprises. A workflow with different consequences may need additional conditions.
An unresolved data scope still allows research and synthetic demonstrations. It does not establish readiness for a pilot using real work. If the source of an answer is unclear, fix that gap first; a stronger model will not close it.
Choose the first testers deliberately
I would start with people who already pay for AI tools out of pocket. Starting with them means the pilot does not begin by persuading anyone to use AI.
That choice has two costs worth naming up front.
First, it is a biased sample. Smooth results from enthusiasts say little about how employees waiting for accounts, or the skeptics, will fare. Treat the outcome as evidence about that group.
Second, personal accounts sit outside company control. Before the pilot, ask what company data has already gone into those tools. Run the pilot itself on accounts and data scopes the company has approved.
Put three departments in the same table
Every case below is synthetic. Owners, sources, and permissions are assumed conditions. No working times or success rates have been measured.
| Comparison | Support policy lookup | Sales meeting notes | Operations form checks |
|---|---|---|---|
| Small task | Draft a reply from approved policy | Draft action items from notes | Find missing or inconsistent fields |
| Input → output | Synthetic policy and question → draft with source and version | Synthetic notes → action items and fields to confirm | Synthetic form → exception list |
| Data and source status | Synthetic data usable; policy owner assigned; version and scope specified | Real customer data scope and responsible owner unconfirmed | Synthetic forms usable; field rules clear |
| Review and handoff | Check against policy; hand off missing answers, unclear versions, and exceptions | Check against notes; invent no customer commitments or action items | Check flagged exceptions; route cases that rules cannot decide |
| Reversible scope | Draft only; approve before sending | Draft only; no automatic email or CRM updates | Flag issues; do not modify source data |
| Alternative to a model | Compare policy search plus manual drafting | Compare a fixed note template | Check required fields, formats, and relationships with rules first |
| Manual baseline | Unknown; record it | Unknown; record it | Unknown; record it |
| Next step in this example | Select a bounded drafting pilot | Resolve the two gaps; use only synthetic demonstrations for now | Measure rules plus manual review; assess a model if unstructured exceptions remain |
Support is selected because this example assumes checkable sources, an owner, and a drafting boundary. It does not establish that every enterprise should start with support. A support workflow with confused policy versions and no reviewer would change the choice.
Sales is deferred because its data scope and ownership remain unresolved, not because AI cannot assist with this work. Operations also has usable data and clear field rules. It is held back for a different reason: fixed rules may be sufficient. Check how far the existing rules go before deciding whether a model adds value on unstructured exceptions.
If all three lack prerequisites, the selection result can be “resolve the gaps first.” Starting a pilot does not require naming a winner today.
Record the manual baseline before the pilot
The manual baseline remains unknown in all three cases because an illustration cannot replace measurement. Before a real-work pilot, record the manual method for the same task scope and compare it with the assisted method.
For each item, retain the task type, input scope, completion criteria, operator, and reviewer, and keep the four measurements below in separate records.
- Active human work and results for the manual baseline.
- Preparation, operation, review, rework, and item-specific maintenance for the assisted method.
- Model waiting time, including whether it actually occupies a person's work time.
- One-time setup, shared data preparation, and tool maintenance. Report them separately; explain any allocation to individual items and avoid double counting.
Include abandoned items and human handoffs, not just successful drafts. Check quality too, including whether completion criteria match the original and whether errors surface later with someone else. Faster drafting does not establish savings if review work has moved to another colleague.
Someone who completes the manual version first may be faster on the assisted version simply because they have already seen the input. Tasks with different difficulty cannot be pooled without qualification. Where feasible, use comparable work and alternate the order; retain sample sizes and differences. Small-sample results support that pilot, not an annual ROI projection.
Acceptance includes declining to answer and handing off
For support drafts, check that citations trace to the correct policy and version, applicability is correct, missing answers are not fabricated, and handoffs include a clear reason. Define normal cases, missing information, policy updates, exceptions, and out-of-scope cases. Retain actual outputs and review results.
These are proposed checks; no model tests have been run here. Passing synthetic tests would not establish real-work performance either. Synthetic tests can expose obvious gaps. A real pilot must also examine the data, work distribution, and cost of human review.
Anthropic's explanation of agent evaluations distinguishes tasks, trials, grading, and outcomes. An outcome is the final state of the environment when a trial ends: the flight reservation exists in the database, whatever the agent claimed in the transcript. A drafting pilot is not an agent task, but the same caution carries over by analogy. Producing a reply is only an output; passing review and being usable at the next handoff meets the agreed completion criteria. The source does not validate this article's three-department selection method.
There are four stop conditions, and the first takes priority. When any of them occurs, stop that scope and investigate.
- Data out of scope. Data is used where it should not be, such as the pilot touching data outside the confirmed scope.
- Unsupported affirmative answers. A scope that requires source-supported answers produces affirmative answers the sources do not support.
- Bypassed approval. An output is sent or executed without human approval.
- Quality below the bar. Faster work cannot compensate for failing an agreed quality condition.
Fill one page per workflow
Write “unknown” where evidence is missing. The table exposes gaps and records the selection reason; it does not calculate a weighted score.
| Field | What to record |
|---|---|
| Workflow and boundary | Department, recurring small task, exclusions |
| Input and output | Data format, desired output, completion criteria |
| Permitted data scope | Synthetic or real data, allowed use, unresolved permissions |
| Existing employee use | Who already uses personal AI tools, which tools, and what company data has gone into them |
| Source owner | Version, applicability, update responsibility; output approver and exception owner |
| Manual baseline | Actual minutes and sample count; unknown until measured |
| Output checks | How correctness and rework are determined |
| Errors and handoff | Missing answers, exceptions, error handling, responsible person |
| Reversible action | How to withdraw a draft; approval before execution |
| Rule-based alternative | Whether search, templates, formulas, or rules suffice |
| Missing prerequisites | Unresolved data, ownership, baseline, or testing |
| Pilot scope and reason | What to try, why it comes first, when to stop (cite the four stop conditions above or define your own) |
| Status | Research, synthetic demonstration, or real-work pilot; whether results exist |
For unresolved data scope, see data and retrieval in enterprise AI adoption. If a draft will lead to sending or writing, use write, send, and approval boundaries to define that next action.
Filling in the table matters more than starting a pilot quickly. Fill it for three candidates and check whether the selection reason still holds. If you share feedback, the useful details are which prerequisite blocked progress, which exception changed the choice, and whether review work matched your expectations. Share only public, de-identified descriptions; actual customer records are unnecessary.
Working on something like this?
I help teams ship AI-native systems — architecture, governable autonomy, and the evidence discipline to back them. One conversation is enough to see whether it fits.
Discuss fit