A useful AI automation pilot starts with a sentence an operations manager can verify: when this event happens, the system prepares this piece of work for this person, and we measure this outcome. If the proposal cannot be expressed that clearly, there is more discovery to do before anyone selects a model.
Consider an illustrative purchase-request workflow. A coordinator receives a request, checks the supporting documents and prepares a summary for an approver. A sensible first pilot might prepare that summary and flag missing information. The approver still decides. The pilot has a beginning, an end and an existing task against which its performance can be compared.
That is a more useful unit of work than a broad instruction to introduce AI across procurement. It gives the team something specific to accept, change or stop.
Find the smallest complete piece of work
Small does not mean trivial. A pilot should complete a useful handoff, such as turning an incoming service request into a draft job record that a dispatcher can check. Extracting a few fields without showing where the result goes next may demonstrate technology while leaving the operating problem untouched.
Write down the trigger, the input record, the expected output and the receiving person. Then describe what happens when a required document is missing, a request is duplicated or the intended owner is absent. These exceptions help reveal whether the proposed boundary is workable.
Our AI process audit framework provides the broader discovery structure. Pilot selection is the next decision: which part of that map should be put to work first?
Compare candidates using evidence
For each candidate, collect recent examples of ordinary work and difficult cases. Ask the people doing the task how long active handling takes, what causes rework and which mistakes have serious consequences. Record the uncertainty in those answers instead of converting estimates into apparently precise savings.
A practical selection review should cover:
- Frequency: does the task happen often enough to generate useful feedback?
- Verifiability: can a reviewer tell whether the result is correct from available records?
- Readiness: can the team lawfully access suitable input data and keep it current?
- Consequence: what happens if the system is wrong, late or unavailable?
- Ownership: who can change the process and resolve its exceptions?
- Integration effort: can the result enter the next stage without repeated copying?
Do not average away a serious weakness. A valuable task with unclear permission to use its data is not ready for a pilot. A frequent task with no accountable process owner needs an ownership decision first. Use the review to make these dependencies visible.
Assign AI a job, then define its authority
Separate interpretation from business rules. Reading an email and proposing a category may suit a language model. Checking whether a mandatory field is present, calculating a total or applying an approved limit should follow explicit rules. The pilot can combine both approaches.
For the purchase-request example, the system could prepare a brief containing the request, supporting evidence, missing fields and a suggested route. It should not acquire permission to approve expenditure merely because it can write a persuasive summary. The person reviewing it needs access to the underlying records, not just the generated prose.
Write the allowed actions into the design: read specified records, create a draft, attach evidence and route it to a named queue. Specify the actions that require approval. This is the operating boundary behind our AI workflow agents approach.
Design the acceptance test before the demonstration
Choose a representative sample that includes incomplete inputs, conflicting information and cases that should be escalated. Keep some examples separate from the examples used to develop the system. Otherwise the team may be measuring how well the pilot handles material it has already been tuned to recognise.
Agree what constitutes an acceptable result for each task. A correct category, a complete summary and a justified escalation are different outcomes. A single overall accuracy figure can hide a serious failure in one of them. Treat permission failures as release blockers rather than balancing them against faster processing elsewhere.
Measure review effort as well. If an approver has to reconstruct every source document to check the generated brief, the pilot may have transferred work rather than removed it. The test should show how a person spots an error and how the corrected record proceeds.
Run with a clear fallback
Begin with a mode appropriate to the consequence of an error. That may mean comparing prepared outputs against completed historical work, or creating drafts alongside the existing process before relying on them. Set a defined review window and collect enough varied cases to support a decision.
Name the person who can pause the automation. Preserve a way to handle work manually, and make paused or failed items visible in the same queue the team uses to manage the process. A silent failure that leaves work unowned is an operational problem even when no incorrect action was taken.
NIST describes its AI Risk Management Framework as voluntary guidance for incorporating trustworthiness into AI design, development, use and evaluation. The practical method in this article is Cyclotron's suggested working approach; it is not a certification or a claim of NIST endorsement.
Make an explicit decision about expansion
At the review, compare accepted outputs, correction effort, exception handling and operating cost with the baseline. Record whether the pilot should expand, remain limited, be redesigned or stop. Attach the evidence and identify who owns the decision.
Expand one boundary at a time. Handling another document format, serving another team and gaining permission to take an external action are separate changes. Each deserves its own acceptance criteria. The first pilot has done its job when the business can explain both the value it produces and the authority it has been given.

