01
Pick a Task with a Clear Beginning and End
Choose one repeatable task such as summarising an approved document, drafting a response from known facts, or finding relevant guidance. Name the user, input, output, and completion condition. Avoid a first brief that asks AI to manage an entire department.
Record the current baseline where possible: how long the task takes, what people check, and which errors create rework. The pilot should compare against that existing process. A faster draft is not necessarily a better result if review takes longer.
02
Prepare an Evaluation Set Before Tuning
Collect representative examples of the task, including incomplete information, ambiguous requests, and cases the system should decline. Use authorised or anonymised material. Keep some examples separate from the development process so they remain useful for evaluation.
Define what a good result means. Reviewers may check factual correctness, relevance, completeness, source references, tone, and whether the output is ready for the next step. A simple written rubric is often more useful than a single broad satisfaction score.
03
Decide What the System May Access and Do
Specify the source documents, business records, and tools that the pilot can use. Access should follow the project’s requirements, and the integration should not expose more information than the task needs. Treat instructions found inside external content as data to inspect, not authority to change the workflow.
For an agent, distinguish reading information, drafting a proposed action, and executing it. A person can review a draft before a message is sent or a record is changed. A narrow set of permitted actions makes failures easier to understand and recovery easier to plan.
04
Measure the Whole Task
Track the time spent preparing inputs, waiting for results, reviewing outputs, and making corrections. Record recurring failure types as well as successful examples. Include provider usage and the operational work required to maintain the source material.
A knowledge assistant should be evaluated on unanswered questions as well as answer quality. Correctly recognising that an answer is unsupported may be preferable to producing a fluent guess. A source link also needs to support the answer, not merely look relevant.
05
Agree the Decision at the End
Before the pilot begins, identify what evidence would support continuing, revising the approach, or stopping. The result may be a narrower use case, better source documents, more review controls, or a conclusion that conventional automation is sufficient.
Document what remains untested. A pilot used by a few staff with a small document set has not established behaviour at every scale, for every user, or under all production conditions. The next stage should address the missing evidence rather than silently treating it as solved.
06
Plan Production as a Separate Responsibility
Moving forward may require authentication, permissions, monitoring, fallbacks, cost limits, content maintenance, and support. Someone must own provider changes and evaluate whether a new model behaves differently on the important tasks.
Training also matters. Users should know the tool’s intended role, when to verify a result, and how to report a problem. With that structure, a pilot becomes a practical learning exercise that informs an investment decision rather than an isolated experiment.



