A pilot should answer a decision: is this workflow useful and supportable under the proposed scope? It needs representative cases, an agreed baseline, named reviewers, and acceptance criteria written before the results are known.
Write the decision and boundary
State what the pilot will prepare or change, which cases it includes, which cases it excludes, and who owns the outcome. “Automate our inbox” is too broad. “Prepare a reviewed routing record for project enquiries in one mailbox” gives the team something to assess.
Define permitted actions and account access. Use the minimum scope needed for the pilot, with external sends and system commitments behind the agreed review controls. Document how the workflow can be paused and how the team will handle work while it is paused.
Preserve a baseline and a case set
Measure the existing process using the same business outcome you intend to measure later. Record active handling time, waiting time, rework, and unresolved cases. Include normal variation, not just the records most likely to succeed. Keep difficult cases available so changes can be assessed against a consistent set.
In an illustrative request-routing pilot, the set includes complete requests, missing customer references, forwarded messages, duplicate submissions, attachments that cannot be read, and messages containing two requests. The process owner defines the expected result for each case before reviewing the automation’s output.
Make acceptance criteria observable
| Area | Question to answer | Evidence |
|---|---|---|
| Correctness | Does the output match the agreed case decision? | Labeled cases and reviewer corrections |
| Control | Can an unapproved action proceed? | Rejection and changed-version tests |
| Usefulness | Does the team spend less effort completing the same outcome? | Handling and review time measurements |
| Recovery | Can failure be resolved without duplicate effects? | Timeout and rerun tests |
| Ownership | Who handles each unresolved condition? | Queue and runbook handoff |
| Cost | Is recurring usage within the agreed assumption? | Measured processing and support costs |
Set the required threshold or outcome for each criterion with the process owner. Avoid inventing a universal accuracy target. A field used for internal sorting and a field used to commit a purchase may require different controls.

A pilot answers one written decision.
- Baseline
- Observe the current work
- Case set
- Include failures and exceptions
- Acceptance
- Agree observable checks
- Handoff
- Assign support and recovery
View data
| Evidence | Meaning |
|---|---|
| Baseline | Observe the current work |
| Case set | Include failures and exceptions |
| Acceptance | Agree observable checks |
| Handoff | Assign support and recovery |
Illustrative operating model. Apply your organization’s controls.
Download imageTest behavior, not only successful runs
Microsoft documents design checks, flow testing, and static results for mocking action outcomes in Power Automate. It also warns that resubmitted runs can duplicate records or emails. Those features support testing, but a successful platform status does not establish that the business result is correct. Microsoft’s cloud flow testing guidance.
Test denied access, invalid input, unexpected category, unavailable integration, duplicate event, rejected approval, expired review, and timeout after a possible write. For each, specify the expected state, the information preserved, and the person who receives the next task.
Separate observation from controlled execution
An observation stage can prepare suggestions alongside the existing process without sending or committing them. Compare the suggestions with real decisions. Then move to a controlled stage where authorized reviewers approve defined actions. The stages should have different permissions and clear entry criteria.
Record manual interventions during both stages. If the builder repeatedly repairs records behind the scenes, include that effort in the result. Otherwise the pilot can appear reliable while depending on support that will not be present during rollout.
Review results without changing the question
Summarize case coverage, correct outcomes, exceptions, review time, waiting time, support interventions, and recurring cost. Identify failures by cause and consequence. If the scope or implementation changes, record the version and rerun the relevant cases rather than blending incompatible results.
Choose one of three outcomes: proceed with a defined rollout scope, revise and retest named gaps, or stop because the evidence does not justify the approach. Stopping is a valid pilot result when it prevents a larger unsupported commitment.
Prepare the handoff before rollout
Provide the process owner with a runbook, account and permission map, exception queue, change process, monitoring owner, and pause procedure. Use the readiness tool to identify gaps, the ROI tool to update assumptions, and the decision checklist to record unresolved conditions. Rollout should be approved against that concrete operating plan.