When launching a new procurement platform, the natural instinct is to target the largest spend categories first to prove immediate ROI. However, this is the most common mistake organizations make when designing an intake pilot. A pilot is not a value exercise; it is a proof of concept designed to test whether your routing, data capture, and approval mechanisms actually work. By testing the system against the most complex categories right out of the gate, you risk conflating software failure with bad master data or internal politics. To ensure success, your intake pilot must focus on categories that are testable, not just important.
TL; DR
- Pick categories that will succeed, not categories that matter most. A pilot proves the process works, not that it saves money.
- Good candidates have high volume with low variation, a settled approval path, and clean master data.
- Capital equipment, professional services, and preference-driven categories belong in the rollout rather than the pilot.
- Three to five categories, with duration set by request volume rather than by calendar.
- See how Merlin Intake handles procurement requests. Request a demo.
Pick phase-1 categories that will succeed, not categories that matter most. The pilot is not a value exercise, it is a proof that the process works, and a pilot that fails on the hardest category proves nothing except that you chose badly.
Most intake pilots are scoped by spend, on the reasonable-sounding logic that the biggest category should go first. That reasoning is wrong for a pilot specifically, because the largest categories are usually the ones with the most exceptions, the most stakeholders, and the least standardization.
This is written for procurement leaders scoping a phase-1 rollout who have been asked why the pilot is not covering the categories that matter. A 30-day proof of concept is the usual shape.
What is a pilot actually testing?
That requests arrive complete, route correctly, and reach the right owner without manual intervention.
That is a process question, not a savings question. If the process works on three categories it will work on thirty with more configuration. If it fails on three, adding categories makes the failure larger rather than clearer, and harder to attribute to a cause.
Figure 1 contrasts the two approaches. Scoping by spend tests the process against its hardest conditions first, which is the opposite of what a pilot is for. You test whether the mechanism works before you test whether it works under pressure, which is the ordering every other kind of engineering uses.

Figure 1: Scoping by spend tests the process at its hardest first.
Gartner projects that by 2027, 70% of procurement intake requests will be AI-assisted.
Gartner was forecasting intake automation broadly rather than pilot scoping. It matters here because a pilot is where the routing behind that automation is proved, and proving it on the hardest categories first tests the wrong thing.
Which categories make good phase-1 candidates?
Three properties, and a good candidate has all three.
High volume, low variation. Enough requests to generate signal within weeks rather than quarters, with each request resembling the last. Software renewals, standard IT hardware, and office consumables usually qualify.
A settled approval path. Everyone already agrees who approves these and in what order. A category where approval is contested will produce a pilot that fails on governance, and the failure will be attributed to the platform.
Clean master data. Suppliers, items, and cost centers already exist and are reliable for this category. Master data problems surfacing during a pilot get read as intake problems, which is both unfair and difficult to argue against once the impression forms.
Categories meeting all three are usually unexciting, and that is the point.
There is a political argument for the easy categories that is worth making explicitly. A pilot that succeeds creates permission for the harder phases. A pilot that struggles, even for reasons unrelated to the platform, spends credibility that the difficult categories will need later.
This is not risk aversion. It is sequencing, and it is the same logic that puts intake before sourcing in an automation program: build the thing that makes the next thing easier.
Hackett measured the effect of generative AI on top-quartile operations rather than pilot scoping. The bearing on this argument is that returns of that size are realized at scale rather than in a pilot, which is the case for choosing pilot categories that establish the mechanism quickly.
Which categories should wait?
Capital equipment, professional services, and anything with strong internal preference.
Capital equipment has low volume and long cycles, so a pilot would end before enough requests accumulated to say anything. Professional services carries scope ambiguity and worker classification questions that need real design rather than a pilot configuration. Categories with strong preference, where individuals have relationships with specific suppliers, produce resistance that has nothing to do with whether the platform works and everything to do with what it makes visible.
Each of these belongs in the rollout. None belongs in the pilot.
The distinction worth holding is between categories that are hard to configure and categories that are hard politically. The first can be worked through. The second will consume the pilot’s credibility while it is being resolved.
One more property is worth checking before committing: whether the category has an engaged owner. A category where somebody cares about the outcome will surface problems early and help fix them. One where nobody owns the spend produces silence, and silence during a pilot is indistinguishable from success until the rollout says otherwise.
How many categories, and for how long?
Three to five categories, and long enough to see the volume drop.
Three is enough to test that routing differentiates rather than defaulting. Five is about the limit before configuration effort delays the start, and a pilot that starts late loses the momentum that justified it.
Duration should be set by request volume rather than by calendar. You need enough requests to see the pattern hold after the initial enthusiasm fades, which usually means several weeks past the point where the pilot appears to be working. Figure 2 shows the shape. Pilots reporting success at week four frequently look different at week eight, when real request variety arrives and the early adopters stop being the only users.
Merlin Intake operates as the front door for procurement requests and works inside Microsoft Teams and Slack, which matters for pilot design because adoption is not a separate change management workstream when the tool is where people already are.
Set the success criteria before the pilot starts and write them down. A pilot evaluated after the fact will be evaluated against whichever numbers looked best, which is how programs acquire a reputation for reporting success while nothing visibly changes.
Agree in advance what result would mean the approach is wrong. If no result could mean that, the pilot is a rollout with a smaller scope and should be described that way.

Figure 2: What week four reports, and what week eight shows.
How do you know the pilot succeeded?
By the numbers that predict scale, not the ones that describe the pilot.
First-pass completeness tells you whether the form design works. Category match rate tells you whether routing will hold as categories are added. Off-channel volume, meaning requests that still arrived by email, tells you whether adoption is real or reported.
That last one is the honest test and the one most often skipped. A pilot with excellent internal metrics and unchanged email volume has not proved that the process works. It has proved that the process works for the requests that entered it.
Then check whether anything broke that the pilot categories were chosen to avoid. If the answer is nothing, the pilot was scoped correctly and the next phase should add difficulty deliberately rather than all at once.
Forrester measured provider satisfaction rather than pilot scoping. It is included because a pilot that reports success while requesters remain dissatisfied is measuring the wrong thing, which is what the off-channel check is designed to catch.
Conclusion
A successful intake pilot creates the political permission and technical foundation needed for a wider rollout. By selecting three to five categories characterized by high volume, low variation, and clean master data, you isolate the process from unnecessary friction. Capital equipment, professional services, and highly contested spend can wait until the foundation is solid. Remember to measure success by tracking first-pass completeness and off-channel email volume, ensuring adoption is real before moving to the next phase.
Ready to see how a streamlined pilot can transform your purchasing experience? Request a demo to explore how Merlin Intake acts as a frictionless front door for all procurement requests.
Frequently asked questions
Q1. How do you choose categories for an intake pilot?
Choose for high volume with low variation, a settled approval path, and clean master data for those categories. The purpose is to test whether the process works, so the categories should isolate the process from every other variable that could cause a failure.
Q2. Should the pilot cover your largest spend categories?
Usually not. Large categories tend to carry the most exceptions, the most stakeholders, and the least standardization, which means a pilot there tests the process under its hardest conditions before establishing that it works at all. Large categories belong in the rollout that follows.
Q3. How many categories should a phase-1 intake pilot include?
Three to five. Three is the minimum to confirm that routing differentiates between categories rather than defaulting. Beyond five, configuration effort delays the start and the pilot stops being a fast test of the mechanism.
Q4. How long should an intake pilot run?
Long enough to see volume settle after initial enthusiasm fades, which usually means several weeks past the point where it first appears to be working. Set duration by request volume rather than calendar, since a low-volume category can run for months without producing enough requests to conclude anything.
Q5. Which categories should be excluded from an intake pilot?
Capital equipment, because volume is too low and cycles too long. Professional services, because scope ambiguity and worker classification need real design rather than pilot configuration. Any category with strong individual supplier preference, because the resistance produced has nothing to do with whether the platform works.
Q6. What metrics show whether a pilot succeeded?
First-pass completeness, category match rate, and off-channel request volume. The first two describe the pilot. The third describes whether adoption is real, and it is the one most often skipped because measuring it means counting requests that bypassed the system.
Q7. What is the most common intake pilot mistake?
Scoping by spend rather than by testability, which puts the process against its hardest conditions before anyone has established that the mechanism works. The second most common is declaring success at week four, before enough request variety has arrived to be informative.
Q8. How do you move from pilot to full rollout?
Add difficulty deliberately rather than all at once. The next phase should introduce one hard property at a time, whether that is contested approval, variable requirements, or weaker master data, so that when something breaks it is clear which property caused it.





















































