← All articles
- Autonomous Sourcing
- Procurement AI
- Category Management
- Risk
The demo runs five steps. Your contract doesn't.
Autonomous sourcing demonstrations show a four- or five-step process. A Strategic or Bottleneck deal runs to fifty or seventy, and accumulated error decides what that costs.
· 6 min read
By EXOS Research Team
Autonomous sourcing tools are usually shown doing the same thing: a request goes out, quotes come back, the software compares them, a purchase order is issued. Four or five steps, start to finish, and it works cleanly on stage. The problem isn't that the demonstration is dishonest. It's that almost nothing in a Strategic or Bottleneck category looks like that, and the gap between the demo and the deal is where the trouble starts.
Even ordinary purchasing isn't five steps
Before getting to strategic categories at all, it's worth looking at plain purchase-to-pay — the transactional end of procurement, the part every vendor in this space treats as the easy case. A study of a real company's purchasing event log, 150,000 logged transactions, found 231 different paths through the process. The single most common one — the clean, linear path a demo would show — covered under 3% of all cases. A separate, unrelated case study at another company found much the same thing: over 230 distinct variants in a comparable process.
That's the baseline, before a single supplier negotiation or cost model enters the picture. The five-step story isn't a simplification of reality. It's a description of a small minority of it.
A strategic deal adds total cost, should-cost, negotiation prep and a compliance check
A genuine Strategic or Bottleneck category deal isn't a longer version of the same five steps — it's several distinct pieces of work stacked on top of a sourcing process that was already more than five steps to begin with. A full total cost of ownership model has to account for six separate cost categories, not just the invoice price. A should-cost breakdown works through materials, labour, overhead, logistics and margin separately. Negotiation preparation — building a real BATNA, working out where the zone of possible agreement actually sits — is its own multi-step exercise, not a single meeting. And where a supplier concentration issue exists, EU rules now require specific checks to be documented before the contract is signed, not after something goes wrong.
Add those up against a standard sourcing process, using nothing but the step counts each of those methods already publishes on its own, and a real Strategic or Bottleneck project lands somewhere in the region of fifty to seventy steps. That's not a number pulled from a single study — it's a sum of parts that can each be checked separately, and the full breakdown is on the methodology page linked below.
Every extra step is another chance for the process to fail
This matters because errors in a multi-step process don't add up — they multiply. If a system gets each individual step right 95% of the time, which is a solid result by current standards, the maths for the whole chain looks like this:
| Steps in the process | Chance the whole chain completes without error |
|---|
| 5 | About 77% |
| 20 | About 36% |
| 80 | About 2% |
A five-step process with a 95%-accurate system fails roughly one time in four — tolerable for a low-value purchase order. An eighty-step process with the same accuracy fails roughly 98 times in a hundred. This isn't a procurement-specific claim; it's the standard way AI researchers currently explain why systems that look impressive on short demonstrations degrade sharply on longer, real-world tasks, and independent studies keep arriving at the same numbers. One recent analysis even found that errors compound faster than simple multiplication predicts, because a model that has already made one mistake earlier in a task becomes more likely to make the next one, not less.
None of this is a measurement of any specific procurement bot's accuracy — that figure isn't published anywhere for this category of tool. It's the general shape of what happens to any multi-step automated process as the chain gets longer, and a Strategic or Bottleneck deal is a long chain.
The supplier on the other end doesn't have to play along, either
Step count is only half of it. The other half is whether the counterparty engages with an automated system at all. A landlord renewing an office lease, a specialised component manufacturer with no real alternative, a cloud provider with a standard enterprise contract — none of them are price-takers responding to a structured bid request. Where a supplier holds genuine leverage, the response that actually moves price is decomposing their cost structure, challenging the specification that locked out other bidders, and building a credible alternative — not sending a better-worded automated message. Where the market is genuinely competitive and any of several suppliers will do, that constraint doesn't apply, and running the process on autopilot is the right call.
This isn't an argument against automation
Routine and Leverage spend — short processes, suppliers who compete on price — is exactly where full automation belongs, and the procure-to-pay and autonomous-sourcing tools built for that zone are the right tool there. Nothing here argues for adding a manual review step to a standard purchase order. The argument is narrower than "AI doesn't work": it's that step count and supplier power, not transaction volume, are what decide whether a category is a good match for autonomous execution or needs a person making the final call. Strategic and Bottleneck categories sit on the wrong side of both variables at once, and higher model accuracy doesn't change that — even a 99%-accurate system, well above what's currently published for long AI-driven tasks, still fails on more than half of an eighty-step chain.
The full step-by-step build-up, the maths behind the failure-rate table, and how this maps onto the Kraljic matrix are on the methodology page: The automation boundary: where accumulated error and supplier power rule out autonomous AI.
Sources: BPI Challenge 2019 purchase-to-pay dataset (231 process variants); Appian/TKE case study (232 variants); Ellram & Siferd (1993) and Ferrin & Plank (2002) on total cost of ownership; A.T. Kearney's strategic sourcing framework (2001); Regulation (EU) 2022/2554 (DORA), Articles 28–30; METR, "Measuring AI Ability to Complete Long Tasks" (2025); Sinha et al., "The Illusion of Diminishing Returns" (arXiv:2509.09677, 2025); Gartner, 26 August 2025. Full citations and the complete evidence table are on the methodology page above.