The automation boundary: where accumulated error and supplier power rule out autonomous AI
By EXOS Research Team
Two variables decide whether a procurement category can be executed autonomously or has to be augmented: the number of sequential steps in the process, and the negotiating power of the supplier on the other side of it. Sequential error compounds as P to the power of n, so a chain of several dozen dependent steps degrades sharply even at high single-step accuracy; and where the supplier holds leverage, an automated counterpart has nothing to press with. Routine and Leverage spend sits below both thresholds and should be automated. Strategic and Bottleneck spend sits above them and should be augmented.
- Applies to
- Any category where the automation decision is open — the boundary runs in both directions
- Unit of analysis
- One category, characterised by step count and supplier power
- Position in the lifecycle
- Before a tool is selected for the category, not after
Autonomous sourcing tools are commonly demonstrated against a five-step transaction: requisition, RFQ, bid collection, price comparison, purchase order. A Strategic or Bottleneck category deal is not five steps. Once total cost of ownership modelling, should-cost decomposition, negotiation preparation and pre-contractual regulatory checks are counted individually, a genuine strategic procurement project runs to several dozen discrete steps, built up here from named, sourced methodologies rather than asserted as a single figure. Two variables determine whether autonomous execution or human-in-the-loop augmentation is the appropriate mode for a given category: the number of sequential steps in the process, and the negotiating power of the supplier on the other side of it.
A five-step demonstration and a two-hundred-and-thirty-one-variant reality
Vendor demonstrations of autonomous sourcing typically show a short, linear chain — request, quote, comparison, award — because a short chain demonstrates well. The published evidence on how procurement processes actually run does not support the five-step model even for transactional purchasing, which every vendor in this category treats as the easy end of the spectrum. An academic analysis of the BPI Challenge 2019 dataset — a real purchase-to-pay event log covering 150,370 logged cases at a single company — found 231 distinct process variants. The single most common variant, the clean linear path a demonstration would show, accounted for 2.8% of cases. A separate, unrelated case study of a comparable finance/procurement process (TKE, via Appian) found 232 variants. Two independent datasets converge on the same order of magnitude: the tidy short-chain story is not a simplification of what happens, it is a description of a small minority of it — and this is before any category-specific complexity is added.
Accumulated error compounds with every additional step
Where a process consists of n sequential, dependent steps and each step is completed correctly with probability P, the probability of the whole chain completing without error is:
Probability of error-free completion = P^n
P is the single-step success rate, expressed as a number between 0 and 1. n is the number of sequential steps in the process.
This identity is not specific to procurement. It is the standard treatment of pipeline reliability in current AI-agent research, where it is used to explain why systems that test well on short demonstrations degrade sharply on longer, real-world workflows. Independent analyses converge on the same figures from the same formula: at 95% single-step accuracy, a ten-step chain succeeds roughly 60% of the time and a twenty-step chain roughly 36% of the time; at 99% single-step accuracy — a high bar by current standards — a hundred-step chain still succeeds only around 37% of the time. METR’s empirical work on long-horizon task completion found frontier models achieving close to 100% success on tasks that take a human under four minutes, falling below 10% success on tasks beyond roughly four hours, with the threshold at which models succeed half the time roughly doubling every seven months over six years of tracking. A 2025 study from Cambridge, Stuttgart and the Max Planck Institute for Intelligent Systems adds a further mechanism: a self-conditioning effect, where a model’s per-step error rate rises as a task lengthens, because errors already present in its own prior output measurably increase the likelihood of further errors — meaning the plain P^n formula is, if anything, an optimistic bound rather than a pessimistic one.
None of this literature measures the accuracy of autonomous procurement-negotiation systems specifically — no source in this section is a study of Keelvar, Pactum or any other named autonomous-sourcing vendor, and the figures above should be read as illustrative accuracy assumptions common in published AI-agent reliability research, not as a measured property of any procurement product. Gartner separately projects that 40% of enterprise applications will include task-specific AI agents by the end of 2026, up from under 5%, and — in the same body of forecasting — that more than 40% of agentic AI projects will be cancelled or paused by the end of 2027. Both figures are worth carrying together; citing adoption without the accompanying attrition forecast would be a one-sided reading of Gartner’s own position.
Where the steps come from: the base process, plus what a Strategic or Bottleneck deal adds
The near-universal reference framework for sourcing — developed by A.T. Kearney in 2001 and taught as the standard model since — describes seven phases. Category and opportunity assessment; supplier market assessment; supplier survey; sourcing strategy; RFP and bid solicitation; negotiation and selection; implementation and benefit tracking. Practitioner descriptions of the framework document three or more named sub-activities within most phases — confirming stakeholders and usage data in phase one, assessing market dynamics and supplier risk in phase two, and so on — which puts the generic skeleton, before any category-specific work, at roughly 20–25 discrete steps on its own.
A Strategic or Bottleneck category deal adds several technical modules on top of that skeleton, each independently documented:
Step count of a Strategic or Bottleneck deal: base sourcing skeleton plus technical modules| Module | What it requires | Documented step count |
|---|
| Total cost of ownership modelling | Six cost categories — Quality, Management, Supply, Service, Communication, Price — each requiring separate data-gathering and estimation | 8–12 |
| Should-cost / specification decomposition | Six cost drivers — materials, labour, conversion cost, overhead, logistics, profit — each sourced separately | 6–10 |
| Negotiation preparation (BATNA/ZOPA) | Objectives, fact base, BATNA, ZOPA, risk terms, approval gates — several of which are themselves multi-step | 6–10 |
| DORA pre-contractual concentration checks | Concentration measurement, single-versus-sole-source classification, alternative-supplier verification, exit-strategy documentation, criticality assessment | 5–8 |
| Multi-departmental approval gates | Sequential sign-off across legal, finance, IT/security and the business owner | 4–8 |
| Total, base skeleton plus modules | | 49–73 |
This is a constructed sum, not a single citation, and it is presented as one deliberately: a reader can verify the total by checking each row rather than taking the figure on trust, which is the same auditable-arithmetic standard the rest of this section uses. The base sourcing framework is a macro model rather than an audited task list — the sub-activity counts above are drawn from practitioner descriptions of it, not from a published task-level breakdown A.T. Kearney itself has issued. The should-cost and negotiation-preparation module counts reflect converged practitioner methodology rather than a single peer-reviewed source.
A second variable moves independently of step count: what the supplier will accept
Step count measures internal process complexity. It says nothing about whether the party on the other side of the negotiation will engage with an automated counterpart at all, and the two variables move independently — the deep-dive on this second variable, with its own evidence base, is supplier power and automated negotiation. A landlord renewing an office lease, a specialised component manufacturer with no qualified alternative, and a hyperscale cloud provider are not price-takers responding to structured bid requests — each holds leverage the buyer’s software does not change by existing. Where supplier power is high, published procurement practice describes the applicable response as category engineering: relaxing over-specification, decomposing the supplier’s cost structure, establishing a costed and credible alternative, and restructuring multi-year indexation terms — not automated counter-offer generation. Where supplier power is low and alternatives are genuinely interchangeable, this constraint does not apply, and autonomous, high-volume negotiation is the better-matched tool.
Two variables, four positions, and where each sits on the Kraljic matrix
Step count and supplier power are not a new classification competing with the Kraljic portfolio matrix already used across this site — they are the mechanism that explains why the matrix’s boundary sits where it does. Kraljic sorts categories by financial impact and supply risk; step count and supplier power are how that position translates into an automation decision for a given category.
Automation mode by Kraljic quadrant, with the system class matched to each| Kraljic quadrant | Typical step count | Typical supplier power | Appropriate AI mode | Representative system class |
|---|
| Routine | Low | Low — many substitutable suppliers | Automate | Transactional & Intake AI |
| Leverage | Low to medium | Low — competitive market, buyer holds leverage | Automate | Autonomous Sourcing |
| Bottleneck | Medium to high — regulatory and continuity checks dominate even where contract value is small | High — few or no qualified alternatives | Augment | Scenario-Based Procurement Analytics |
| Strategic | Highest — every module in the table above is typically active | Highest | Augment | Scenario-Based Procurement Analytics |
EXOS Coexistence Overlay: Kraljic Portfolio Matrix with automation mode by quadrant. Automation fits the lower-left; human-in-the-loop augmentation fits the upper-right.
This table is the classification underlying the EXOS Coexistence Overlay, published on the procurement iceberg and the Kraljic matrix and set out in commercial terms on win your Strategic and Bottleneck categories. Directional descriptions (low/medium/high) are used rather than fixed numeric ranges per quadrant, because no source in this section measures step count broken down by Kraljic quadrant specifically — the 49–73 total in the table above is an aggregate figure for a Strategic or Bottleneck category taken together, not a per-quadrant measurement.
What this does not argue
This is not an argument against automation generally, and the boundary runs in both directions. Routine and Leverage spend — low step count, low or no supplier leverage — is where full automation is the correct match, and procure-to-pay and autonomous-sourcing platforms exist specifically for that zone; nothing above recommends adding human-in-the-loop review to a standard-specification, competitively-supplied purchase order. Nor does higher single-step accuracy make the boundary disappear: the worked figures above show that even a 99%-accurate system, well above what current published benchmarks report for long-horizon agentic tasks, still fails on more than half of an 80-step chain. The boundary moves with accuracy, but at every accuracy level published to date, it does not reach into genuinely Strategic or Bottleneck territory.
Where this runs in EXOS
The scenario chain typically used for a Quadrant 4 (Strategic/Bottleneck) category: TCO Analysis, followed by Should-Cost / Cost Breakdown Analysis, followed by Negotiation Preparation (BATNA/ZOPA), with Single vs Multi-Source Analysis where concentration risk is present. Ongoing exposure between projects is tracked by the Inflation Monitor and Risk Monitor platforms. See the six-class taxonomy this table draws on at Types of AI systems in procurement: a classification.
Sources
- A.T. Kearney, Strategic Sourcing (2001; restated 2004) — the seven-phase reference framework. Sub-activity counts per phase are drawn from practitioner descriptions of the framework, not from a Kearney-published task-level breakdown.
- Ellram, L.M. and Siferd, S.P., "Purchasing: The Cornerstone of the Total Cost of Ownership Concept," Journal of Business Logistics, 1993 — the six TCO cost categories.
- Ferrin, B.G. and Plank, R.E., "Total Cost of Ownership Models: An Exploratory Study," Journal of Supply Chain Management, 2002 — hidden-cost categories in technology procurement.
- Regulation (EU) 2022/2554 (DORA), Articles 28–30 — pre-contractual concentration-risk obligations.
- BPI Challenge 2019 dataset; independent analysis in arXiv:2409.11294 — 231 process variants in a real purchase-to-pay event log, dominant variant = 2.8% of cases.
- Appian/TKE case study — 232 variants, comparable finance/procurement process.
- METR, Measuring AI Ability to Complete Long Tasks (2025) — long-horizon agentic task success rates and the reliability-horizon doubling trend.
- Sinha, A., Arun, A., Goel, S. et al., "The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs," arXiv:2509.09677 (2025), NeurIPS 2025 — the self-conditioning effect.
- Gartner, press release, 26 Aug 2025 — 40% of enterprise applications to include task-specific AI agents by end 2026; and Gartner’s separate projection that over 40% of agentic AI projects will be cancelled or paused by end 2027.
- Kraljic, P., "Purchasing Must Become Supply Management," Harvard Business Review, 1983 — the portfolio matrix.
Should-cost driver structure (materials, labour, conversion cost, overhead, logistics, profit) and the BATNA/ZOPA preparation checklist (objectives, fact base, BATNA, ZOPA, risk terms, approval gates) reflect converged procurement-advisory practice rather than a single peer-reviewed source, and are presented as such.