Gopinath Polavarapu on why frontier AI benchmarks don’t matter in supply chain

Every few months, a new AI model claims the top spot on the leaderboards, and every few months, supply chain leaders feel the same pull to chase it. But what determines whether AI works in procurement isn’t how it scores on a reasoning test. It’s whether the model fits the shape of the work, comes with a governance structure that can survive an audit, and understands the business it’s being dropped into. Get those three things right and the choice of the ‘best’ model becomes much less important.

The governance gap

Every decision in procurement eventually needs a paper trail – if a system flags a supplier as compliant, recommends an approval route, or raises a contract risk, someone (an auditor, a regulator, an internal control team) may need to retrace exactly how that conclusion was reached. This gap is wider than most organizations realize: a 2026 Compliance Week survey found that 83 percent of organizations now use AI tools, but only 25 percent have strong governance frameworks around them.

Gopinath ‘GP’ Polavarapu
Gopinath ‘GP’ Polavarapu

It’s easy to conflate model-level safety with process-level governance, but they aren’t the same thing. Today’s frontier models have safeguards such as behavioral guardrails and data-handling protections built in, but a perfectly well-behaved model can still be slotted into a workflow that produces no audit trail at all, simply because nobody built traceability into the process around it. The EU AI Act makes this concrete: its obligations around high-risk AI systems, such as automatic record-keeping, traceability of outputs and demonstrable human oversight, describe an architecture, not a model configuration, and the compliance burden lands on the organization deploying the system, not the model vendor. Governance of this kind needs to be designed in from the outset.

Traceability is only half of the problem, though. The other half is integration, since no model, however capable, arrives already knowing your business. It doesn’t know your specific ERP configuration, your approval hierarchy, or years of accumulated policy exceptions unique to your organization. Building and maintaining that context is a separate undertaking altogether, and it’s where supply chain leaders should be focusing their evaluation efforts.

Why frontier models don’t fit routine procurement

None of that shows up on a leaderboard, which is why the industry’s obsession with frontier benchmarks looks misplaced. Every time a new frontier AI model launches, the web is flooded with benchmark comparisons and demo reels, and a familiar debate reignites about whether machines have reached some new height of reasoning ability. Very little of it speaks to whether the model fits the way supply chain teams work day-to-day.

Frontier reasoning models have been designed with software engineering projects, open-ended research, and multi-step planning in mind, work that requires a model to explore dead ends, backtrack, and rethink its approach across long sessions – and the benchmarks that rank them measure precisely that. Most day-to-day supply chain operations, by contrast, are largely repetitive and rule bound. Reconciling purchase orders against invoices and receipts, checking new suppliers against a fixed compliance checklist, or pushing approvals through the same delegation chain every time does not need freewheeling problem-solving. What Source-to-Pay needs is consistency: reliable output, high throughput, and costs that don’t swing wildly from one transaction to the next.

Applying reasoning-heavy models to routine transactional work also carries a direct cost. Claude Fable 5, for example, runs at $50 per million output tokens, exactly double Claude Opus 4.8’s $25 standard rate. At enterprise transaction volumes, a model that reasons through a routine task for thirty seconds before responding is not cheaper simply because it reaches the same answer. The correct unit isn’t cost per token; it’s fully-loaded cost per complete transaction, measured at your volume peak. In practice, disciplined context caching alone can reduce consumption on a high-volume inference workload by roughly two-thirds, with no change to output quality. That saving comes from engineering rather than from model choice.

Two-layer architecture that works

None of this means frontier-level reasoning has no place in supply chain and procurement. It’s genuinely valuable for judgement-heavy tasks such as spotting spend patterns across fragmented category data, shaping sourcing strategy, assessing risk across complex multi-tiered supplier relationships, and negotiating where the counterparty is adaptive and the optimal move isn’t obvious. The mistake is assuming that raw capability translates into value everywhere, including in high-volume transactional work that looks nothing like those problems. The smarter approach uses frontier intelligence selectively, where judgement earns its cost, alongside a governed, auditable, ERP-native AI system that handles the routine, high-frequency bulk of transactions. Each layer does what it was built for, on a common substrate that logs and traces every decision the same way.

A competitive edge

A new, better model will appear every few months, but when frontier-grade intelligence becomes abundant and nearly free, it stops being a real differentiator for anyone. What remains scarce is everything that intelligence must plug into: clean, reconciled, permissioned procurement data, deep integration with the systems of record where transactions settle; an audit trail that satisfies a regulator and human oversight positioned where it changes outcomes.

Supply chain teams that have built their processes around governance, traceability, and ERP-native integration won’t need to chase the latest leaderboard.  For them, it will be as simple  as plugging the better model in and carrying on, business as usual.

Gopinath ‘GP’ Polavarapu
www.jaggaer.com

Gopinath ‘GP’ Polavarapu is a seasoned technology executive and AI strategist who has spent over 20 years turning emerging technologies into billion-dollar growth stories. As Chief Data & AI Officer at JAGGAER, he leads the enterprise-wide AI vision for one of the world’s largest Source-to-Pay platforms, driving innovation through agentic, predictive, and generative AI.