Why Agentic AI Fails at Enterprise ROI, and What Software Factories Do Differently
xSquad Team
Why Agentic AI Fails at Enterprise ROI, and What Software Factories Do Differently
Agentic AI fails to deliver enterprise ROI because most organizations deploy it as a faster way to do existing tasks instead of redesigning the workflows around it. The gap between experimentation and scaled impact keeps widening. In McKinsey's 2025 global survey, 88 percent of respondents said their organizations regularly use AI in at least one business function, but only 23 percent reported scaling agentic AI anywhere in the enterprise and just 39 percent saw any EBIT impact from AI source. The small group that succeeds, roughly 6 percent of respondents classified as AI high performers, treats agentic AI as a catalyst for workflow redesign rather than a coding assistant source. High-performing software factories close the same gap by combining end-to-end workflow redesign, outcome-based measurement, and senior ownership with delivery discipline.
Why do most agentic AI initiatives fail to show enterprise ROI?
The first reason is that pilots rarely leave the lab. McKinsey found that 62 percent of respondents are at least experimenting with AI agents, yet only 23 percent are scaling an agentic system anywhere in the enterprise and most of those are scaling in only one or two functions source. Agent use is most common in IT and knowledge management, but in any given business function no more than 10 percent of respondents say their organizations are scaling agents source. The technology is present, but it is not woven into the operating model.
The second reason is that AI is still measured like a cost experiment rather than a value engine. Only 39 percent of McKinsey respondents attributed any level of enterprise EBIT impact to AI, and most of those said AI accounted for less than 5 percent of their organization's EBIT source. Deloitte's 2026 Global Technology Leadership Study adds a sharper point: 79 percent of tech leaders cite driving business outcomes as their top priority, yet 42 percent report low or no ROI on AI investments source. Boards expect transformation; delivery teams are still proving point solutions.
The third reason is timeline mismatch. Deloitte's 2025 AI ROI survey found that 85 percent of organizations increased AI investment in the past 12 months and 91 percent plan to increase it again, but only 6 percent reported payback in under a year and even among the most successful projects just 13 percent saw returns within 12 months source. Agentic AI makes the problem worse. Among organizations already using agentic AI, only 10 percent said they are currently realizing significant ROI source. The tools are bought as quick wins; the value only arrives after structural change.
What is the real barrier: models or operating models?
The barrier is rarely the model. The barrier is the operating model. In Deloitte's 2026 study, 81 percent of tech leaders said their current operating model can deploy and govern AI enterprise-wide, yet 75 percent also said their organization must change its operating model within the next 12 to 18 months to drive greater value source. That disconnect shows that enterprises know how to launch AI, but not how to run it.
McKinsey's research confirms that workflow redesign is the single strongest differentiator. AI high performers are nearly three times as likely as others to say their organizations have fundamentally redesigned individual workflows, and half of those high performers intend to use AI to transform their businesses source. High performers are also much more likely to define processes for human validation of model outputs and to track KPIs for AI solutions source. The model layer is a commodity; the context layer, the approval layer, and the measurement layer are not.
This is why a code factory that scales software delivery without headcount is not only a staffing play. It is a new operating model that embeds agents into planning, building, reviewing, and shipping so that speed is not a side effect, but the design.
How do high-performing software factories redesign workflows?
Software factories that capture value do four things differently from enterprises stuck in pilot mode.
They start with end-to-end coverage rather than isolated use cases. McKinsey's survey of nearly 300 publicly traded companies found that top AI-driven software performers are six to seven times more likely than peers to scale to four or more use cases across the product development life cycle, from design and coding to testing, deployment, and adoption tracking source. Nearly two-thirds of those leaders report four or more use cases at scale, compared with only 10 percent of bottom performers source.
They create AI-native roles inside the life cycle. More than 90 percent of software teams in the same survey use AI for refactoring, modernization, and testing, saving an average of six hours per week source. Product managers spend less time on feature delivery and more on design, prototyping, quality assurance, and responsible AI practices. Engineers focus on full-stack fluency, spec communication, and architectural trade-offs source. Roles are redefined around orchestration, not replaced by agents.
They align senior leaders with ownership, not sponsorship. McKinsey's AI high performers are three times more likely than peers to strongly agree that senior leaders demonstrate ownership of and commitment to AI initiatives source. Deloitte's study makes the same point from the other direction: tech leaders now need to orchestrate the entire C-suite, not just collaborate with it, because AI accountability is enterprise-wide source.
They treat consolidation as a prerequisite for AI. WEX, a global fintech, spent years with more than 300 Azure DevOps organizations, fragmented pipelines, and no unified visibility into code ownership. After consolidating onto GitHub Enterprise and deploying GitHub Copilot across 1,700 seats, WEX saw approximately 30 percent higher developer productivity, roughly 60 percent ROI on Copilot licenses, and roughly 99 percent reduction in deployment cycle times source. The factory could not run until the factory floor was unified.
What delivery practices separate top performers from pilots?
The top quintile of AI-driven software organizations in McKinsey's research saw 16 to 30 percent improvements in team productivity, customer experience, and time to market, plus 31 to 45 percent improvements in software quality source. The gap between top and bottom performers was 15 percentage points source. Three practices explain most of the difference.
- Upskilling that mirrors real work. Fifty-seven percent of top performers use hands-on workshops and one-on-one coaching, compared with only 20 percent of bottom performers. Top performers integrate AI into code reviews, sprint planning, and testing cycles rather than relying on static documentation source.
- Impact measurement tied to outcomes. Seventy-nine percent of top performers track quality improvements and 57 percent track speed gains, while bottom performers focus mainly on adoption metrics that show little correlation with performance source.
- Incentives aligned with AI-enabled behaviors. Nearly eight in ten top performers link generative AI goals to product manager and developer reviews, compared with just 10 percent of bottom performers for developers and none for product managers source.
Thomson Reuters illustrates the same loop at enterprise scale. After running two focused GitHub Copilot pilots with more than 100 engineers, the company measured 46 percent faster task completion, 39 percent improvement in code quality, and 68 percent of developers reporting a positive experience source. When the program scaled beyond 2,000 seats, pull request durations fell 45 percent and the number of changes per PR rose 44 percent source. The gains came from structured enablement, community champions, and continuous measurement, not from buying licenses alone.
What governance and measurement habits drive ROI?
Governance without speed stalls. Speed without governance breaks. The factories that capture ROI treat both as part of the same operating model.
Deloitte's 2025 tech value survey found that 74 percent of surveyed organizations invested in AI and generative AI in the past year, and more than 95 percent expected moderate to significant value increases in the coming year source. Yet the same survey showed that near-term ROI expectations vary sharply by maturity level. Forty-five percent of respondents expect a near-term ROI, under three years, from basic automation, while 60 percent expect ROI from more advanced multi-agent and process-reimagination levels to take longer source. Setting the right time horizon for each layer prevents the classic mistake of judging a multi-year transformation by quarterly metrics.
Leading organizations also differentiate measurement between generative and agentic AI. Eighty-five percent of Deloitte's AI ROI leaders explicitly use different frameworks or timeframes for generative versus agentic AI source. Generative AI is often assessed on efficiency and productivity, while agentic AI is assessed on cost savings, process redesign, risk management, and longer-term transformation source. The metrics must match the mechanism.
Finally, high performers centralize what should be centralized and distribute what should be distributed. McKinsey found that risk, compliance, and data governance are often fully centralized, while tech talent and solution adoption use a hybrid model source. The goal is not a center of excellence that becomes a bottleneck; it is a repeatable delivery system that can absorb new models and new use cases without rebuilding the governance layer each time.
How quickly should enterprises expect agentic AI ROI?
Enterprises should expect agentic AI ROI to take two to four years for a typical use case, with significant returns from advanced automation often arriving in three to five years. Deloitte found that most executives reported satisfactory ROI on a typical AI use case within two to four years, far longer than the seven to twelve month payback period usually expected for technology investments source. Among agentic AI users, half expect returns within three years while a third anticipate three to five years source. The shortest path to that horizon is to use generative AI for near-term productivity gains while building the data, governance, and change-management foundations that agentic AI needs.
Can a small team build a high-performing AI software factory?
Yes, if the team is structured around outcomes rather than head count. McKinsey's research on AI-driven software development highlights Cursor, the fast-growing AI-native start-up, as an example of a team that operates as an internal lab for AI-driven engineering workflows and scales end-to-end use cases with a lean team source. The lesson is not to copy Cursor's tooling but to copy its operating principle: small pods with clear ownership, agents as collaborators, and continuous feedback loops between builders and reviewers. A small team can outperform a large one when its workflows, roles, and metrics are designed around AI rather than inserted into legacy processes.
When should a company use an external software factory instead of building internally?
A company should use an external software factory when the internal operating model is the bottleneck and the business cannot wait for a two-year redesign. Many enterprises already know why they struggle: they experiment with agents but do not scale them, their leaders are stretched across an expanding tech C-suite, and their budgets are split between running legacy systems and funding AI pilots. An external factory provides the pre-built workflow layer, the validation gates, and the senior oversight that turns agentic AI into shipped production code in days rather than quarters.
This is the core of the AI-native operating model: not more agents, but the right agents embedded in the right steps of the delivery pipeline, measured by business outcomes, and reviewed by experienced engineers. xSquad ships production code in 48 hours using autonomous AI dev squads with senior human oversight because the factory, not the model, is the product.
Agentic AI will not deliver enterprise ROI until enterprises stop asking which model to buy and start asking which workflows to redesign. The organizations that win will be the ones that build, or partner with, software factories designed for that question from day one.
Ready to Scale Your Development Team?
See how xSquad can help you ship production code in 48 hours, not 6 months.
Related Articles
Why Do Nearly Two-Thirds of Enterprises Experiment With AI Agents but Fewer Than 10% Scale Them?
10 min read
Scaling & EfficiencyAI-Native Operating Model: How Agent Factories Actually Deliver 5-6x Speed at Enterprise Scale
11 min read
Scaling & EfficiencyCode Factory: Scale Software Delivery Without Headcount
10 min read