Agentic AI is often evaluated through the promise of faster processing, lower labor costs, and greater automation. The harder calculation begins after a pilot moves into production, when integration work, model and infrastructure spend, exception handling, governance, maintenance, and the cost of a wrong action enter the equation.
In this TNGlobal Q&A, Atul Arya, Founder and Chief Executive Officer of Blackstraw.ai, discusses the evidence management teams should require before scaling an agentic system, the operational costs that belong in an AI business case, and why human approval should remain in place for consequential decisions.

What evidence should management require before scaling an agentic AI system?
A Gartner CDAO boardroom I moderated highlighted a stark reality: only about 5 percent of generative AI pilots reach production with real impact, a figure reinforced by MIT research citing a 95 percent failure rate. That aligns with what we observe across the ecosystem.
A primary reason is that many boards do not define what working actually means. Before taking an agentic system to a board for a scaling decision, there are three things we ensure are settled before deployment:
- A real baseline: Clear visibility into what the process costs today in dollars, hours, and error rate, measured directly so it can be compared after deployment.
- Business metrics rather than model metrics: Accuracy and latency matter to engineering, while cycle time, cost per transaction, exception rate, and customer impact are the measures a board needs to evaluate.
- Human-in-the-loop triggers: A clearly defined boundary for when a human must step in.
If a pilot cannot show real return on investment, it is a demonstration rather than a program, and boards should treat it that way.
How should enterprises calculate the ROI of an agentic AI deployment?
A common mistake is treating agentic AI like a one-time software build instead of an operating system with ongoing running costs. Labor savings and faster processing are the easier half of the calculation, and the part that looks good in a presentation.
The harder half is accounting for the total cost stack:
- Upfront and system costs: Integration with legacy infrastructure and baseline architecture.
- Ongoing operational spend: Continuous model inference, monitoring, maintenance, and periodic retraining.
- Human-in-the-loop realities: Exception handling, manual rework, and active reviewer oversight.
True ROI requires netting total business gains against the full operational cost stack over a multi-year horizon, not only the initial licence tag or one strong quarter after launch.
In our FinOps work on agentic systems, we consistently find that 20 percent to 35 percent of AI operating spend can be recovered by fixing untracked prompt usage and overlapping agent-call costs. Many enterprises do not model these costs until systems are live.
What did the claimed $16 million ROI from a credentialing automation project represent?
The client operated on legacy, commercially licensed analytics software that was expensive, rigid, and slow to adapt. We executed a phased, zero-downtime migration to a cloud-native, open-source stack.
The financial baseline and attribution were separated as follows:
- Infrastructure modernization: The baseline was established using existing annual licensing and infrastructure expenditure. Eliminating commercial licences accounted for the majority of the reported $15 million to $16 million in annual savings. That portion is an architectural-modernisation story rather than an AI story.
- AI and automation: AI and automation contributed by improving analytics and reporting cycle times by 40 percent to 50 percent through faster pipelines, automated deployment, and new self-service capabilities.
Infrastructure modernization generated the primary cost reduction, while the AI layer compounded that value by making the platform faster to use as well as cheaper to run.
Why do technically successful pilots fail to create operational value in production?
A pilot succeeds because it operates on curated data, a narrow scope, and close team supervision.
Operational bottlenecks typically emerge in five areas:
- Data readiness: Pilots run on clean, ideal sample data. Production requires handling the messy, fragmented reality of enterprise data at scale, a gap many organisations underestimate.
- Process ownership: A pilot has an executive sponsor, while production needs a permanent business owner. Without ongoing ownership, a system can drift until performance degrades without notice.
- Legacy integration and edge cases: Transactions that deviate from the standard path require the highest degree of human judgment, yet pilots rarely test these exceptions or stress-test legacy infrastructure.
- User adoption: A technically sound system that does not integrate smoothly into existing workflows will be bypassed or routed around.
Technical success in a pilot proves feasibility. Delivering operational value in production requires solving for messy data, missing governance, and workflow friction that pilots deliberately isolate.
How does greater autonomy change the economics and risk profile?
A copilot’s worst failure is a bad recommendation. An agentic system’s worst failure is an action already taken, such as a payment issued, a record altered, or an incorrect email sent to a customer.
That shift affects both economics and governance:
- Economic impact: Autonomy requires organisations to price in the cost of rework, remediation, and liability rather than evaluate only standard model-execution costs.
- Governance architecture: Risk management must move upstream into predefined permissions rather than rely only on downstream auditing.
Autonomy should scale with reversibility. Any action that is legally binding, financially consequential, customer-facing, or subject to regulatory oversight should retain a human in the loop, even if full automation is technically feasible.
Full autonomy should not be granted because a demonstration went well. It must be earned over time and validated transaction by transaction.
What should enterprises budget for after deployment?
Enterprises generally budget well for the build phase but plan poorly for ongoing operations. Post-deployment, we ask clients to explicitly plan for five recurring operational categories:
- Detecting model and data drift before degradation appears in business outcomes.
- Refining prompts, logic, and models as source data and business rules evolve.
- Tracking performance against the original business metrics.
- Securing a broader attack surface, because an autonomous system capable of taking action carries higher exposure than one that only recommends.
- Defining explicit ownership for exception handling, system maintenance, and incident response.
A realistic baseline for long-term viability is 15 percent to 25 percent of the original build cost annually. If that recurring line item is absent from the initial business case, the case was incomplete from the start.
Which execution challenges are particularly important in Asia Pacific?
Having established our first delivery center outside the United States in Chennai, I have watched this ecosystem closely for years. This is based on direct operational experience rather than theory.
- Regulations vary sharply across jurisdictions. India’s Digital Personal Data Protection Act, Singapore’s Personal Data Protection Act, and China’s Personal Information Protection Law do not align. Data-residency architecture, which may be a secondary consideration in North America, needs to be established upfront.
- Legacy infrastructure across Asia Pacific is often more fragmented, although not necessarily less modern. Enterprise banking and telecommunications stacks have been built over decades using localised systems that were never designed to interoperate.
- Language and dialect variation changes the definition of production-ready data for customer-facing implementation.
- The region has deep core engineering talent, but disciplines such as MLOps and production AI governance remain scarce across APAC.
These factors are not reasons to slow execution. They require organizations to design for regional realities from the outset.
Which workflows are most likely to justify autonomous or multi-agent systems?
The workflows that justify autonomous or multi-agent systems are high-volume, rule-governed, and carry a high cost of delay. Repeated transactions combined with material financial consequences are where multi-agent architectures can deliver rapid ROI.
Enterprises tend to overestimate agentic AI in low-volume, judgment-heavy decisions such as senior negotiations, complex legal interpretation, and strategic resource allocation. In these domains, the effort rarely yields a positive return, and one confident but erroneous autonomous action can cost far more than the labour it saves.
Less than 10 percent of agentic AI and Model Context Protocol pilots make it to production, and the barrier is rarely the core technology. Projects stall because use cases are too small to justify the initial ROI conversation, organisations fail to account for the operational change management needed for adoption, or technical teams remain focused on framework debates rather than production deployment.
Editor’s note: This Q&A has been lightly edited for clarity and TNGlobal house style. The substance of the interviewee’s responses has been preserved.
Share your perspective: TNGlobal welcomes contributed insights and expert commentary from across Asia’s technology and innovation ecosystem. Submit a contribution for editorial consideration, or explore more conversations in our TNGlobal INSIDER and TNGlobal Q&A and Interviews archive.
Nebius’ Dr. Ilya Burkov on making safe healthcare AI the easy option [Q&A]

