When a CFO asks, “What did we actually get for this investment?”, a dashboard filled with adoption rates, interaction volumes and positive sentiment is not enough.
These indicators can help explain whether people are using an AI agent. They do not prove that the agent has reduced costs, improved service, accelerated revenue or strengthened operational performance.
That distinction is becoming increasingly important. Gartner predicts that more than 40% of agentic AI projects will be canceled by the end of 2027 because of escalating costs, unclear business value or inadequate risk controls.¹
For organizations moving from experimentation to production, the challenge is no longer simply proving that an agent works. It is demonstrating that the agent creates measurable, repeatable and governable business value.
Why traditional ROI frameworks fall short
Traditional automation generally follows predefined rules. Its value can often be measured through transaction volumes, labour substitution or reductions in handling time.
Agentic AI is different.
An AI agent may interpret a goal, develop a plan, use several tools, retrieve information, make decisions and complete multiple steps across an end-to-end workflow. Human involvement may vary according to the task, the confidence of the system and the level of risk.
This makes isolated measures such as time saved on one step unreliable. A faster activity does not create value when it introduces rework, moves a bottleneck downstream or requires extensive human review.
Organizations must also distinguish genuine agents from conventional assistants and automation tools. Gartner describes “agent washing” as the practice of rebranding chatbots, robotic process automation and AI assistants as agents without substantial agentic capabilities.¹
The market remains early. Gartner reports that 17% of organizations have deployed AI agents, while more than 60% expect to do so within the next two years.² That gap between current deployment and future ambition makes disciplined measurement essential.
Establish the baseline before deployment
A credible business case begins before an agent goes live.
For every proposed workflow, document the current state, including:
- End-to-end completion time
- Cost per completed outcome
- Work volumes
- Error and rework rates
- Escalation rates
- Service-level performance
- Customer or employee outcomes
- Revenue influenced by the workflow
- Compliance and operational incidents
The comparison should use the same scope, definitions and measurement period before and after deployment. Without a reliable baseline, improvements can easily be attributed to an agent when they were caused by changes in demand, staffing, seasonality or another technology initiative.
Measure four dimensions of value
An effective measurement model should examine productivity, financial performance, quality and risk, and workforce experience.
Productivity
Productivity should be measured at the workflow level rather than at the level of an isolated AI-assisted task.
Useful measures include:
- End-to-end cycle time
- Tasks completed per employee
- Percentage of work completed without intervention
- Human escalation rate
- Time spent reviewing agent outputs
- Rework created by incorrect or incomplete outputs
- Capacity redirected to higher-value activities
Time saved is only valuable when the organization can explain what happens to the recovered capacity. It may allow employees to serve more customers, reduce backlogs, complete analysis sooner or focus on higher-value work.
McKinsey says its experience indicates that initial uses of agents to support employees and automate tasks can produce company-level annual productivity improvements of 3% to 5%.³ These gains should be treated as directional evidence rather than a guaranteed benchmark for every deployment.
Financial performance
Financial measurement should capture the full cost of producing a successful outcome.
A practical metric is cost per completed task:
Cost per completed task = total operating cost divided by successfully completed tasks
Relevant costs may include:
- Model and inference charges
- Software licensing
- Integration and infrastructure
- Data preparation
- Monitoring and evaluation
- Security and governance
- Human review
- Exception handling
- Maintenance and improvement
- Change management and training
A useful expansion measure is an agent value multiple:
Agent value multiple = financial value created divided by total agent cost
Financial value may include verified cost savings, additional margin, accelerated revenue or avoided losses. Each component should have an agreed owner, calculation method and evidence source.
Quality and risk
An agent that completes work quickly but produces unreliable outcomes can destroy value.
Quality and risk measures should include:
- Accuracy against an approved standard
- Error severity
- First-time completion rate
- Policy compliance
- Unsupported or fabricated outputs
- Security incidents
- Unauthorised actions
- Reversals and remediation costs
- Performance drift over time
Human intervention is not automatically a sign of failure. For sensitive or consequential tasks, review may be a deliberate control. The objective is to define when intervention is required and determine whether the agent behaves consistently within that boundary.
Workforce experience
Employee sentiment should not be treated as proof of ROI, but it remains an important operational indicator.
PwC surveyed 308 senior executives in the United States in May 2025. Seventy-nine per cent said agents were being adopted in their companies. Among organizations adopting agents, 66% reported measurable productivity value.⁴
Organizations should measure whether agents:
- Reduce repetitive work
- Improve access to information
- Increase employee capacity
- Create new review or administration burdens
- Affect role clarity
- Improve or weaken confidence in decisions
- Change training and skill requirements
These measures help leaders understand whether the technology is improving the operating model or simply moving work from one part of the organization to another.
Separate activity from business value
Many commonly reported metrics describe activity rather than outcomes.
Examples include:
- Number of prompts
- Number of agent interactions
- Registered users
- Login frequency
- Percentage of employees with access
- Employee satisfaction scores
These figures may support adoption analysis, but they cannot independently demonstrate financial or operational value.
Metrics that are more likely to withstand executive scrutiny include:
- Cost per successful outcome before and after deployment
- End-to-end cycle-time reduction
- Error and rework rates
- Revenue accelerated or influenced
- Backlog reduction
- Service-level improvement
- Risk incidents and avoided losses
- Percentage of recovered capacity redirected to priority work
A balanced scorecard should show activity, outcomes and risk together. High adoption combined with weak results is not success. Lower adoption within a well-selected workflow may produce significantly greater value.
Account for people and process change
Technology alone does not determine the success of an AI transformation.
BCG’s 10-20-70 principle suggests that approximately 10% of transformation effort should focus on algorithms, 20% on technology and data, and 70% on people and processes.⁵
This does not mean every organization should divide its budget according to those percentages. It highlights that operating-model redesign, capability development, governance and change management often determine whether technical performance becomes business value.
Measurement should therefore include indicators such as:
- Training completion and demonstrated proficiency
- Adoption within the intended workflow
- Process compliance
- Decision rights
- Accountability for outcomes
- Employee confidence
- Speed of issue resolution
- Business-owner participation
When measurement focuses only on technical performance, it overlooks much of the work required to scale safely and successfully.
Build the expansion case through three stages
Prove
Select a high-volume workflow with a clear owner, measurable baseline and defined outcome.
The first deployment should establish whether the agent can improve the entire workflow without creating unacceptable quality, cost or risk.
Scale
Apply the proven measurement and governance model to adjacent workflows.
Scaling should not mean copying the technology without adapting the controls. Each workflow may have different data requirements, risk thresholds, review rules and definitions of success.
Transform
Redesign the process around the combined strengths of people and agents.
The greatest value may not come from automating the existing process. It may come from removing unnecessary steps, changing decision points, connecting previously separate workflows or creating a new service model.
PwC’s 2026 AI performance research found that 20% of 1,217 surveyed companies captured 74% of the AI-driven returns.⁶ This concentration suggests that value depends on focused execution rather than the number of disconnected pilots an organization launches.
Create evidence that decision-makers can trust
A boardroom-ready agentic AI business case should clearly show:
- The business problem
- The pre-deployment baseline
- The target outcome
- The agent’s defined scope
- Total implementation and operating costs
- Financial and operational benefits
- Quality and risk performance
- Human oversight requirements
- The measurement period
- The accountable business owner
- The conditions required for expansion
It should also distinguish measured results from forecasts. Forecasts may support an investment decision, but they should never be presented as realised value.
Agentic AI has the potential to create meaningful productivity improvements. McKinsey says early uses can generate annual company-level productivity gains of 3% to 5%.³ The opportunity is significant, but it depends on selecting the right workflows, redesigning the work and measuring outcomes consistently.
The organizations that move successfully from pilot to production will not be those with the most agent demonstrations. They will be those that can show, with credible evidence, where value was created, what it cost and how risk was controlled.
Insentra helps organisations establish the measurement, governance and operating foundations required to turn agentic AI pilots into credible investment cases.
Explore our thinking at AI Momentum.
Sources






