Fluency
Back to Blog
Donald La

Donald La

Why AI Transformation Projects Stay Stuck in Pilot

Most enterprise AI pilots follow the same path. The company launches an initiative, the pilot shows promise, the adoption numbers look good, and then the decision to scale never gets made.

Six months become twelve while leadership debates whether to scale it, refine it or shut it down. In IBM's 2025 CEO study, only 25% of AI initiatives had delivered the expected return, and only 16% had scaled across the enterprise.

The technology usually works and people use it. What stalls is the decision. Nobody can say before deployment which workflows will return the investment, yet the business case finance asks for can't be built without an answer.

The use case selection trap

AI transformation typically begins with a hypothesis about a department. For example, if customer service takes too long the plan is to roll out a chatbot for customer service. The observation is accurate, but a department is too broad to be a use case.

Customer service is made up of distinct workflows like first-call resolution, escalation handling, account updates, technical troubleshooting and billing disputes. Each follows a different pattern and delivers different value when automated.

Simple account inquiries might resolve in a minute, so automating them saves almost nothing. Exception requests keep an agent busy for an hour across multiple systems. That hour is worth automating. A plan for "a chatbot for customer service" doesn't say which of those workflows it's for. Without knowing which one, there's no way to say what the plan is worth.

Without a picture of how work flows, leaders choose the most obvious or most painful workflow. Neither one tells you where automation will deliver the most ROI. The most painful workflow may depend on human judgment AI can't supply. The most obvious one may already run efficiently.

RAND's study of why AI projects fail put this kind of mistake first among its root causes. The people commissioning the work often can't say which problem AI is meant to solve. What ends up getting built is measured against the wrong target, or it doesn't fit the way the work happens.

The guessing becomes obvious when finance asks for the expected return. Nobody can answer with confidence, because nobody knows how long the workflow takes today or where the time goes, which is the baseline an AI use case needs.

Decision paralysis at scale

Pilots often succeed by their own metrics. Adoption reaches 75%, users report satisfaction, the technology works as designed, yet leadership still can't say whether it should be scaled.

At that point, the question is whether cycle times fell and quality rose, and whether the freed time went into other work or disappeared.

Adoption numbers don't answer either of those questions. MIT's study of enterprise AI found that over 80% of organizations had explored or piloted tools like ChatGPT and Copilot. Nearly 40% had deployed them. Those tools, the report concluded, "primarily enhance individual productivity, not P&L performance."

Measuring performance needs an accurate picture of work before and after automation. Most enterprises don't have the "before" part of the picture. They measured system activity and tool usage then surveyed sentiment. None of those show how the work ran.

Running several pilots at once makes the decision harder. A company tests AI in customer service, claims processing, underwriting and fraud detection. Some show promise. All of them request funding to scale.

Finance asks which one returns the most, and operations can't compare them because each pilot measured different things and defined success in its own way. The usual response is more analysis. Consultants study the pilots and committees debate the findings. By the time a framework is agreed, AI capabilities have moved on and the original results are stale.

The heavy lift assumption

The paralysis comes from an assumption that new tools need months of change management and training, and processes need redesign before anything can run.

If that's true, choosing the wrong use case can be a waste of time and resources. It seems safer to pilot longer and gather more data before committing. In IBM's survey, 64% of CEOs said the risk of falling behind drives investment in some technologies before they understand the value. That's how a company ends up with pilots it can't evaluate and a scaling decision it can't make.

Deploy tools without a picture of the workflow and the integration feels chaotic. Scale an initiative without a baseline and the return can't be proven. Both outcomes confirm that AI is hard and argue for another round of piloting, which is how the assumption proves itself.

What breaks the pilot trap

Enterprises need baseline visibility before the pilot begins. System logs record what happened inside the systems they cover, and process documentation describes how the work is supposed to run. Consultant interviews add what people remember.

A baseline needs execution data: a record of how tasks flow, where decisions happen, what consumes the time and how much a process varies from one run to the next.

With that record, teams identify the workflows consuming the most time, carrying the most exceptions or varying the most. The business case starts from a measured number and names a target. It gives the pilot a success metric other than adoption, such as research time per case or the wait at each handoff, measured before and after. The scaling decision gets easier too, because the data shows where AI impacted operational performance and where adoption was high while the impact was zero.

In one Fluency deployment in private markets investment operations, cash flow reconciliation went from 60 minutes to 10. That figure only exists because the baseline was recorded before anything changed.

Fluency is foundational to how we think about AI investment. It gave us the clarity to define what to build and how to prove the return.

— Jesse Gill, CTO, Johns Lyng Group

How does Fluency track the ROI of AI automations?

Fluency is a work intelligence platform that shows large enterprises where AI will have the biggest impact by finding the best work to automate first. It observes work at the point of execution, deploys AI agents, and measures results against an observed baseline.

A desktop agent captures work across the applications a team uses, with no integrations and no survey round, and builds the work ontology, a living map of how the enterprise runs. From that map, Opportunities ranks candidate workflows by time saved, frequency, rework rate and estimated cost. The choice of what to pilot rests on measured hours.

The baseline for each candidate exists before any automation gets deployed. Observation continues afterward. You can answer whether the work changed without running a survey.

The scaling decision then compares initiatives on the same measures: throughput, completeness, quality and rework rate, before and after. It shows which pilots delivered value and which only moved a bottleneck somewhere else.

For a team stuck at the pilot stage, the baseline is what's missing, and the usual tools can't provide an accurate one. Process mining only sees work inside a system of record. Consultants and surveys produce a snapshot of what people remember.

From pilot to production

The organizations that get past the pilot stage measure the work rather than the tool. Because they know how a process performs before AI touches it, they can show which initiatives had enough ROI to scale.

Measuring the work also speeds up the decision. A pilot with a baseline can be judged in weeks, early enough to shut a bad one down while the technology it used is still current. Good ones get scaled on evidence.

Put the measurement before the funding decision. Measure the workflow, write down the number the pilot has to move, and then approve it. The pilot ends with a comparison against that number, and so does the decision to scale AI in your organization.

Useful AI starts with understanding the work.

Fluency shows you where AI will return value before you deploy it.

The new way to deploy AI across your enterprise.