Fluency
←Back to Blog
Donald La

Donald La

How to Identify and Rank the Work Worth Automating

Most organizations have a long list of work they could automate with AI, but aren't sure which of it will pay off. In fact, MIT found that 95% of enterprises get zero return from generative AI tools, despite a $30–40 billion corporate investment.

The lack of ROI from AI tools comes down to using software that doesn't fit how work actually runs day to day, or not understanding the true process. Teams often skip comprehensively measuring how a process works, opting to automate workflows based on workshops and interviews or process mining logs. Both of these methods miss key steps and don't provide the base operational data needed, so adding in AI automation tools just speeds up a broken process.

This guide gives you a six-step method for getting that data and putting it to use. By the end, you'll know how to rank every automation candidate by the hours it will return, match each one to the right kind of automation, and measure the result against a baseline finance can check.

TL;DR

  • Build an inventory of how work happens now from observation. Interviews miss what people forget. System logs miss everything outside connected tools.
  • Score every task on time saved, frequency, rework rate, and estimated cost. High-frequency small tasks yield more recoverable hours than rare large ones.
  • Classify each task as programmatic automation, an AI agent candidate, or a human improvement. Scripts handle fixed rules, AI agents handle variable inputs, and people retain judgment.
  • Rank the backlog by score, merge the duplicates, and keep the order out of the hands of whoever asked loudest.
  • Redesign the workflow before automating it. Remove the steps that exist for no current reason, then automate what's left.
  • Establish a baseline throughput, turnaround, completeness, quality, and rework before deployment, then measure the same five after, without surveys or adoption dashboards.

Why automation programs pick the wrong work

Most teams still find automation candidates through manual discovery: a workshop, a round of interviews, and a list ranked by opinion. Every figure on that list comes from what people remember about their work, not from what they do. That's how programs end up automating the wrong work. It shows up in three places.

Self-reported data distorts the backlog

Workshops run on self-reported data. People estimate how long a task takes and how often they do it, but these estimates are usually too low. In Fluency pilots, teams took 45 to 50 minutes on tasks they'd reported at 20 to 30. People time the main task and forget the steps around it, like moving data between systems, chasing approvals over email, and fixing bad records.

The task list is skewed, too. People bring up long, infrequent tasks because those are the ones they remember. They rarely mention the five-minute task they do repetitively. Every business case built on this list inherits both problems. Teams spend months automating rare, painful work and skip the repetitive work they spend most of their time doing.

Static interviews capture obsolete workflows

Manual discovery records a process as it ran on the day of the workshop. But automations take months to build, so by the time an automation is live, the software has been updated, the team has been reorganized, or a workaround somebody invented has become the official method. The steps before and after the automation also keep changing once it's live. For example, if procurement moves to a new purchasing system and purchase order numbers change format, an invoice-matching automation built for the old format stops finding matches. Someone has to rebuild it. Until they do, it returns less than it promised.

Unobserved edge cases break live deployments

An automation built from what people describe, rather than from the work as observed, misses the unscripted exceptions that are forgotten, causing the automation to break once live. MIT found enterprise AI tools mostly fail from "brittle workflows, lack of contextual learning, and misalignment with day-to-day operations."

When these tools fail, teams without an observed record have no way to diagnose why. They can't tell whether they picked the wrong work, built for an outdated workflow, or automated a step the process already dropped.

Without that visibility, the next automation they try to build repeats the mistake. As the same report notes, organizations keep investing in static tools that can't adapt to how work happens. AI pilots consistently show a massive gap between what leaders believe is happening and what actually occurs at the desk.

Six steps to identify and rank automation opportunities

Following these six steps gives you a clear, repeatable framework to identify where AI will pay off, prove the ROI before writing code, and protect your automation budget from expensive guesses.

Step 1: Build an inventory of how work happens now

Write an inventory list that has every repeated task in scope, which team does it, how often it runs, how long it takes, and which systems it touches. There are three common ways to build this:

What it capturesWhat it misses
Interviews and workshopsThe steps people remember and the problems they still noticeWork the team has stopped noticing, and everything outside the sample of people in the room
System logs (process mining)Every transaction a connected system recorded, with timestampsThe work between systems: spreadsheets, email, browser tools, and manual workarounds
Continuous observation via desktop agentWork as it runs across every application, including the steps between systems, updated as the work changes. Keyed to an artifact carried from the start to end of a process.Work that never happens on a screen, such as a phone call or a conversation at a desk.

Process mining tools only log the status updates recorded in major software like ERPs or CRMs. They completely miss the work between those system check-ins: the emails, spreadsheets, and manual workarounds where up to 70% of execution time lives. An inventory built solely from logs starts from the small fraction of work that system records can see.

In a shared finance operation of 357 people across six teams, observed by Fluency for four weeks, only 15% of the work belonged to a defined process. The other 85% ran outside of one. An inventory built from process documentation or system logs would've only covered that 15%, computing every score and ranking on an incomplete picture. Process documents follow the steps of a workflow linearly, and they fail to identify and document the exceptions and variants of a process that automations and agents will handle. Because a ranking can only order what the inventory contains, traditional discovery misses the high-value automation candidates hiding outside official workflows.

Continuous observation captures this unmanaged work across every browser tab, desktop app, and system handoff, building an inventory that covers the entire range of activity rather than the minority that hits a system log.

The full version of that record, a work ontology, maps every process, actor, and handoff from observed activity. It gives leadership a live map of how work actually gets done, ensuring every dollar goes to the true operational bottlenecks rather than unverified assumptions.


Step 2: Score every task on four dimensions

Score each task in the inventory on time saved, frequency, rework rate, and estimated cost. A task is one repeated piece of work, like matching an invoice to a purchase order. Each invoice that arrives is one case of the task.

Time saved

Measure the hands-on minutes a run of the task takes. Include switching between tools, hunting for files, and picking the work back up after an interruption. People forget those minutes when they estimate, so don't ask them. Take the time from the observed record you built in step 1, or time a few runs yourself.

Leave out the hours a case sits in a queue waiting for someone to pick it up. The team isn't working on the case during the wait, so it adds nothing to labor cost. Record it in its own column instead. An automation picks up each case the moment it arrives. On every step it takes over, the wait disappears and the case reaches the next person sooner.

Frequency

Work out how many times the task runs in a week across the whole organization. Each run is a case. If the task is matching invoices to purchase orders, every invoice is one case. Where a system logs arrivals, like invoices received or tickets opened, use its count. If nothing logs them, have the team tally arrivals for a week.

Frequency is the key variable in deciding what to automate. Five minutes on 400 cases a week is 33 hours. Two hours on one case a month is half an hour a week. People complain about the two-hour task, so it tops a ranking built from interviews. The five-minute task returns more than 60 times the hours.

Rework rate

Next, find the share of cases handled more than once. A case counts as reworked when it's sent back, corrected, re-keyed, or re-approved. Look for cases reappearing in the queue after someone marked them done.

Rework adds hours the frequency count misses, because frequency counts each case once, however many times it's worked. To fold rework into the count, use this formula:

Total runs per week = Unique cases per week × (1 + Rework rate)

A high rework rate also tells you the kind of automation the task needs. Cases usually come back because something in them didn't fit the standard procedure. Tasks whose inputs vary from one case to the next suit an AI agent better than a script.

Estimated cost

Convert the hours to dollars:

Weekly hours = Total runs per week × Hands-on minutes per run ÷ 60

Annual cost = Weekly hours × 52 × Fully loaded hourly rate

Ask finance for the fully loaded hourly rate. Salary alone leaves out benefits, payroll tax, and overhead. Priced on salary, the return comes in low, and the automation looks like it has less ROI.

In the finance operation example, the inventory held about 3,900 hours of work a week. About 2,900 of those hours sat in tasks where the annual cost beat the cost of building an automation, so they went on to be classified and ranked. The other 1,000 sat in tasks where the automation would cost more than it saved. Those stayed as they were.

Now every task in the inventory has a weekly-hours figure and an annual cost. The CFO can compare each one with the cost of building it and decide which automations are worth funding.


Step 3: Classify each task as programmatic automation, AI agent, or human improvement

Every scored task fits into one of these three classes:

  • Programmatic automation: Fixed rules on structured inputs, where the same input always produces the same output. A script or an integration handles it with no language model, like matching an invoice to a purchase order when the purchase order number is printed on the invoice.
  • AI agent: Inputs a person currently has to read and interpret before any rule can apply, because they vary from one run to the next. Invoice triage across changing vendor formats and case routing from a free-text description both belong here.
  • Human improvement: Judgment is the point, so the task stays with a person, and the fix is a shorter path to the decision (a template, a rebalanced queue, or the method the fastest team already uses). Payment approvals above a threshold and disputed claims are typical cases.

The score says whether the task is worth fixing, while the class says what the fix is. Get the class wrong, and you pay either way, for an agent doing work a script could handle or for a script pointed at work that needs a person's judgment.

Use the rework rate from step 2 to split the first two classes. A low rate means the inputs are predictable enough for a script. A high rate means each input has to be read before a rule can apply, which is what an agent is for. For example, a script can match an invoice when the purchase order number is printed in the same field every time. A stack of supplier invoices, each laid out in the supplier's own vocabulary, needs an agent to read each one before the match can run.

Some of the highest-scoring work in the inventory belongs in the third class. A payment approval or a disputed claim costs hours because a person has to weigh it. Give the decision to an agent, and some come out wrong, while a person is still accountable for each one. The decision stays with the person while time is saved from the work around it. Gathering the inputs, chasing the approvals, and re-keying the result belong to the first two classes. Once a script or an agent takes that work, the person spends the time on the decisions, and the team handles more approvals and claims with the same people.


Step 4: Rank by evidence

Merge the duplicate tasks before you sort. An observed inventory records each team's version of the same work, so the same task shows up more than once under different names. Left in, the repeats double-count the hours and overstate the total. In the finance operation example, merging took 241 tasks down to 157. Their annual cost added up to about $9.1 million.

Then sort the inventory by annual cost with the highest first. Without the number, the order gets set by who asked loudest or who is most senior, with no way to check it. With the number, a department head's request names a task already in the inventory, with an annual cost next to it. If the task sits near the top, it goes ahead as asked. If it sits near the bottom, you can show the department head the tasks above it and what each one costs. The order is set by annual cost. Every request, whoever makes it, is checked against the same ranking.


Step 5: Redesign the workflow before you automate it

Before anything is built against a ranked automation candidate, remove the steps that exist for no current reason. Look in the observed version of the workflow for an approval that was added after one bad invoice and never removed, a report that's still produced for a manager who left, or a re-keying step that exists because two systems were never connected.

Inefficiency accumulates when workarounds become permanent, variants multiply across teams, automations degrade, and work outlives its purpose.

For each step in your chosen workflow, ask what would break if it stopped. If nothing would, remove it. Automate the version that's left so it has fewer steps to build and fewer to maintain.


Step 6: Baseline first, then measure the change

Record five baseline measurements prior to automation:

  • Throughput: How many cases the team completes in a week
  • Turnaround: How long a case takes from arrival to completion, including the time it spends waiting
  • Completeness: How many of them finish end to end without being abandoned or handed back
  • Quality: How many are right the first time
  • Rework: How many get touched again after they were marked done

Observe all five for at least one full cycle of the workflow to cover any variations. Write the numbers down before the automation goes live. This is the baseline. Without it, nobody can say afterward whether the automation changed anything, and the decision to scale stalls.

After deployment, measure the same five factors on the same workflow, from the same kind of observation. The difference between the baseline and the new record is the return. In one Fluency deployment in private markets investment operations, cash flow reconciliation went from 60 minutes to 10. The 50 minutes saved on every reconciliation is the return, and it's visible because the baseline was recorded before anything changed.

Compare throughput and rework with the baseline to find the hours saved. If the team completes more cases a week with the same people and touches fewer of them a second time, the difference is time the team no longer spends on the task. Compare turnaround to see how much sooner each case finishes. Most of that speed comes from removing the queue time you recorded separately in step 2. Queue time was never counted in the hours, so turnaround is the only place you'll see it. Then check quality and completeness to make sure the faster work is still right the first time and finished end to end.

To get the deployment's ROI, price the hours saved at the fully loaded rate from step 2 for the annual saving, then set that against what the automation costs to build and run. Every figure in that calculation comes from observed work. When you ask finance to fund the next automation, you can show them what the last one returned.

Run the method continuously with Fluency

Fluency is a work intelligence platform that shows large enterprises where AI will have the biggest impact by finding the best work to automate first. It observes work at the point of execution, deploys AI agents, and measures results against an observed baseline.

Run by hand, these six steps produce one inventory and one ranking, both out of date as soon as the work changes. Fluency runs all six continuously on a desktop agent, deploys in under an hour with no direct system access, and captures work across systems, teams, and apps. The inventory from step 1 comes out of that capture and stays current.

Opportunities scores every task it finds on time saved, frequency, rework rate, and estimated cost, then classifies each one as programmatic automation, an AI agent candidate, or a human improvement, and ranks the backlog by ROI. Automations builds AI agents from the observed workflow, maintains them as the work changes, and measures each deployment against the baseline it observed before launch. The finance operation example is one such deployment. Four weeks of observation produced its inventory, its scores, and its ranked backlog.

Fluency observes work, not workers. It doesn't capture keystrokes or screen recordings, nothing it collects ties back to an individual, and deployments run under SOC 2 Type I and II controls.

“Fluency is foundational to how we think about AI investment. It gave us the clarity to define what to build and how to prove the return.”

— Jesse Gill, CTO, Johns Lyng Group

Find the work worth automating in your operation

An automation program using these steps funds the tasks with the highest return. It measures each deployment against a baseline, so finance sees what the build returns before the next one is funded. Because the inventory is rebuilt from the work as it runs, it still holds when a process changes.

Every score, class, and rank in the method is computed from the inventory. Building one is the step most programs don't do well. Fluency builds the inventory from observation across systems, teams, and apps, including the work between systems, and keeps it current as the work changes.

Request a Fluency demo to see how the inventory, the scores, and the ranking come out of observed work, and what the first deployed agents look like.

FAQs

How do you identify automation opportunities?

Observe how work happens across the applications a team uses, list every repeated task with its time per run, frequency, and rework rate, and rank the list by the hours and cost each task would return. Interviews and workshops turn up the work people remember, and the system logs the work inside connected systems. Observation covers both, plus the work between systems.

Which tasks should be automated first?

The first candidates are the tasks that take the most hours a week (time per run multiplied by frequency), carry a high rework rate, and run on fixed rules because they return the most hours, and a script can automate them. Variable-input work goes to an AI agent, and anything where judgment is the point stays with people and gets a shorter path instead.

How do you measure the ROI of an automation before building it?

The ceiling on what an automation can return is the observed time per run, multiplied by runs per week, at the cost of the people doing the work. Cost depends on the class, a script for fixed rules, or an AI agent for variable inputs. After deployment, measure throughput, turnaround, completeness, quality, and rework against the same baseline to confirm the number.

Can you find automation opportunities without task mining?

Yes, continuous work observation finds them without a task mining project. It captures execution across every application as it happens and keeps the inventory current as processes change. Task mining records selected users during a defined project window, so its findings describe the scope and the weeks it was pointed at.

Useful AI starts with understanding the work.

Fluency shows you where AI will return value before you deploy it.

The new way to deploy AI across your enterprise.