Fluency

Baseline

Last updated

A baseline is the measured record of how a process ran before anything was changed, used as the reference point for whether a deployment improved it. A baseline that stops updating stops being a reference point.

What Baseline actually means

Every claim about improvement is a comparison, and the baseline is the other half of it. Without one, a post-deployment number is a number rather than a result: forty minutes of handling time is neither good nor bad until it is set against what the handling time was before, measured the same way, across the same population.

The distinction that matters is between a measured baseline and an estimated one. Most programs have an estimate, produced in a workshop or from a subject-matter expert recollection, and it is usually wrong in the flattering direction. A reduction measured against an estimate cannot be separated from the estimate being inaccurate, which is why those numbers do not survive scrutiny from finance.

The second failure is staleness. A baseline captured once at pilot stage describes a business that no longer exists within a quarter or two, because teams reorganize, systems get replaced, and volumes shift. Comparing a current number against a frozen baseline attributes ordinary business change to the deployment, in whichever direction happens to be convenient.

This is why a baseline is better understood as infrastructure than as a document. If it is captured continuously from observed execution, it reflects the current business, it covers the variants rather than an idealized path, and the same evidence base serves the next deployment as well as the last one. A snapshot depreciates from the day it is taken.

Examples

Measured against an estimate

A program reports handling time down from ninety minutes to fifty. The ninety came from a workshop estimate. Observed pre-deployment handling time was fifty-five, so the reported improvement is mostly the estimate being wrong.

A baseline that moved on its own

A close process is compared against a baseline captured eighteen months earlier. Two entities have been divested and a system has been replaced since. The comparison credits the deployment for structural change it had nothing to do with.

A baseline covering the variants

Rather than one average, the baseline records the distribution: the median case, the exception path, and the volume of each. Post-deployment, the exception rate can be checked separately from the median, which is where regressions usually hide.

Frequently asked questions

A baseline is the measured record of how a process ran before anything was changed, used as the reference point for whether a deployment improved it. It should cover handling time, exception rate, rework volume, and the distribution across process variants rather than a single average.

Because capturing one requires measuring the process before touching it, and the incentive is to deploy first. What usually exists instead is an estimate from a workshop, which is not a measurement and cannot support a claim about improvement.

Because the business keeps changing. Teams reorganize, systems are replaced, and volumes shift, so a baseline frozen at pilot stage stops describing the current business within a quarter or two. Comparing against a stale baseline credits the deployment for ordinary business change.

A process map describes the shape of a process. A baseline quantifies its performance. A map tells you there are twelve steps; a baseline tells you how long they take, how often they fail, and how frequently the exception path is used.

When it is captured from observed execution rather than assembled from interviews, an initial baseline is available within days and improves as more execution is observed. Reconstructing one retrospectively after a deployment has shipped typically takes two or three quarters longer, because the pre-deployment data no longer exists.

Ready to deploy AI across your enterprise?

Discover how Fluency can help you continuously deploy AI.