Skip to content

Method

How I measure results

Every number on this site invites a fair question: measured how, against what, agreed by whom. This page answers it. It sets out the discipline I apply before work starts, then takes each headline result I publish and states plainly what it claims and what it does not. If a figure here reads smaller than a consultant's brochure, that is deliberate.

By Ashish Kumar Agnihotri·Last reviewed

A result nobody agreed the definition of before the work started is an opinion with a number attached to it.

Each of my four published results is a specific, bounded claim. Throughput is not quality. Billings protected is not revenue earned. Cycle time is not headcount.

In a new engagement the diagnostic exists to fix the baseline, name the owner of each metric, and get client sign-off on the definition before anything changes.

In depth

What you need to know about how operations results are measured.

The order that makes a number defensible

Measurement is not something you do at the end of a project. It is the first hour of work, and it runs in a fixed order. First, agree the definition. What exactly is being counted, in what unit, with what inclusion rule. "Quality improved" means nothing. "The proportion of campaigns passing pre-launch QA on first submission, measured monthly across the full book" means something, because two people reading it will count the same thing. Second, fix the baseline. Pull the last three to six months of actuals before anyone touches the process. If the data does not exist, say so in writing and build the measure first; a baseline reconstructed after the fact is a negotiation, not a number. Third, name an owner. One person, by name, accountable for the number every month. Metrics owned by a department are owned by nobody. Fourth, decide what does not count. This is the step most engagements skip and the one that decides whether the final figure survives scrutiny. Are re-opened tickets counted twice. Do cancelled items leave the denominator. Does a fix applied within the same day count as a miss. Write the exclusions down at the start, because at the end everyone has an incentive to remember them differently. Fifth, get client sign-off on the definition, not just the result. Someone on the client side, usually finance or the process owner, signs that this is the right way to count. Then the number is theirs as much as mine, and it holds up in a board pack.

Leading and lagging indicators, and why the difference matters

A lagging indicator tells you what happened. Cycle time last month, defect rate last quarter, billings written off last year. It is authoritative and it is late. A leading indicator tells you whether the mechanism you installed is actually running. Percentage of items passing through the new checkpoint. Number of exceptions raised and closed. Proportion of approvals completed inside the agreed window. These move within days, not quarters. A result that moves only a lagging number is unproven. Quality scores rise for many reasons: an easier month, a change in client mix, a softer sampling method, a team that has learnt what the auditor looks for. If the lagging measure improved and no leading measure moved, you have a coincidence, not a cause. The improvement will regress the moment attention shifts. So I instrument both. The leading measures tell me whether the control is being used. The lagging measure tells me whether using it produced the outcome. When both move together, and the leading one moves first, you have something that will still be there in twelve months. This is also why the Operations Governance Scorecard holds five to nine metrics rather than thirty: enough room for the outcome and the mechanism behind it, not enough room to hide.

Republic World: what roughly 400 stories a day claims

The claim: newsroom throughput rose to roughly four times its previous level, reaching in the region of 400 stories a day, alongside the addition of two news desks. What it is. A volume measure. Published output per day, counted against the newsroom's own prior output over a comparable period. The two additional desks are part of the change, not a separate result, and any honest reading of the throughput figure has to hold them in view. What it is not. It is not a quality claim. Throughput says how much moved through the system; on its own it says nothing about accuracy, editorial standard or audience response. It is not a revenue or traffic figure. And it is not a productivity-per-person figure, because capacity was added. Why I still publish it. Throughput is the right measure for what the problem was: a newsroom that could not move enough work through its process in a day. Stating it as throughput, and only throughput, is what makes it usable evidence rather than a slogan.

USD 20M+ in makegoods: protected, not earned

The claim: at a global advertising network, a three-step makegoods QA framework protected more than USD 20M in client billings. What it is. An exposure measure. Billings that were at risk of being credited back, and were not, because errors were caught before delivery rather than after invoicing. The unit is money the business kept, and it comes from the same billing records the finance team already ran. What it is not. It is not revenue generated. Not one rupee of new business is claimed here. It is not profit, because it takes no view on the cost of running the QA framework. And it is not a permanent annual figure being restated year on year. The honest caveat. Prevented loss is inherently counterfactual: you are counting things that did not happen. That is why the definition has to be tight and the method has to be the client's own. The defensible version counts errors actually intercepted at a named checkpoint, valued at the billing value of the affected work, against a documented prior credit rate. The indefensible version estimates what might have gone wrong. I only publish the first kind. The same discipline applies to the earlier APAC makegoods work, where the measure was a reduction from USD 200K to zero. That is a floor being reached, not a growth number.

95% to 99%: a quality score, across a stated population

The claim: at an enterprise media operation, quality moved from 95% to 99% across more than 2,000 campaigns and 450 clients. What it is. A quality score against a defined standard, measured over a stated population. The population matters as much as the percentage; a four-point move across 2,000 campaigns is a different fact from a four-point move across twenty. What it is not. It is not a revenue claim, a margin claim or a client-retention claim. It says errors fell. It does not say what that was worth, because attributing revenue to a quality movement requires assumptions I am not willing to make on someone else's behalf. Why the last four points are the hard ones. Going from 80% to 95% is usually a training and checklist problem. Going from 95% to 99% is a systems problem: the remaining failures are rare, scattered and each one has a different cause. You cannot get there by trying harder. You get there by changing what the standard means, where it is checked, and who owns the exception. That is the substance of the work, and it is why the figure is worth stating at all.

Two months to fifteen days: cycle time, nothing more

The claim: at a multi-entity enterprise, the billing approval cycle went from roughly two months to fifteen days, across 75 entities and more than 2,000 employees. What it is. A cycle-time measure. Elapsed calendar time from the start of the approval chain to final approval, measured across the whole entity set rather than a pilot group. Cycle time is one of the cleanest measures in operations because the start and end events are usually already timestamped in a system, which means the baseline is a query rather than an estimate. What it is not. It is not a headcount saving. Nobody's job is being claimed here. It is not a cost figure, and it is not a cash-flow figure, though a shorter approval cycle plainly has implications for both. It also is not a quality claim: approvals moving faster is a different thing from approvals being better. The part worth borrowing. In most multi-entity approval chains, the majority of elapsed time is queue time, not work time. The work takes hours; the waiting takes weeks. The improvement comes from removing handoffs and deciding what does not need approval at all, not from asking approvers to be quicker.

Why client names are withheld, and what you can ask for instead

One client is named on this site. The others are described by role and scale only, and one global advertising network is not named at all. The others are described by role and scale only, and one global advertising network is not named at all. That is a contractual and professional position, not modesty. An advisor who names confidential clients to win the next piece of business will name you to win the one after that. What this costs you as a prospect is real: you cannot ring the client and ask. So here is what you can ask for instead, and what I will do. Ask me to walk you through the mechanism. Not the outcome, the mechanism. How the makegoods framework's three steps sit in the delivery process, what each checkpoint catches, what it deliberately lets through. Anyone can memorise a result; only someone who built the thing can answer the fourth follow-up question. Ask how the baseline was established in each case, and who signed it off. Ask what did not work in the engagement, and what I would do differently. Ask what the number would have looked like under a stricter definition. Ask for a reference conversation. Where a former client is willing, I will arrange one directly, on their terms and with their consent. That is a slower path than a logo wall and a more useful one. And read the two published frameworks. The 3-Step Makegoods QA Framework and the Operations Governance Scorecard are set out in enough detail that you can judge the thinking without taking my word for any number.

How I would measure results in your engagement

The fixed-fee diagnostic exists largely to do this properly, before there is any incentive to make a number look good. During the diagnostic I establish what is currently measured and what is merely reported, which is rarely the same thing. I pull baselines from your own systems, using your own definitions where they exist and writing definitions where they do not. Every proposed metric gets a named owner on your side, agreed in the room. I write down the exclusions. The output includes a short scorecard, five to nine metrics, mixing leading indicators for the mechanisms we install with lagging indicators for the outcomes you actually care about. Your finance or process lead signs off on the definitions. From that point the numbers are produced by your systems and your people, not by me, which is the only arrangement under which they mean anything. Then the retainer runs against that scorecard monthly. If a metric does not move, that is reported as plainly as if it had. I would rather hand you a scorecard with two flat lines and a credible explanation than a deck of green arrows nobody can reproduce. The point of the discipline is that at the end of the engagement, the numbers belong to you and survive my leaving.

Questions

Common questions about how operations results are measured.

It does not argue against them. It bounds them. A number without a stated population, definition and exclusion rule is unusable to a sceptical buyer, and rightly ignored. Stating that the makegoods figure is billings protected rather than revenue earned makes it a smaller claim and a far more credible one. If you are evaluating advisors, the ones whose figures shrink under questioning are the ones to worry about.

That is common in operations-heavy businesses, particularly around quality and cycle time, and it is not a blocker. The first piece of work becomes building the measure itself: defining the metric, finding or creating the data source, and running it for a period before any change is made. It costs a few weeks at the start. The alternative is reconstructing a baseline after the fact, which produces a number nobody in the room believes, including you.

You cannot avoid it entirely, so you design for it. Pair every volume measure with a quality measure and every speed measure with an accuracy measure, so that gaming one degrades the other visibly. Keep the scorecard small enough that each metric gets real attention. Audit the definition periodically, not just the number. And watch for the classic tell: a metric that improves smoothly while the complaints, escalations or rework behind it stay flat.

Leading indicators should move within weeks, because they measure whether the new control is being used. Lagging outcome measures take longer and depend on your natural cycle: a monthly quality score needs at least a quarter to distinguish signal from noise, and a billing cycle-time change is only proven once it has run across a full period for every entity. Anyone promising a lagging number to move in the first month is describing a coincidence.

You do, and by design. The metrics run on your systems, are produced by your people and are signed off by your process owner. My role is to define them properly, install the mechanism behind them and hold the review discipline while it beds in. If the scorecard stops working the month after I leave, the engagement failed regardless of what the numbers said while I was there.

Where a former client is willing, yes, arranged directly and with their consent. What I will not do is name confidential clients on a public page or in a pitch. In the meantime, the more revealing test is to question the mechanism rather than the outcome: ask how a framework's steps work, what each one catches, what it lets through, and how the baseline was established. Depth of answer is harder to fake than a logo.

It would be, if the definition changed mid-measurement, which is exactly why the definition is fixed and signed off before the work starts and is not touched afterwards. If a standard genuinely needs to change during an engagement, the correct treatment is to restate the baseline under the new definition and report both series. Changing the ruler and keeping the old starting point is the most common way operations improvements are overstated.

An agreed metric set of five to nine measures with written definitions and exclusions, a baseline for each drawn from your own systems, a named owner per metric, and identification of the leading indicators that will tell us early whether the changes are working. It is a fixed-fee piece of work, and it deliberately happens before any retainer begins, so the baseline is set at a point where nobody yet has a reason to flatter it.

The next step

A short conversation settles most of this — and a fixed-fee diagnostic settles the rest.

Call