Why most hospital benchmarking produces nothing
Hospital benchmarking usually arrives as a slide showing your average length of stay against an industry figure, and it usually changes nothing. The reason is structural: the number being compared has been produced by two different definitions, over two different patient populations, in two different care contexts, and everyone in the room knows it. The conversation therefore becomes a debate about the comparison rather than about the performance.
The failure is not that benchmarking is useless. It is that a comparison only carries information when the things being compared are alike in every respect except the one under examination. Get that right and a benchmark is one of the most powerful tools available, because it converts an abstract question — are we good at this? — into a concrete one: why does the unit next door discharge two hours earlier than we do?
Everything that follows is about earning the right to make the comparison. Choose the cohort deliberately, adjust for the differences you can measure, be honest about the ones you cannot, and design the exercise so that the natural response to a gap is investigation rather than justification.

Choosing a comparable cohort
A comparable peer shares the characteristics that mechanically drive the metric. For length of stay, that means similar specialty mix, similar proportion of surgical to medical admissions, similar critical care capability, and similar payer mix — a hospital with heavy PM-JAY or ESIC volume faces different discharge dynamics than one that is largely cash and private insurance. For OPD throughput, it means similar appointment model, similar walk-in proportion, and similar consultant availability patterns.
Bed count is the most commonly used and least informative comparator. A 200-bed hospital doing high-acuity tertiary work and a 200-bed hospital doing routine secondary care have almost nothing in common operationally. If you must use a size proxy, occupied bed-days or annual admissions is more meaningful than sanctioned beds, which in many Indian hospitals reflects a licence rather than an operating reality.
Be explicit about what you cannot match. Teaching status, referral patterns, ownership model, and the local competitive landscape all shape performance and are rarely comparable. Naming these openly at the start of the discussion prevents them being deployed later as a blanket excuse, which is their usual function.
Characteristics that should match before you compare
- Specialty and case mix, not just bed count
- Surgical to medical admission ratio
- Critical care and emergency capability
- Payer mix, including scheme and corporate volumes
- Referral role: primary catchment versus tertiary referral centre
Case-mix and severity adjustment in plain terms
Case-mix adjustment answers a single question: how would our numbers look if we treated the same kinds of patients as the comparator? The simplest honest method is indirect standardisation. Split both populations into strata — by diagnosis group, age band, and severity marker — compute the comparator's rate within each stratum, apply those rates to your patient counts, and you have the number you would expect if you performed exactly like them. Compare your observed number to that expected number.
The resulting observed-to-expected ratio is far more interpretable than a raw difference. A ratio above one means you are longer, costlier, or higher on whatever the metric is than your case mix predicts; below one means the opposite. It also degrades gracefully — if a stratum has very few patients, the ratio is unstable and you can say so, rather than presenting a difference that is really noise.
Severity markers are the hard part in Indian hospitals, because a well-populated coded severity field is not universal. Usable proxies include ICU admission during the stay, ventilator hours, number of active comorbidities on the problem list, emergency versus elective admission, and admission-time vitals or early warning score. None is perfect; all are better than comparing crude averages and pretending the populations are identical.

What unadjusted comparison does to behaviour
An unadjusted league table does not merely mislead; it teaches people the wrong lesson. If a unit is measured on crude mortality without severity adjustment, the rational response is to avoid the sickest patients — refuse the difficult transfer, discharge before the risky period, refer out the complex case. That is not a hypothetical failure mode; it is the predictable consequence of the incentive you created.
Unadjusted comparison also destroys the credibility of the whole measurement programme. Clinicians can identify a mismatched comparison faster than an analytics team can defend one, and once a set of numbers has been shown to be unfair, every subsequent number from the same source is treated as suspect. The reputational cost of one badly constructed benchmark is high and durable.
The rule that protects you is simple: if you cannot adjust, do not rank. Present the distribution, present the case-mix differences alongside the metric, and let the discussion be about mechanisms. A chart showing four units with their patient profiles beside their numbers produces a better conversation than a sorted list with a red cell at the bottom.
Internal branch-to-branch benchmarking
The most useful benchmarking a multi-branch group can do is against itself, because you control the definitions on both sides. The same system, the same metric logic, the same coding conventions, the same data collection responsibilities — the comparison is clean in a way that no external benchmark can be. If branch A completes discharge in three hours and branch B in seven, the difference is real and the reason is findable.
Findable is the key word. Internal comparison is valuable precisely because you can go and look: sit in branch A's discharge process, read the actual sequence of steps, and identify what they do differently. Perhaps the pharmacy delivers take-home medicines to the ward rather than making families queue; perhaps the billing clearance runs in parallel with the summary rather than after it. That mechanism is transferable in a way that an external number never is.
Multi-branch analytics does require that definitions be genuinely centralised rather than nominally shared. Where each branch computes its own figures in its own spreadsheet, comparison is comparing accounting habits. A platform such as HealUDoc that computes branch metrics from one shared model removes the definitional argument, leaving the operational one — which is the argument worth having.

Making internal comparison actionable
- One metric definition computed centrally, not per branch
- Same reporting calendar and cut-off time across sites
- Case-mix profile shown alongside every branch figure
- A named owner at each branch who can explain their number
- A route to observe the leading branch's actual process
Acting on a gap instead of explaining it away
When a gap appears, the default organisational response is to find a reason it does not count. Sometimes that reason is legitimate and the adjustment was inadequate. More often it is a reflex, and the tell is that the explanation arrives before anyone has examined the underlying cases. A useful discipline is to require that any dismissal of a gap be supported by case-level review, not by a general statement about the patient population.
Structure the follow-up as an investigation with a deadline and an owner. Pull twenty consecutive cases from the lagging unit, walk the process, identify where time or cost accumulates, and compare that walk against the same twenty-case walk at the comparator. This converts a benchmark into a clinical audit, which is the mechanism that actually changes practice — the benchmark's job was only to point at where to look.
Close the loop by re-measuring after the change, using the same adjusted method. A benchmarking programme that generates gaps but never demonstrates closure trains the organisation to treat benchmarks as an annual ritual. One that shows a unit moving from an observed-to-expected ratio of 1.3 to 1.05 after a specific process change has proved its own worth, and the next gap gets taken seriously.
“The number never changed anyone's mind. Sending two of our consultants to spend a morning watching the other branch's discharge round changed it in a week.”
Building benchmarking into the operating calendar
Benchmarking works when it is periodic and expected rather than commissioned in response to a problem. A quarterly cycle suits most metrics: one month to compile and adjust, one meeting to review gaps and assign investigations, and the remainder of the quarter to act before the next cycle. Anything faster and the underlying processes have not had time to move; anything slower and the findings go stale.
Keep the metric set small. Four to six adjusted metrics that the leadership genuinely acts on beat thirty that get skimmed, and a short list forces the prioritisation conversation to happen at the design stage rather than in the meeting. Retire a metric once it has stabilised across all units and add a new one in its place.
Finally, publish the method alongside the results, every time. The cohort definition, the adjustment approach, the exclusions, and the known limitations should travel with the numbers. It costs a page and it removes the single most common way benchmarking discussions get derailed — the suspicion, usually unspoken, that the comparison was constructed to produce a particular answer.



