Skip to main content
AI & Automation10 min read

Evaluating AI Screening Tools as a Buyer, Not an Enthusiast

Screening AI is sold on accuracy figures from populations that may not resemble yours. What validation evidence to ask for, how workflow placement determines whether it helps, and the per-study economics that decide it.

Dr. Ananya Chatterjee

Clinical AI Governance Advisor

#ai screening tools hospital#clinical ai validation#ai vendor evaluation healthcare#screening workflow design#ai economics per study
Evaluating AI Screening Tools as a Buyer, Not an Enthusiast

The accuracy figure is the least useful number offered

Screening AI is marketed with a headline accuracy or sensitivity figure, and that figure on its own is close to uninterpretable. It depends entirely on the population it was measured in, the reference standard it was measured against, the prevalence in that population, and the operating threshold chosen. The same model produces very different numbers across those choices without anything about the model changing.

Prevalence matters more than buyers expect. A tool with strong performance in a high-prevalence population will produce a much larger proportion of false positives when deployed in a screening context with low prevalence, because most positives in a low-prevalence setting are false regardless of how good the model is. That is arithmetic rather than a defect, and it determines the workload the tool creates.

So the questions to ask are about the study rather than the number: in whom, against what reference, at what prevalence, at what threshold, and how does performance change if the threshold moves. A vendor who can answer those readily is a different proposition from one who repeats the headline figure.

A headline accuracy figure decomposed into population, reference standard, prevalence and operating threshold
A headline accuracy figure decomposed into population, reference standard, prevalence and operating threshold

Validation on a population that resembles yours

The most important question for an Indian hospital is whether the tool has been validated on a population resembling its own, on equipment resembling its own. Models trained and validated elsewhere can perform differently on different demographics, different disease spectra, different equipment and different acquisition protocols, and the degradation is not predictable from the original figures.

Ask for validation evidence from Indian populations specifically, and read what is offered rather than accepting its existence. A study on a few hundred cases from one tertiary centre is not strong evidence for deployment in a different setting with different case mix. Absence of local validation is not necessarily disqualifying, but it makes local evaluation before full deployment essential rather than optional.

Ask also about equipment. Imaging AI in particular can be sensitive to the acquisition device, protocol and processing, and a tool validated on modern equipment may behave differently on an older machine that remains entirely adequate clinically. This is a specific question worth asking about your actual equipment rather than in general.

Validation evidence worth requesting and reading

  • Population studied, including demographics and care setting
  • Prevalence in the validation population versus your screening context
  • Reference standard used to establish ground truth
  • Equipment and acquisition protocols the tool was validated on
  • Whether any Indian validation exists, and its size and setting

Regulatory status, verified rather than asserted

Establish the tool's regulatory position in India rather than accepting a general claim of approval. Software intended for a medical purpose falls within the medical device framework, and what matters is the specific product, its stated intended use, and its registration or licence status here — not that it holds a clearance in another jurisdiction, which does not transfer.

Ask for the specifics in writing: the registration or licence particulars, the classification, and the intended use exactly as stated in the registration. Then check that the intended use as registered matches the use you propose. A tool registered as an assistive aid to a reporting clinician and deployed as an autonomous screening filter is being used outside its stated purpose, which matters both regulatorily and clinically.

Record what you established and when in your governance file. Regulatory status changes, product versions change, and a decision taken on the basis of a check nobody documented cannot be revisited sensibly. This is the same discipline as any other clinical procurement and it is applied far less consistently to software.

Workflow placement determines whether it helps at all

The same tool produces entirely different value depending on where it sits. As a triage aid reordering a worklist so that likely-abnormal studies are read sooner, it can shorten time to diagnosis without changing any clinical decision, which is a low-risk and genuinely useful placement. As a second reader flagging studies a clinician has already reported, it catches misses at the cost of additional review. As an autonomous filter deciding which studies a clinician sees at all, it carries a fundamentally different risk profile.

Decide the placement explicitly and match it to the evidence and the regulatory position, because vendors will often demonstrate the most impressive placement while the defensible one is more modest. Worklist prioritisation is where most hospitals should start: the benefit is real, the failure mode is a study read in the usual order rather than sooner, and nothing depends on the model being right.

Then consider what the tool does to the clinician reading. A flag on a study influences the person reading it, in both directions — raising suspicion where the tool flagged, and potentially lowering it where the tool did not. That effect is real and is a reason to think carefully about whether outputs are shown before or after the clinician forms their own impression.

The same model placed as worklist triage, as second reader, or as an autonomous filter, with very different risk
The same model placed as worklist triage, as second reader, or as an autonomous filter, with very different risk

We deployed it as a prioritisation tool rather than a reading aid. The radiologists stopped objecting immediately, because nothing about their reading changed. It just reordered the queue.

Head of imaging at a mid-size hospital

The economics, per study rather than per licence

Reduce the commercial proposition to a cost per study and compare it against what the tool changes. That comparison has to include the downstream cost of positives: a screening tool that flags cases generates follow-up investigation, and in a low-prevalence setting most of those will be false positives, each carrying a real cost in investigation, clinician time and patient anxiety.

Model that explicitly before deployment. Expected volume, expected positive rate at your prevalence and chosen threshold, and the cost and capacity implication of investigating them. Hospitals routinely deploy screening tools without this and are then surprised by the workload created, which is entirely predictable arithmetic.

Then be clear about what the benefit actually is and whether it is monetisable. Earlier detection, reduced misses and faster turnaround are real benefits and may not produce revenue. If the case rests on quality rather than return, say so plainly to the board rather than constructing a financial case that will not survive review, which is how these projects lose credibility.

What the economic model must include

  • Cost per study at realistic volume, not at list price
  • Expected positive rate at your prevalence and threshold
  • Downstream investigation cost and capacity for those positives
  • Integration and ongoing operating cost, including the owner's time
  • Whether the benefit is monetisable or is a quality argument

Piloting on your own data before committing

Run any screening tool retrospectively on your own historical cases before deploying it prospectively, where the licence permits. It requires no clinical risk, uses cases whose outcomes you already know, and answers the only question that matters: how does this perform on our population, our equipment and our case mix.

Include the cases the tool will find hard. A retrospective evaluation on a clean consecutive series tells you about ordinary performance; deliberately including known difficult cases, poor-quality studies and the conditions that are commonly confused tells you where it fails. The failure pattern matters more than the aggregate, because it determines what you must warn clinicians about.

Then evaluate prospectively in a limited placement with monitoring before extending. Retrospective performance is a necessary condition and not a sufficient one, since prospective use introduces workflow effects, real-time data quality and the influence on clinician reading that a retrospective run cannot capture.

What has to be in place before it goes live

A screening tool in production needs an owner, a monitoring arrangement, a defined response when it fails, and a route to withdraw it. Monitoring should track the flag rate against expectation, since a drift in flag rate is usually the first observable sign that something has changed in the model, the equipment or the population.

Establish before deployment what happens when the tool disagrees with a clinician, who reviews such cases, and how a suspected error in the tool is reported and escalated. Clinicians need a route to report that it is producing wrong output that does not require them to raise a general complaint, and that route needs to reach someone who can act.

Record the whole deployment decision in your AI governance register: what was evaluated, the validation evidence, the regulatory position established, the placement chosen and why, the economic model, the owner and the monitoring arrangement. That record is what makes the next evaluation faster and what allows anyone to reconstruct, in two years, why this tool is running in this way — which is a question that will be asked.

Deployment record capturing evaluation, regulatory position, placement, economics, owner and monitoring arrangement
Deployment record capturing evaluation, regulatory position, placement, economics, owner and monitoring arrangement
Share this article
Back to all articles

Keep reading

Related articles

See HealUDoc in action

From EHR to analytics, watch how one platform runs your entire hospital. Book a personalized walkthrough with our team.