What the policy has to decide before anyone drafts it
A hospital AI governance policy is not an ethics statement. It is a decision-rights document. It answers four questions: who may propose an AI tool, who may approve one, what has to be true before it touches a patient, and who is accountable when it goes wrong. Everything else, including principles, values and references to national guidance, is context wrapped around those four answers. Policies that lead with principles and never name a decision-maker get adopted by a board and then ignored by everybody who has to act on them.
The second thing to settle is scope. Does the policy cover only tools touching clinical decisions, or also those in finance, human resources and marketing? Does it cover a feature inside your existing HMS that the vendor describes as intelligent? Does it cover a consultant using a consumer chatbot at home to draft a summary? A deliberately narrow scope is defensible. An unstated scope is not, because every borderline case then becomes an argument at the meeting.
Third, decide what the policy is anchored to. The ICMR ethical guidelines for the application of artificial intelligence in biomedical research and healthcare set out guiding principles that map reasonably onto hospital practice, and the DPDP Act 2023 supplies your data obligations. Citing both is useful. Treating either as a substitute for your own decision rules is not, because neither of them will tell you who signs off a sepsis alert in your emergency department.

Who sits on the committee, and who chairs it
The committee needs enough authority to stop a project and enough breadth to see what a project touches. In practice that means a clinician chair with standing, usually the medical director or a senior consultant with an informatics interest, rather than the IT head. When IT chairs, the committee is read as a technology forum and clinicians stop attending. When a clinician chairs and IT presents, it reads as a clinical governance body, which is precisely what it has to be to work.
Keep it small enough to actually meet. Eight to ten members who turn up will outperform a twenty-name constitution that never reaches quorum. Legal, procurement and biomedical engineering are usually better handled as standing invitees called in for specific items than as permanent seats. What you cannot skip is a nursing voice and a medical records voice, because those two functions absorb most of the workflow consequences of anything the committee approves.
Meet on a fixed cycle, monthly or six-weekly, and publish an agenda deadline. AI proposals arrive with commercial urgency attached, and an ad hoc committee will end up approving things in a messaging group between meetings. A fixed calendar with a documented fast-track route for genuinely urgent items is the only structure I have seen hold up under that kind of pressure over a full year.
Seats worth filling on an AI governance committee
- A clinician chair with the authority to halt a live deployment
- Clinical informatics, to present and challenge technical assessments
- The person carrying DPDP accountability for the hospital's data
- Quality and accreditation, holding the NABH and indicator implications
- Nursing and medical records, who absorb the workflow consequences
The use-case register is the actual policy
If you write only one artefact, write the register. It is a single table listing every AI capability in use or proposed across the hospital, with enough fields that a stranger can tell what it does, who owns it, what data it touches and whether it was approved. Committees without a register spend their meetings rediscovering what already exists. Committees with one spend their meetings making decisions, which is the difference between governance and a discussion group.
Populate it by amnesty rather than by audit. Announce that anything declared in the next six weeks will be assessed on its merits, and that anything discovered afterwards is treated as unapproved. You will find more than expected: a radiology plug-in bundled into a PACS upgrade, a chatbot on the website procured by marketing, a transcription tool one department bought on a corporate card. All of these become governable the moment they are visible.
The register also becomes your evidence. A NABH assessor asking how you control clinical software, a DPDP question about processing purposes, and a board asking what the hospital is spending on AI are all answered from the same table. Keep it in one place with a named owner and a review date on every row, and resist the very strong pull towards letting it grow into a general project tracker.

Fields every use-case register row needs
- Tool name, vendor, and the specific module or version deployed
- Clinical or administrative purpose, written in one plain sentence
- Autonomy tier and who reviews the output before it is acted on
- Data classes processed, lawful basis, and where processing happens
- Named owner, approval date, next review date and retirement trigger
Approval gates from proposal to production
Gate one is intake: a one-page proposal naming the problem, the decision the tool would change, the sponsor and the expected users. A surprising number of proposals die here, correctly, because the author cannot name a decision that would change. Gate two is risk assessment. Is this a medical device under the Medical Device Rules 2017, what is the licensing position, what data does it touch, and does it warrant a data protection impact assessment before anything else proceeds?
Gate three is procurement and contract, where change-control terms, logging obligations and exit rights get negotiated while you still have leverage. Gate four is evaluation: shadow running or a bounded pilot with pre-agreed success criteria, a fixed end date and a named analyst who will produce the result. Gate five is the production decision, which is a genuinely separate decision requiring an operational owner, a budget line and a monitoring plan rather than the pilot simply continuing.
The single most valuable rule in the entire policy is that a pilot cannot become production by inaction. Set an end date at gate four and make the tool stop on that date unless gate five has been cleared. Hospitals that skip this accumulate a long tail of half-owned tools that nobody formally approved, nobody monitors and nobody can switch off without triggering an argument with a department head.
Autonomy tiers and the human in the loop
Define tiers and place every use case in one. Tier one is informational: the tool surfaces data or drafts text and a human authors the clinical content. Tier two is assistive: the tool proposes something specific and a clinician confirms or rejects it before anything happens. Tier three is autonomous with oversight, where the tool acts within a bounded rule set and a human reviews afterwards. Tier four is fully autonomous, which almost nothing in a hospital setting should be.
The tier drives everything downstream: how much validation you demand, what you log, what you tell the patient, and what your consent position is. It also disciplines the vendor conversation usefully. Asking which tier a product operates at in this specific workflow moves the discussion from capability claims to accountability, and it exposes the gap between what a tool can technically do and what you actually intend to permit it to do.
Be honest that the human in the loop degrades under load. A clinician confirming forty suggestions an hour is not meaningfully reviewing them by the twentieth. If your safety argument rests entirely on human review, you have to design for the conditions under which that review really happens, which means the night shift, the high-volume OPD and the junior registrar, rather than the conditions present during the vendor demonstration.
“Our policy said a doctor reviews every AI suggestion. Then we watched the OPD at eleven in the morning. Review was a click. We had to redesign the tool to show fewer, better prompts before the policy meant anything at all.”
Monitoring duties once a tool is live
Approval is the beginning of the obligation rather than the end of it. For every production tool the policy should name what is monitored, how often, by whom, and what threshold triggers escalation. At minimum that means utilisation, override rate, a periodic sample review by a clinician, incident reports involving the tool, and any vendor notification of change. Four of those five are cheap to run. The sample review is the one that costs and the one that finds things.
Override rate is the most informative single number you will collect. A tool nobody overrides is either excellent or invisible, and you cannot tell which without asking the users directly. A tool overridden most of the time is generating work rather than removing it. Track it by user and by unit, because a tool that works well in one department and poorly in another is common, and the pattern disappears inside a hospital-wide average.
Route incidents through the channel you already have. AI-related events should enter the same incident reporting system as medication errors and near misses, carrying a flag rather than living in a parallel process. A separate AI incident register will not be used by anyone. What you do need to add is a question on the existing incident form asking whether a decision-support or automation tool was involved, so the flag gets set at the point of reporting.

Retirement criteria, and why nobody writes them
Every policy has an approval section and almost none has a retirement section, which is exactly why hospitals accumulate tools. Write the triggers in advance, because deciding to switch something off is far easier when the criterion was agreed while everyone was still optimistic. The triggers are not exotic: performance below an agreed floor, vendor end-of-support, a lapsed licence, sustained low utilisation, an unresolved safety incident, or the disappearance of the workflow the tool served.
Retirement needs a procedure as well as a trigger. Who tells the users, how far in advance, what replaces the function, what happens to historical outputs already sitting in the record, what data comes back to you and in which format, and how the integration is unwound without breaking something adjacent. Exit terms negotiated at gate three make all of this survivable. Exit terms discovered at retirement make it expensive and slow.
There is one more thing worth writing down, and it is the hardest to get agreed. State that a tool may be retired because it is no longer needed, not only because it failed. Teams become attached to systems they helped build, and a policy that permits removal only after a failure quietly guarantees that removal only ever happens after a failure. A documented review date against each capability turns that into a scheduled conversation rather than a confrontation.
Retirement triggers to write into the policy
- Sustained performance below the floor agreed at the approval gate
- Vendor end-of-support, licence lapse or unresolved regulatory change
- Utilisation below the level that justifies the monitoring burden
- A safety incident that remains unresolved after an agreed period
- The underlying workflow, service line or data source ceasing to exist


