10 Hospital AI Tasks: What Should AI Agents Never Touch?

Ten tasks are written on a hospital front-office whiteboard. A vendor has offered to hand some of them to an AI agent and the administrator wants to know which ones.
Before reading on, sort them in four buckets, categorising them in as functions that run on its own, run with a person approving, need fixing before anything run it and never goes to an agent.
Sending appointment reminders to patients who have opted-in.
Acknowledging a new website enquiry within seconds.
Drafting a reply to "how much does a knee replacement cost?".
Qualifying a new enquiry and routing it to the right department.
Asking a patient for a Google review after a visit.
Chasing enquiries that arrive by WhatsApp, phone, web form and walk-in, logged in four different places by whoever is free.
Deciding who gets a fee concession, depending on who is at the desk that day.
Releasing a lab result to a patient.
Approving a prescription refill request.
Replying to a WhatsApp message that reads: "My father has chest pain and is breathing fast".
I'm Boudhhayan Duttaa, founder of Batti Jalao, an AI-led healthcare marketing growth consultancy based in Guwahati. We build content and marketing systems, such as BattiLynk AI, BattiSense and other custom agentic AI, across healthcare sub-segments and we run BattiOps, a fixed-price Agentic AI readiness audit for healthcare operators. This piece sorts the list with the same two-question rubric BattiOps uses, shows where the answers are less obvious than they look and sets out what a register of excluded tasks should contain.
Why Do Most Healthcare AI Pilots Fail Before the First Message Is Sent?
Roughly 95% of enterprise Agentic AI pilots deliver no measurable P&L impact, according to MIT NANDA's 2025 research while Gartner expects more than 40% of Agentic AI projects to be cancelled by the end of 2027. Deloitte's reading of the MIT findings puts the cause in workflow and organisational policy, not in the model.
Put that next to the whiteboard and the pattern is easy to see. A pilot that skips the sorting turns a list of tasks into a production system by accident. Every item gets the same treatment, so the reminder and the chest-pain message end up with the same level of oversight. BattiOps begins with the sort instead and it begins with a 20-minute qualification call that can end with us telling you honestly that you are not a fit.
The Rubric: Two Questions, Multiplied
Every process gets two scores. Suitability: Can an agent run this reliably? Risk: What happens if the guardrails fail? Priority is the two multiplied and each process lands in one of four buckets namely Deploy Now, Pilot, Redesign First or Do-Not-Automate.
The second question is easier to answer with a lens from ICMR's 2023 ethical guidelines for AI in healthcare. The guidelines start from the premise that AI cannot be held accountable for its decisions, so humans stay the final decision-makers and they ask organisations to consider whether a decision made by an AI can be reversed. In plain terms, if this goes wrong, could it be undone?
The Reveal, Bucket by Bucket
Deploy Now: Tasks 1 and 2. Both are structured, repeatable and reversible. A reminder that goes out an hour late costs nothing to correct and an acknowledgement says only that a message arrived. The condition applies to the first task is that the patient must have a consent record for that specific purpose, which is a DPDP question before it is an AI one. Even here, the agent runs with a person checking its work for 30 days.
Pilot: Tasks 3, 4 and 5. Suitable, but sensitive, so the agent drafts and the staff approves until it has earned more freedom. A cost reply is a commitment that depends on the patient's case. Routing an enquiry is usually recoverable, but a message that describes symptoms should never be routed on autopilot. The review request is the subtle one in which the patient whose discharge went badly should not receive a polite ask for five stars and only a person knows which patient that is.
The Pilot bucket has a catch most pilots miss. 'Staff approves' only protects anyone if the approval is a real gate. Research on human oversight of healthcare AI, published in Mayo Clinic Proceedings: Digital Health, points to a documented tendency for people to over-trust computer output. A receptionist clicking approve on forty drafts an hour is not oversight. That is why a 'Pilot' process graduates only after a clean run, on evidence and not because it has been quiet for a month.
Redesign First: Tasks 6 and 7. This is the most underrated bucket, because the answer is not no, rather it's not yet. An agent cannot tidy a process nobody has defined. Enquiries logged in four places would simply be automated badly and a fee-concession rule that lives in one person's head would be reinvented by the agent every time, differently. The fix is one log, one owner and one written rule. Once that exists, these tasks can be scored again and may move up.
Do-Not-Automate: Tasks 8, 9 and 10. Releasing a lab result and approving a prescription refill are clinical acts and BattiOps names them explicitly in a written Do-Not-Automate register. The third is the hard one.
In February 2026, researchers at Mount Sinai published a structured test of ChatGPT Health in Nature Medicine. The tool under-triaged just over half of the clear emergency scenarios. In one asthma case it identified early warning signs of respiratory failure in its own explanation and still advised the patient to wait. The study used clinical vignettes, not real conversations and it tested a general-purpose consumer tool, not a hospital's own assistant. The lesson that carries over is that the narrower you're in your definition, the more useful the results would be. Nobody should find out how an agent handles an urgent message by waiting for a real one. BattiOps scopes this out on purpose with no autonomous emergency communications, no clinical or diagnostic AI and no automating clinical decisions.
What a Register of Excluded Tasks Should Contain
A list of 'Never' items is only useful if someone can act on it. Whatever format a hospital chooses, each entry should answer four questions:
What is the task and which named person owns it?
Why is it excluded? Be specific about the harm and whether it can be undone.
What, if anything, may an agent do near it? For example, whether it may hand a message to a human faster than a person would have spotted it.
When is the exclusion next reviewed and by whom?
ICMR's guidelines ask organisations to define legal responsibility before adopting AI for clinical use. A register is where that definition lives in writing. It is also the one document a vendor's pitch has no reason to include, because a pitch sells what a tool can do and a register lists what it will not.
From the Sort to the Build
A sorted list is a diagnosis and a diagnosis that stops at a report is only half the job. A BattiOps audit runs in two to three weeks and ends with a 12 to 15 page report and a 60-minute walkthrough. The deployment step covers only what scored safe and the work is checked by a human for 30 days. We work on surfaces we control end to end including WhatsApp, the web, Google Business Profile and the CRM. Anything that needs the hospital's own HIS, EMR or LIS is handed to the hospital's vendor, as a rule. If the work is commissioned within 90 days, 50% of the audit fee is credited against the first deployment wave, so the diagnosis is not a dead end.
Common Mistakes
Mistake One: Sorting by what a tool can do instead of what happens when it is wrong.
Mistake Two: Treating 'Staff approves' as a safeguard without ever checking whether the approvals are real.
Mistake Three: Letting a vendor's demo set the list, when the hospital's own whiteboard should.
Mistake Four: Writing the exclusions once and never reviewing them as the tools change.
Is This Only Relevant to Hospitals With a Large Operations Team?
No. A three-person front desk has a shorter whiteboard, not a different one. The buckets are the same and so are the two or three tasks that should never reach an agent. Smaller teams are often more exposed, because nobody is employed to ask the question.
💡The whiteboard was never the risk. Leaving the fourth bucket unwritten was.
Note on the examples: The ten-task list and its placements are an illustrative composite written to show the rubric at work, not the output of a real client audit. A real BattiOps audit scores each process against the hospital's own volumes, owners and tooling and may place the same task differently.




