Ranjeet Randhawa← Writing

Essay

AI Hype in Healthcare

By Ranjeet Randhawa

AI does not make clinical decisions. Governed humans do.

I have said that sentence in rooms full of vendors, buyers, and physicians for years now. It usually lands one of two ways. The operators nod. The vendors change the subject.

I have spent 25 years in healthcare technology. Through P&S Network, the clinical operation I have worked with for over eight years, I have watched more than 100,000 clinical review decisions move through a governed process — utilization review, peer review, the full apparatus. That is not a metaphor. That is files. Dates. Reviewer names. Guideline citations. Outcomes.

So when someone tells me their AI "makes clinical decisions," I do not reach for an opinion. I reach for a question: show me the audit trail.

The demo is always beautiful

Here is the pattern I have watched repeat for a decade. A vendor walks into a health plan or a UR organization with a beautiful demo. The model reads a chart in seconds. It summarizes the record. It recommends approve or deny, with citations that look authoritative. The buyer — usually an executive who has never sat through a utilization review decision — sees speed and thinks scale.

Nobody in the room asks the only question that matters: when this is wrong, whose name is on it?

Because it will be wrong. Not always. Not even often. But sometimes, in the cases where "sometimes" means a person does not get care. And in that moment, "the model recommended it" is not an accountable answer. It is not even an answer.

I am not against the technology. I build with it every day. I am against the story being sold around it — the story where the machine decides and the human nods. That story has a body count in the literature, and I will get to it.

What the files teach

California publishes something unusual and valuable: an annual report on its Independent Medical Review program, the system that resolves disputes when utilization review denies a treating physician's request. The numbers are public, and they are instructive.

In the report covering 2024 data, treatment denials were overturned 12.7% of the time overall. But look at which categories get overturned most: evaluation and management services (23.1%), programs (22.2%), and behavioral and mental health services (20.1%). The 2025 data tells the same story in a different order: program services at 18.5%, behavioral and mental health at 18.3%, evaluation and management at 16.8%. (California DWC IMR annual reports; DIR newsline summary)

Think about what those categories have in common. They are the decisions where individual patient context matters most. A program of care is not a single procedure — it is a sequence, and its necessity lives in the trajectory. Behavioral health turns on history, presentation, and judgment calls that resist checkboxes. An evaluation is, by definition, someone looking at the whole patient and forming a picture.

These are precisely the decisions where a shortcut is most dangerous. And they are precisely the decisions that get overturned at two to three times the average rate — after a human reviewer actually reads the file.

That is what 100,000 files taught me before I ever read the state data: the hard cases are hard because they are human. Automating the assembly of the evidence is a gift. Automating the judgment is a liability wearing a costume.

The science of the nod

There is a name for what happens when a confident system meets a tired human: automation bias. It is not a new finding, and it is not a matter of opinion.

Goddard, Roudsari, and Wyatt's 2012 systematic review in the Journal of the American Medical Informatics Association found that clinicians overrode their own correct decisions in favor of erroneous technology advice between 6% and 11% of the time. (Goddard et al., JAMIA 2012)

Lyell and Coiera's 2017 follow-up in the same journal found something more actionable: automation bias is mediated by cognitive load and verification complexity. The harder it is to check the machine's work, the more the human defaults to trusting it. (Lyell & Coiera, JAMIA 2017)

Read that again, because it is the whole design problem in one sentence. The bias is not a character flaw. It is a load problem. And every vendor demo you have ever seen increases the load it claims to reduce — fluent, confident output that takes real effort to verify, handed to a reviewer who has forty more files in the queue.

In 2023, researchers demonstrated it in mammography: AI-generated BI-RADS suggestions measurably shifted radiologists' performance — the automation bias effect, live, in cancer screening. (Dratsch et al., Radiology 2023)

So when a vendor tells you their system keeps "a human in the loop," ask what the loop is actually made of. A human presented with a confident recommendation and a forty-file queue is not a reviewer. They are a rubber stamp with a pulse. The loop has to be designed — for verification, not deference. That means the system shows its work, cites to source, flags what is missing, and makes disagreement cheap. If disagreeing with the machine takes ten clicks and agreeing takes one, you do not have human oversight. You have theater.

When the story goes wrong

The worst version of this story is not hypothetical. It already happened, at the largest health insurer in America.

UnitedHealth's subsidiary naviHealth used an algorithm called nH Predict to estimate how many days of post-acute care — nursing homes, rehab — elderly Medicare Advantage patients should need. According to the class-action complaint filed in Minnesota federal court in November 2023 (Estate of Gene B. Lokken et al. v. UnitedHealth Group), those predictions became de facto coverage cutoffs. Employees were allegedly instructed not to deviate from the model's projections. When patients and families appealed and actual humans reviewed the cases, the denials were reversed roughly 90% of the time. The catch: only about 0.2% of affected patients ever appealed. (Case background; complaint coverage)

Ninety percent wrong when a human finally looks. Two-tenths of a percent ever get that look.

That is the hype story carried to its conclusion: the machine decides, the human nods, and the accountability evaporates into a proprietary black box. Note the detail that should chill every buyer: the company's defense was that the law already requires humans to make the final call. The plaintiffs' allegation was that the humans were not permitted to disagree. A human in the loop who cannot say no is not governance. It is decoration.

California, to its credit, wrote the rule down plainly years ago. Labor Code section 4610: no person other than a licensed physician competent to evaluate the specific clinical issues may modify, delay, or deny a treatment request for reasons of medical necessity. (Cal. Labor Code § 4610) The law understood before the hype cycle did: the decision belongs to a licensed human, full stop.

Five questions for any AI vendor

I keep a short list. It fits on an index card. Any vendor selling AI into clinical decisions should answer all five without blinking:

  1. Who makes the final call, and is their name recorded? Not a role. A name. On the decision. If the answer is "the physician reviews it," ask how many they review per hour and what happens when they disagree.

  2. Can every guideline citation be traced to its source? A citation to "MTUS" is not a citation. Show me the chapter, the section, the text. If the model paraphrases a guideline, I want the original one click away.

  3. Is the analysis built from this patient's records — or from a pattern? Population statistics are useful context. They are not a medical opinion about the person in front of you. The nH Predict failure was a pattern-matching machine standing in for an examination.

  4. What happens when the evidence is missing? This is the question that separates tools from toys. A serious system says "the record does not contain X, and the decision cannot be made without it." A hype system fills the gap with fluent confidence. Guess which one the demo shows.

  5. What do you measure? Overturn rates. Appeal rates. Disagreement rates between the model and the reviewer — and whether disagreement is trending toward zero, which would tell you the loop has collapsed into deference. If they measure only speed and volume, they are selling throughput, not judgment.

Zero hype

I call my posture zero hype, and people sometimes hear cynicism. It is not. It is the only honest way to build in a domain where the output of your system is someone's care.

Here is what zero hype looks like in practice. Deterministic rules where rules suffice — the boring, auditable kind, tens of thousands of them, each one traceable to a guideline and a human who wrote it. Machine learning where it adds leverage — extraction, classification, assembly — never as the decider. Every decision carrying the evidence it rested on, the guideline it cited, and the name of the human who made it. An audit trail you could hand to a regulator, a judge, or the patient's family without flinching.

That is slower to build than a demo. It is also the only thing worth building.

The hype cycle will move on — it always does — and what will be left standing are the systems that kept their receipts. Twenty-five years in, I have watched enough cycles to know the ending. The vendors selling "AI that decides" are selling a shortcut through the one part of the process that cannot be shortcut: accountability.

AI does not make clinical decisions. Governed humans do. Everything else is marketing.


References

  1. California Division of Workers' Compensation. Independent Medical Review Annual Report, 2025. dir.ca.gov
  2. California Department of Industrial Relations. DIR newsline release 2026-43. dir.ca.gov
  3. Goddard K, Roudsari A, Wyatt JC. "Automation bias: a systematic review of frequency, effect mediators, and mitigators." JAMIA 19(1):121–127, 2012. academic.oup.com
  4. Lyell D, Coiera E. "Automation bias and verification complexity: a systematic review." JAMIA 24(2):423–431, 2017. researchers.mq.edu.au
  5. Dratsch T, et al. "Automation Bias in Mammography: The Impact of Artificial Intelligence BI-RADS Suggestions on Reader Performance." Radiology, 2023. doi.org/10.1148/radiol.222176
  6. Healthcare Uncovered. Background on the UnitedHealth AI denial lawsuit. healthcareuncovered.substack.com
  7. CRBC News. Coverage of the complaint. crbcnews.com
  8. California Labor Code § 4610. california.public.law
← Back to writing