Everyone is selling AI that makes decisions now. I run clinical operations, so let me tell you what I have learned from watching a hundred thousand of them get made: the decision is the easiest thing to automate, and the most dangerous thing to automate.
The industry has this backwards. Vendors start with the model and work outward toward governance. We started with the operation and worked inward toward the model. That ordering is the whole essay.
What deterministic means here
Claimatix runs on 12,366 deterministic rules. Deterministic means the most boring thing in software: the same inputs produce the same outputs, every time. No temperature. No sampling. No mood.
These rules constrain every critical step of a claim decision — what evidence is required, what the deadlines are, which guideline applies, who is qualified to make the call. The AI sits on top of that foundation and does what it is good at: reading a thousand pages of records in seconds, drafting the summary, flagging the pattern across cases. It proposes. The rules dispose. A human decides.
People hear "12,366 rules" and picture bureaucracy. It is the opposite. The rules are what make the system fast. When the guardrails are deterministic, you do not need a committee to approve what the machine did — you can see exactly what it did, because it does the same thing every time.
The rules were earned, not theorized
I did not sit in a room and invent 12,366 rules. They came out of eight years of running real clinical operations with P&S Network — more than 100,000 clinical decisions, each one with a paper trail, each one with a human being on the other end waiting for treatment.
Every rule exists because something went wrong once. A deadline got missed and a valid denial became invalid. A record arrived incomplete and the decision went out anyway. A guideline got applied to the wrong body part. Each failure got turned into a rule, and each rule got tested against the next ten thousand cases.
This is the part the AI vendors skip. They train on data; we operated on consequences. A model trained on historical decisions learns the patterns of the past, including the mistakes. A rule written from a failure prevents that failure from recurring. Those are different things, and in a regulated setting only one of them is defensible.
You cannot audit a vibe
Every claim decision in our system gets a decision-level audit trail: every input, every reviewer, every communication, every outcome. If a regulator, an IMR reviewer, or a judge asks why a decision was made, we can reconstruct it exactly.
Now imagine the same question asked of a pure large-language-model pipeline. Run the same case twice and you can get two different answers. Ask why, and the honest answer is a probability distribution over tokens. "The model said so" is not a defense. It is not even an explanation.
This is not a philosophical objection. It is an operational one. In California workers' compensation, an untimely or procedurally defective utilization review decision loses its protection — the dispute goes to a judge instead of independent medical review, and medical necessity becomes a litigated question. The procedure is the product. A system that cannot reproduce its own decisions cannot survive in that environment.
Determinism is what makes the audit trail possible. Reproducibility is what makes the audit trail meaningful.
The law already agrees with me
California figured this out before most vendors did. Labor Code section 4610 says plainly that no person other than a licensed physician competent to evaluate the specific clinical issues may modify, delay, or deny a treatment request for reasons of medical necessity (ca.elaws.us/law/lab_sec.4610).
Then in 2024, the state went further. SB 1120 — the Physicians Make Decisions Act, signed by Governor Newsom and effective January 1, 2025 — requires that when health plans use AI, algorithms, or software tools in utilization review, only a licensed physician or qualified healthcare professional may make the medical necessity determination. No AI tool may deny, delay, or modify care based on medical necessity, and the determination must be grounded in the individual patient's own medical record — not "solely on a group dataset" (Senator Becker's office, LexBlog analysis).
Read that carefully. The law does not ban AI. It bans AI as the decider. It requires the decision to be traceable to a specific patient's record and a specific licensed human. That is determinism-first architecture, written into statute. We built our system to satisfy a law that did not exist yet, because the operation demanded it before the legislature did.
The backdrop is ugly enough to explain the politics. In November 2023, a lawsuit accused UnitedHealthcare of using AI to deny claims. The bill's author cited Kaiser Family Foundation data showing more than 49 million claims denied nationwide in 2021, with less than 0.2% appealed (InsuranceNewsNet). When denial machinery scales faster than human oversight, legislatures step in. They always do.
The automation bias problem is real, and it is measurable
Here is the uncomfortable finding that every clinical AI deployment has to face. Goddard, Roudsari, and Wyatt's 2012 systematic review in the Journal of the American Medical Informatics Association found that clinicians overrode their own correct decisions in favor of erroneous computer advice in 6 to 11 percent of cases — and that erroneous advice was 26 percent more likely to be followed when it came from a clinical decision support system than from no system at all (JAMIA 19(1):121–127, DOI: 10.1136/amiajnl-2011-000089).
Think about that. A trained clinician, looking at a case, reaches the right answer — and then the machine confidently suggests the wrong one, and the clinician switches. Six to eleven percent of the time. The more confident the system sounds, the worse it gets.
Lyell and Coiera's 2017 systematic review, also in JAMIA, found the mechanism: automation bias tracks cognitive load and verification complexity. The harder it is for the reviewer to independently verify the machine's output, the more the reviewer defaults to trusting it (JAMIA 24(2):423–431, DOI: 10.1093/jamia/ocw105).
This is the sentence that turns a research finding into an engineering discipline: if cognitive load mediates the bias, then reducing cognitive load reduces the bias — and cognitive load is something you can design.
That is what the deterministic layer does. It does not ask the physician to verify a black box. It hands the physician a case where the routine work — completeness checks, deadline math, guideline matching — has already been done by rules that behave the same way every time. The physician's job narrows to the part only a physician can do: the clinical judgment. Verification gets cheap. Cheap verification gets done. Bias goes down.
Telling people to "be careful" does not work. Designing the work so that care is the path of least resistance does.
Where the AI actually belongs
I am not anti-AI. I run an AI company. The question was never whether to use it — it was where to put it.
AI belongs in the places where its strengths matter and its weaknesses do not: reading ten thousand pages of medical records and producing a faithful chronology. Drafting the first pass of a summary for a human to correct. Finding the pattern across a hundred thousand files that no human has time to read. Extraction, summarization, triage, draft — leverage, not authority.
AI does not belong in the places where its weaknesses are fatal: the final determination, the denial, the thing with a patient's name on it. In those places you need the same answer twice, a reason you can write down, and a licensed human willing to sign it.
NIST now has federal vocabulary for this. The Generative AI Profile (AI 600-1, July 2024), a companion to the AI Risk Management Framework, names the exact failure modes: confabulation — fabricated information presented as fact — and human-AI configuration, meaning inappropriate reliance on AI outputs (NIST.AI.600-1). The standards bodies have caught up to what anyone running clinical operations already knew: the risk is not that the AI is useless, it is that the AI is useful enough to trust and wrong enough to hurt.
The order of operations
So here is the thesis, stated plainly:
Determinism first, AI second. The system must be auditable, reproducible, and governed before it is intelligent.
Build the rules from the operation, not from the whiteboard. Constrain every critical step with logic that behaves the same way twice. Capture every input and every handoff. Then — and only then — let the models do what they are good at, inside the boundaries the rules set, with a licensed human making the final call.
AI does not make clinical decisions. Governed humans do. The deterministic layer is what makes the "governed" part true instead of aspirational.
Everyone else is trying to make the model trustworthy. We made the system trustworthy, and then gave it a model. That ordering is the moat, and it is the only one I would bet a patient's treatment on.
References
- California Labor Code § 4610. ca.elaws.us
- Office of Senator Josh Becker. "Governor Signs Physicians Make Decisions Act" (SB 1120), September 30, 2024. sd13.senate.ca.gov
- LexBlog. "California Implements New AI and Software Regulations for Insurers," October 9, 2024. lexblog.com
- InsuranceNewsNet. "New California law prohibits using AI as basis to deny health insurance claims." insurancenewsnet.com
- Goddard K, Roudsari A, Wyatt JC. "Automation bias: a systematic review of frequency, effect mediators, and mitigators." JAMIA 19(1):121–127, 2012. academic.oup.com
- Lyell D, Coiera E. "Automation bias and verification complexity: a systematic review." JAMIA 24(2):423–431, 2017. researchers.mq.edu.au
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework: Generative AI Profile (NIST AI 600-1), 2024. nvlpubs.nist.gov