AI Could Fix Utilization Management — Or the Industry Could Waste It. Here’s the Fork in the Road.

The Original Problem: Judgment at Scale

Clinical medicine is clinical medicine. Any well-trained physician, looking at a patient’s chart, can articulate two things: how sick this patient actually is (severity of illness), and how intensive the treatment they’re receiving needs to be (intensity of service). This is not exotic knowledge — it’s the daily bread of hospital medicine.

The problem was never clinical judgment. It was scale.

Health plans process thousands of hospital admissions a day. Having board-certified, specialty-specific physicians personally review every single case to determine whether it meets criteria for inpatient-level care is not clinically difficult — it’s economically impossible. So the industry made a labor-arbitrage decision: hire utilization review (UR) nurses to do the first pass, and have them approve inpatient stays when the patient is sick enough and the treatment intensive enough.

It didn’t take long to discover the obvious limitation: nurses, however skilled, don’t carry the same depth of clinical judgment as physicians when it comes to adjudicating complex, borderline cases. So the industry built scaffolding around that gap — utilization management criteria like InterQual and MCG. Strip away the branding and these guidelines are essentially structured libraries: clinical scenarios for a given condition, paired with corresponding intensity-of-service thresholds. When a case matches one of the recognized combinations, it clears for inpatient approval.

This is the system as it exists today: nurses using guideline logic as a proxy for physician judgment, at a scale no physician workforce could sustain.

A $5.3 Trillion Industry Built on Misaligned Incentives

U.S. national health expenditures now total $5.3 trillion, consuming 18.0% of GDP. There is broad, cross-partisan agreement that this system is bloated — layered with administrative processes that add cost without adding value, and that this layering has only grown over time.

Utilization management is a textbook example. It exists to answer a clinical question, but it has evolved into a compliance and adjudication apparatus: nurses applying criteria, physicians conducting peer-to-peer reviews, case managers preparing appeals, medical directors adjudicating second-level appeals, and independent review organizations handling external appeals — all to answer a question a single well-trained physician could often resolve by reading the chart.

The uncomfortable truth is that most of the players in this ecosystem are not actually incentivized to shrink it. When more money flows through the system, more entities profit from the flow itself — regardless of whether that flow reflects genuine value. Look at how MLR (medical loss ratio) rules interact with vertically integrated health plans, how intracompany eliminations can obscure the real cost of services shuffled between related entities, and how revenue recognition practices can be used to inflate reported numbers. None of these dynamics reward shrinking total cost of care. They reward growing it, or at minimum, not disturbing it.

Real technological progress is supposed to break patterns like this. The value proposition of any transformative technology in healthcare should be straightforward: reduce waste, shrink the administrative footprint, lower premiums, and improve outcomes — all at a lower overall cost. That is the promise AI arrived carrying.

What AI Could Actually Do

There is a genuine, achievable opportunity here — not speculative, not five years out. A well-trained clinical AI model, fed a case, could in principle:

  • Understand the clinical picture the way a well-trained specialty physician would — not just pattern-match keywords, but synthesize the actual clinical narrative.
  • Understand the spirit and intent of InterQual or MCG criteria, not merely their literal text — recognizing that these guidelines were written as decision-support heuristics for a specific clinical scenario, not as rigid, exhaustive rulebooks meant to be parsed like a contract.
  • Separate correlation from causation in a patient’s presentation — distinguishing findings that are clinically meaningful drivers of intensity of service from findings that are simply co-occurring.

If built and deployed correctly, this could fundamentally restructure utilization management. It could dramatically reduce the number of UR nurses employed on both the hospital and payer sides. It could support a direct EMR interface between providers and payers — so that reviewers (human or AI) are working from the full longitudinal chart rather than a faxed excerpt. It could cut down peer-to-peer reviews, next-level appeals, and external appeals. And it could compress the entire revenue cycle toward something resembling instant approvals and near-immediate payment.

This is not a fantasy. It’s a coherent, achievable end state — if the technology is pointed at the actual problem.

Where This Could Go Wrong

Here’s the risk worth naming plainly.

If healthcare defaults to the path of least resistance, both clinical and non-clinical executives will steer AI in the opposite direction from the one described above. Instead of training models to understand the spirit and intent of InterQual or MCG criteria — and to separate correlation from causation the way a skilled physician would — the temptation will be to build AI that adheres to the exact letter of the guidelines instead. That would be a meaningfully different, and meaningfully worse, design goal. Literal, rigid rule-matching is precisely what leads to misapplied criteria and wrong decisions, because clinical reality routinely doesn’t fit neatly into a guideline’s enumerated scenarios. A guideline is a heuristic written by humans anticipating common patterns — it was never meant to be treated as an exhaustive, letter-perfect legal text.

There’s a second version of this risk, compounding the first. Where understanding of the technology is limited — or trust in it is low — the path of least resistance is an “assist” tool bolted onto existing case managers, rather than a system genuinely trusted to reason through cases. That would squander what this technology is capable of. It would be automation applied at the margins of an already-inefficient process, rather than a redesign of the process itself.

If the industry defaults to both of these — letter-matching over spirit, and assist-only deployment over genuine reasoning — the outcome is predictable: the same pattern of improper approvals and improper denials continues. The appeals machinery — peer-to-peers, second-level appeals, external reviews — keeps expanding. The status quo doesn’t change. And there’s now an additional layer of cost sitting on top of it, in the form of the AI tooling itself, which was supposed to be the thing that shrank the system.

The Metric That Will Eventually Matter

If this path is the one taken, providers won’t sit still — they’ll adapt, learning to generate appeals at a scale that matches whatever volume of AI-assisted denials comes from payers. That’s an arms race, not a resolution, and it’s worth heading off before it becomes the default equilibrium.

Health plan executives have a chance to get ahead of this by internalizing one principle now: the success of these AI agents should never be measured by initial denial rate. A high initial denial rate is trivially easy to produce and tells you almost nothing about whether the underlying clinical judgment was correct. The metric that actually matters — the one that reveals whether an AI system is making sound clinical determinations or just efficiently generating disputes — is the rate of denials that get overturned on appeal.

An AI that denies aggressively up front but gets reversed at a high rate on appeal hasn’t reduced administrative waste. It’s manufactured more of it, just shifted downstream and dressed up as innovation. If the industry wants AI to actually bend the cost curve rather than add a new layer to it, that’s the number that should be on every executive dashboard — not how many cases got flagged on day one.

The technology to do this right exists. Whether the incentives in this industry will allow anyone to build it that way is a separate question entirely.

Disclaimer: These are my own views, shaped by my experience in clinical utilization review — not the position of any employer, client, or organization I’m affiliated with.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top