A sensitivity model for AI in the health workforce, and the uncomfortable thing it says.

I built a model. You can open it and change every assumption yourself. It takes an illustrative 50,000-person integrated health provider, decomposes the workforce into eleven job families, applies task-level AI exposure ranges drawn from time-motion studies and federal occupational data, and then does the thing most AI-workforce analysis refuses to do: it separates what AI could technically touch from what an organization could actually bank.
Those turn out to be very different numbers. Here is what falls out.
Technical exposure is ten times the cashable number
Under the settings I consider honest — twenty percent realization of technical potential, half of released time schedulable, a seventy-five percent displacement threshold — 15,500 full-time-equivalents of work are technically addressable. That is the number you see in consulting decks, and it is not wrong. It is just not money.
Apply adoption friction. Discard the released minutes too scattered to become a schedulable block. Keep only the roles where one person’s affected share clears the threshold at which their remaining work can be bundled elsewhere. What survives is 1,489 roles.
Under ten percent. That is the conversion rate from AI can do this task to this changes the org chart.
If you are being pitched a market size derived from thirty-five percent of healthcare work is automatable — a real and defensible figure, and one my own model cites — understand that the operative number is closer to three percent of headcount, and that the gap is not pessimism about the technology. It is arithmetic about how work is distributed among people.
The model only makes money by not replacing someone
This is the finding I did not want.
The P&L bridge takes baseline operating profit of $450M on $15B of revenue — a three percent margin, which is a normal provider-group life — and adds $155M of realizable labor value, less $26M of AI running cost. Operating margin goes from 3.0% to 3.9%. Call it a twenty-nine percent improvement in operating profit. That is an enormous number for a business that lives on three points.
Now look at where it comes from. Every dollar is cloneable roles multiplied by their loaded cost. The 3,100 FTE of released capacity — the thing everyone means when they say AI will give clinicians their time back — contributes exactly nothing to the bridge.
That is not a flaw in the model. It is the structure of the business case. Released time becomes profit only if it is converted, and the conversions available are: see more patients with the same staff, or employ fewer staff. If you do neither, you have bought a better working day at real cost, which is a fine thing to buy and not a return.
So the honest version of the pitch is not AI gives your clinicians their afternoons back. It is AI lets you stop backfilling. Health systems rarely fire; against roughly 1.9 million annual healthcare openings nationally, a three percent role reduction disappears into attrition inside eighteen months. Nobody gets laid off and the P&L still moves. That is the actual mechanism, and it is worth saying out loud because it changes who you are selling to: a hiring-plan story told to a CFO, not a wellbeing story told to a chief medical officer. Founders who don’t know which room they’re in lose the deal in that room.
Half the value sits in one job family, and it isn’t clinical
Multiply cloneable roles by loaded cost per FTE, family by family. Administration, scheduling and revenue cycle produces about $54M of the $155M — thirty-five percent of the value from sixteen percent of the headcount, and just under half of every cloneable role in the model. Nothing else is close.
Why: administrative work is the only place where all three conditions hold at once. Exposure is high (45–65% of tasks), it is concentrated rather than smeared across everyone, and the residual work is reassignable because no one is personally accountable for it in the way a clinician is accountable for a decision.
This is a back-office model wearing scrubs. Which is not a criticism of health AI, but it is a criticism of how health AI gets positioned. The clinical story raises money and wins conferences. The administrative story is the one in the bridge.
Physicians generate the most exposure and the least value
Employed physicians throw off 825 technically addressable FTE-equivalents — the fourth-largest pool in the model. They yield 33 cloneable roles. A four percent conversion, against sixteen percent for administrative staff.
Two reasons, both well evidenced and both structural.
Fragmentation. A time-motion study of emergency clinicians recorded thousands of discrete tasks in under sixty hours of observation, with clinical-information-system work highly interrupted throughout. Shaving seconds off a hundred interactions does not assemble into a schedulable block. It assembles into a slightly less exhausting day.
Accountability. The residual work a physician performs after AI has drafted, retrieved and pre-populated is precisely the part that cannot be reassigned, because someone must be responsible for it. WHO guidance and the OECD’s health-workforce analysis both land here: clinical roles score high on augmentation and low on automation, and that is a governance fact, not a capability gap that a better model closes.
The physician day in the model gives up 133 gross releasable minutes out of 480. At twenty percent realization that is twenty-seven minutes. Real, valuable, worth paying for — and worth twenty-seven minutes, not a headcount line.
I take this as a genuine ceiling on the ambient documentation category rather than a knock on any particular company. Those products work; I’ve watched them work. They sell minutes to a buyer with limited ability to bank minutes, and their pricing power is bounded by that fact regardless of how good the transcription gets.
The cost assumption is doing more work than the technology assumption
At four dollars per released hour, AI costs $26M against $155M of value — a six-fold return. Good business. At the one-dollar default I originally shipped, the same model returns twenty-four-fold, which is the kind of number that gets a slide screenshotted and then mocked.
The footnote matters more than the number: cost per released hour excludes integration, change management, clinical validation, and vendor minimums. That is to say, it excludes everything that actually consumes a health system’s implementation budget. Four dollars is my attempt at honesty. I would not argue hard against six.
The general lesson is that in this model the denominator is softer than the numerator. People spend their scrutiny on whether AI can really do the task. The variance lives in what it costs to make it stick.
Concentration is the whole argument
The load-bearing assumption is not exposure. It is concentration — whether AI-affected work is spread thinly across everyone or clustered into a specialist subset. The model holds mean exposure fixed while varying concentration, which changes how many individuals clear the displacement threshold, and therefore changes the answer by roughly half.
The OECD makes the same point from the other direction: two people with the same job title can have radically different task exposure, and occupational averages obscure it. Concentration is where that fact shows up in a P&L.
It is also, and I want to be plain, the input I have the least evidence for. The exposure ranges are anchored in measured workflow studies. The concentration values are analyst judgment. Anyone applying this to a real organization should replace them with their own HRIS and activity-log data — which is why every input in the tool is editable, and why I would rather publish the model than the conclusion.
What I take from it as an investor
Value accrues where exposure is concentrated, contiguous, and attached to work that permits substitution rather than only supervision. That combination is common in revenue cycle, prior authorization, scheduling and laboratory workflow. It is rare at the bedside.
The corollary is a filter I now apply before I ask anything about the technology. When a founder shows me a time-savings number, I ask three questions:
Is the saving concentrated in the same roles, or spread across everyone?
Is it contiguous enough to schedule — a usable block, or two-minute fragments between irreducible tasks?
Is the residual work reassignable, or does someone still have to own it?
Three yeses is a business. Fewer is a feature — sometimes an excellent one, occasionally a beloved one, but priced as a feature and defended like a feature.
Ten saved minutes rarely equals one fewer person. It usually equals ten better minutes, which is a real good, and a different product.
The model is fully editable — workforce size, realization, schedulability, displacement threshold, task exposure by activity, and the financial assumptions. Evidence and provenance for every input are documented in the appendix. It is an illustrative benchmark, not any organization’s actual census, and technical exposure is not equivalent to labor elimination.
LifeX Ventures research. LifeX invests in software-driven science and longevity.
You must be logged in to post a comment.