Which Risk Tier Pays Back? Churn Math vs Camden Null

TakeawayDetail
Stratified benchmarks changed who eats the penalty, not the size of the meal.Per JAMA (May 7, 2025; 334(1):85–87; DOI 10.1001/jama.2025.5209), stratified benchmarks in the ESRD Treatment Choices model reduced penalties for dialysis facilities serving higher proportions of low-income or Black patients, yet overall financial penalties under the model increased after implementation.
Tiered early-warning systems are built to suppress false alarms, not to rescue more patients.An npj Digital Medicine framework (article s41746-026-02522-8) targets high false-alarm rates in in-hospital mortality prediction for emergency departments, where Sax et al. (JAMA Netw Open 8:e258498, 2025) documented triage inaccuracy and delayed care for high-risk conditions.
Internal calibration is a poor predictor of deployed calibration.ORDER-DR's referable-risk expected calibration error was 0.049 ± 0.001 on held-out APTOS test data but 0.160 ± 0.008 on 1,744 external Messidor-2 fundus images (Frontiers, 2026-08-03).
Elite discrimination scores rarely survive a change of population.ORDER-DR posted quadratic weighted kappa of 0.8960 ± 0.0049 on held-out APTOS splits versus 0.6423 ± 0.0364 on Messidor-2 with validation-calibrated thresholds, while macro-F1 fell from 0.6832 ± 0.0279 to 0.4578 ± 0.0332.

Five years on, the lesson hasn't stuck. Super-utilizer programs still own the moral high ground and the budget line, while the mid-tier work that actually clears break-even — transitions coaching, follow-up outreach, disciplined escalation thresholds — fights for scraps. The peer-reviewed record keeps complicating the intuitive story: stratified benchmarks inside CMS's ESRD Treatment Choices model shielded safety-net dialysis facilities yet left overall penalties higher than before, and risk models that look sharp on their home data wobble badly once they travel.

Which tier pays back, then? Rarely the one with the best origin story. The tiers that clear break-even tend to share unglamorous traits: modest per-member spend aimed at members whose utilization is still directionally movable, and risk assignments honest enough to survive contact with a new population — because a model whose calibration error balloons from 0.049 to 0.160 across datasets isn't allocating care, it's reallocating it at random.

AHRQ's Medical Expenditure Panel Survey hands payers an uncomfortable pairing of facts. The top 1% of spenders account for roughly a quarter of national health expenditures, yet only about a third of people in the top decile of spending are still in that decile the following year. Both facts are true simultaneously, and the second one quietly guts the business case built on the first.

Tiered glass skyscrapers rising through dawn financial district
Tiered glass skyscrapers rising through dawn financial district

Churn Math

The pitch you will hear in 2026 vendor meetings — our top 5% of members drive half our costs, so that's where care management pays back — commits a category error: it treats spending concentration as proof of addressable savings. Concentration tells you where the money was. Persistence tells you whether it will still be there. When only about a third of the top decile holds its position year over year, a roster built on last year's claims is mostly a list of people already reverting toward their own baseline before anyone dials them.

Regression to the mean, in this setting, is a mechanism rather than a nuisance variable. Intensive programs enroll members at their personal spending peak — that is the eligibility event. Under a true zero program effect, their costs decline afterward almost by construction, because the peak was partly bad luck: an exacerbation, an injury, a billing anomaly. Pre/post vendor dashboards then book the entire decline as savings. Year-one phantom ROI is not a measurement flaw you patch with a better dashboard; it is the default output of enrolling at the peak and benchmarking against the peak. The audit that exposes it costs nothing: request the distribution of each member's own historical spend at enrollment. If the cohort clusters at its maximum, the projected decline is arithmetic, not effect.

Honest accounting then hits a ceiling. Super-utilizers carry 60%-plus 180-day readmission rates, but the drivers are advanced heart failure, frailty, polypharmacy, and housing instability — clinical and social load a care manager can accompany but cannot remove. A high baseline readmission rate looks like opportunity; most of it is disease trajectory, not preventable margin. The addressable slice hides inside a very large number.

Now invert the lens. The levers that actually cut admissions — a completed follow-up visit within 7 days of discharge, medication reconciliation, scheduled chronic-care contact — are protocolized, cheap, and deliverable at low monthly cost to mid-acuity members who still have unaddressed care gaps. Those members sit between the 50th and 80th predicted-cost percentiles with an inpatient stay in the past 12 months, precisely the band funded through CPT 99495/99496 transitional care management. The payback lives below the 80th percentile, not above the 99th, because that is where modifiable risk still outweighs the cost of touching it.

Eight hundred of Camden, New Jersey's highest-utilizing patients were randomized, and the intensive version of care management returned nothing. In January 2020, Amy Finkelstein and colleagues published the randomized evaluation of the Camden Coalition Core Model in the New England Journal of Medicine: 800 high-utilizing members assigned to wraparound care management or usual care. At 180 days, 62.3 percent of the intervention group had been readmitted versus 61.7 percent of controls — an adjusted relative risk of 0.99 (95% CI 0.88–1.11) — with no significant difference in readmissions or total costs. Treat this as the field's cleanest test of the folk theorem that cost concentration proves addressable savings. It returned zero, and no scaled replication has credibly overturned it.

The counter-evidence sits in the same literature and points down-market. According to the Cochrane review of discharge planning led by Gonçalves-Bradley and colleagues, structured, protocolized discharge planning modestly reduces readmissions, with a relative risk around 0.92. Set the two findings side by side and the ranking is unambiguous: light-touch standardization beat bespoke super-user wraparound in the only head-to-head reading of the evidence. The mechanism is unglamorous — medication reconciliation, a scheduled follow-up contact, a named owner for every transition — applied uniformly rather than heroically, which is exactly why it survives scale.

Line itemTop-1% intensive modelRising-risk TCM band (50th–80th percentile)
Entry triggerLast year's spend rankDischarge event + ≥1 inpatient stay in past 12 months
Roster persistenceAbout 1/3 of top decile repeats next year (AHRQ MEPS)Anchored to the index admission, not prior-year rank
Readmission baseline60%+ at 180 days; HF, frailty, polypharmacy, housing instabilityMid-acuity with unaddressed care gaps
Delivery costFour figures per member/month at 15–30 caseloadsLow monthly cost per member
Break-even barRoughly one $15,000 admission prevented per enrollee/year7-day follow-up visit + med reconciliation + scheduled contact
Funding posture12-month gated pilot, pre-committed kill thresholdFund as portfolio core via CPT 99495/99496
Overcast morning along Camden canal towpath weathered brick
Overcast morning along Camden canal towpath weathered brick

The Null That Won't Die

The adoption gap makes this either a scandal or an opportunity, depending on which side of the contract you sit. Medicare already pays separately for the intervention that works: a transitional care management fee per qualifying discharge (CPT 99495/99496) and a monthly chronic care management fee (CPT 99490). Published billing audits find TCM billed on well under 10 percent of eligible discharges. The highest-evidence, lowest-cost lever in the stack is the least pulled. A payer does not need a new product here; it needs claims surveillance that identifies eligible discharges and forces the documentation-and-contact workflow to completion.

The formal basis for refusing to treat "expensive" as one intervention target comes from the National Academies' 2017 report, Effective Care for High-Need Patients, which partitions the high-cost population into biologically and operationally distinct segments — high-cost chronic, catastrophic short-term, and end-of-life. Members who share a cost percentile can share nothing else: a catastrophic admission resolves, a chronic trajectory compounds, and each demands a different play. Segment-match or forfeit the budget.

Working posture for the current contracting cycle: treat the Camden null as settled absent a scaled replication whose confidence interval excludes 1.0, point the audit team at unbilled TCM claims on mid-band discharges, and hold any top-tier program to the gated-pilot discipline described in the rules below.

The winner is Rising Risk served by discharge-triggered transitional care management — the 99495/99496 pair — modeled at roughly 1.2–2:1 ROI. The mechanism is asymmetric cost bases: a TCM episode costs a few hundred dollars per member-month, so two enrolled months run under a thousand dollars, while each avoided admission carries the HCUP-based costing above. Break-even therefore works out to roughly a five-to-seven percentage-point absolute readmission reduction — preventing about a quarter to a third of expected events against the 18–20% baseline. That headroom is genuine, and the tier's defining feature is a fresh discharge: an observed event that predicts near-term utilization far better than any score.

The loser is Top 1% intensive wraparound, modeled near 0.0–0.1:1. Break-even demands eliminating about one admission per enrollee per year — in the very population whose readmission rate the Camden randomized trial showed to be statistically insensitive to intervention (the adjusted ratio of 0.99 documented above). When plausible absolute reduction rounds to zero, the inequality's right side collapses no matter how concentrated the spend; the belief that cost concentration itself proves addressable savings dies exactly here. The verdict is the time-capped, kill-threshold pilot specified in the decision rule — nothing open-ended.

Evidence streamFigureVerdict for the portfolio
Camden Core Model RCT (Finkelstein et al., NEJM, Jan 2020)62.3% vs 61.7% readmitted; adjusted RR 0.99 (95% CI 0.88–1.11)Top-1% wraparound stays ROI-negative; no scale-up without replicated benefit
Cochrane discharge-planning review (Gonçalves-Bradley et al.)Structured discharge planning, RR ≈ 0.92Fund protocolized transitions — the cheapest proven effect
HRRP penalties, FY2025 (KFF analysis)2,273 hospitals fined; up to 3% of Medicare payments; ~$521MCondition-specific bands (HF, COPD, pneumonia, elective joints, CABG) are where reduction is real
TCM billing auditsBilled on well under 10% of eligible dischargesClose this gap first — highest-evidence, lowest-cost lever, and the winner
NAM 2017, Effective Care for High-Need PatientsSegments: high-cost chronic; catastrophic short-term; end-of-lifeNever buy "expensive" as a single target

The 2026 wrinkle lands mid-table: the CMS-HCC risk model v28 reaches full phase-in this year, compressing risk scores for several conditions and reshuffling which members land in which tier. Rebuild all four strata on 2026-normalized scores before comparing ROI across program years — percentile cutoffs tuned on the old score regime are stale inputs. The failure mode is documented outside claims data: according to the Frontiers-published ORDER-DR evaluation, referable-risk calibration error worsened from 0.049 ± 0.001 on internal APTOS test data to 0.160 ± 0.008 on external Messidor-2, and its core methodological claim is that threshold performance, rank ordering, and calibration must be evaluated as complementary dimensions. Confirm v28 preserves who-sits-above-whom even as it compresses score levels, then execute step 7 of the Patient Population Risk Stratification and Resource Allocation Study verbatim: "Benchmark against industry standards. Compare internal risk stratification and resource allocation metrics with benchmarks to evaluate performance."

The Null That Won't Die — Which Risk Tier Pays Back? Churn

The Four-Tier Payback Table

Two edge cases close the table. Exhaustiveness: the four tiers leave the 81st–94th percentile stretch unassigned, and according to Wikipedia's stratified-sampling entry, strata must be collectively exhaustive and mutually exclusive — every member in exactly one. Add an explicit residual stratum or fold 81st–94th into Stable Low by fiat, or every ROI denominator leaks. Discrimination: according to the Cross-Vendor NHANES LLM Benchmark, cardiovascular-disease risk was the hardest pattern to predict, with F1 spanning only 0.853–0.885 across all evaluated models — score-only tiering is weakest exactly where chronic-condition flags cluster, which is why Rising-Risk entry anchors on the observed discharge rather than the percentile alone. Name tiers the way Mayo names mSMART — Mayo Stratification for Myeloma and Risk-adapted Therapy — so each stratum label states the action it authorizes. This quarter: re-run the inequality on v28-normalized strata, fill all five cells for every row, and let the finished table — not vendor decks — set the 2026 funding split.

Tier (operational definition)Cost per member per monthBest effect size (randomized or matched)Break-even admission avoidance2026 regulatory tailwindVerdict
Top 1% — above 99th percentile; ≥3 admissions/prior yearHighest-cost staffing model in the portfolio; intensive multidisciplinary wraparoundCamden RCT null: adjusted readmission ratio 0.99 (Finkelstein et al.; see "The Null That Won't Die")Eliminate ~1 admission per enrollee per yearNegative — v28 compression shrinks the above-99th-percentile censusGated pilot only, pre-committed kill threshold
High Episodic — 95th–99th percentile; 1–2 admissionsBelow Top 1% staffing intensity; above a TCM episode's costNo RCT isolates this band; matched estimates ride regression to the meanA proportional slice of a thin 1–2-event annual base; closes only near zero costNeutral — band membership churns under v28 rescoringNo standing program; capture via discharge trigger only
Rising Risk — 50th–80th percentile; ≥1 chronic condition; ≥1 recent dischargeA few hundred dollars per episode-month (TCM service cadence)Modeled 1.2–2:1 ROI against an 18–20% baseline readmission probability~5–7 percentage-point absolute reduction (derived: two-month episode under $1,000 ÷ the HCUP admission costing)Positive — v28 compression swells the 50th–80th band; TCM separately billableFund — the only row that clears the inequality
Stable Low — below 50th percentileNear-zero; automated outreach onlyNothing addressable — no meaningful event base to preventMathematically unreachable; low baseline probability caps the savings termNone requiredZero care-management dollars

Read the decision rule as a ranking of ignorance, not a ranking of proof. Discharge-triggered transitional care management for the rising-risk band does not rest on overwhelming evidence; it rests on being the only option whose failures have been measured honestly. That distinction matters, because everything below is a reason the rule could fail in your particular book.

Start with the limits of the evidence. The Camden null covered above came from a single safety-net system serving a high-poverty, clinically chaotic population, and no multi-site randomized trial has overturned it — but none has confirmed top-tier savings either. The asymmetry favors the band; it does not prove the band. Meanwhile, most contrary evidence arrives as vendor case studies reporting completers rather than everyone assigned, which inflates results mechanically. And any pre/post savings claim on a cohort selected for extreme prior spend runs straight into regression to the mean: those members drift back toward average cost on their own, flattering whatever program touched them. This is why the durable pitch — put managers where the money concentrates — keeps failing. Spend concentration is a selection artifact, not evidence of addressable waste.

Variance across cases is the second gap. A band-average return conceals offsetting subgroups: a member discharged home with a caregiver, one sent to a skilled nursing facility, and one with untreated behavioral health needs rarely respond the same way, and the mix shifts by market. Rural plans often lack clinic capacity to complete the post-discharge visit inside the billing window at all; academic centers route discharges through resident clinics where documentation fails audit. The rule prices a portfolio, not a person.

The rule breaks under three identifiable conditions. Operationally: if no clinician can reliably see the member within the TCM visit window, the code never bills and the economics evaporate — and the per-episode revenue at stake is modest, running in the low hundreds of dollars under CMS's physician fee schedule, a figure that moves annually, so verify the current rate rather than borrowing a vendor's math. Measurement-wise: claims-runout lag misassigns percentiles, quietly dropping true rising-risk members out of the band. Governance-wise: the capped top-tier pilot only bites if the comparator group and kill threshold are locked in writing before enrollment; vendors renegotiate metrics otherwise. Override the band boundary only for documented clinical acuity — never for spend rank.

The Four-Tier Payback Table — Which Risk Tier Pays Back? Churn

What the Data Doesn't Tell You

Before funding anything this plan year, run the six stress tests below against your own recent quarters of discharge-destination mixes and TCM billing denials.

A null result constrains; it does not settle. Camden leaves four genuine openings for a defender of the super-utilizer tier, and in 2026 a fifth is widening faster than peer review can track. None of the five manufactures what that tier lacks — positive evidence of savings. Each licenses exactly one response under the decision rule above: a gated pilot with a pre-committed kill threshold, never a standing portfolio line.

The concessions are real. Camden was one city — Newark, New Jersey — inside a single safety-net ecosystem; it ran before COVID scrambled discharge patterns; and it was powered for readmissions, not total cost of care. Variants adding housing-first or addiction-treatment components were never randomized anywhere. That concession buys permission to test, not to deploy: an unrandomized variant is a hypothesis, and a hypothesis gets the capped pilot, not a budget line.

Payer-side attribution is weaker than the trial's. Most published payer "savings" come from pre/post trend comparisons with no matched control, and a cohort chosen because costs spiked improves the next year even untouched — regression to the mean alone can produce double-digit apparent first-year reductions. The oldest slide in the deck — a sliver of members drives half the spend, so that's where care management pays back — mistakes concentration for addressable savings, the exact inference the null above refuted for the top tier. Demand the comparison design in writing before go-live, not after the first quarterly review: propensity-matched controls, or a stepped-wedge rollout whose later waves serve as concurrent controls.

Books differ, too. Medicaid super-utilizers skew toward behavioral health crises and homelessness; Medicare Advantage super-utilizers toward frailty and polypharmacy. Camden's population sits nearer the first profile, so its null transfers to a Medicaid book more plausibly than to an MA book — and neither direction transfers for free. As Wikipedia's stratified-sampling entry notes, when subpopulations vary, sampling each stratum independently "could be advantageous"; grade each book separately instead of pooling a blended ROI that flatters whichever tier happens to be trending well.

Stress testWhat failure looks likeGuardrail
Mean reversionSavings shown pre/post with no control armRequire a matched-comparison or stepped-wedge design
Completer biasDenominator counts only members who finishedContract for intent-to-treat reporting
Capacity shortfallVisits billed late or outside the TCM windowTrack days-to-first-visit weekly during ramp-up
Stale percentilesBand membership frozen between claims cyclesRecompute percentiles at each claims runout
Subgroup maskingOne blended return quoted for the whole bandStratify results by discharge destination
Threshold erosionPilot metrics reopened mid-flightLock comparator and kill threshold pre-enrollment
What the Data Doesn't Tell You — Which Risk Tier Pays Back? Churn

What the Camden Null Can't Tell You About Your 2026

Last come the 2026-specific unknowns no evidence base has absorbed. CMS-0057-F, the Interoperability and Prior Authorization final rule, set the prior-authorization API deadline at January 1, 2026 — now in force — with patient- and provider-access APIs due January 1, 2027 and payer-to-payer exchange in 2028. HHS ended the public health emergency in May 2023, so any evaluation still anchored to pre-emergency baselines describes a population that no longer exists. And v28 reaches full phase-in for Medicare Advantage risk adjustment in 2026, recomposing predicted-cost percentiles beneath every tier table. Each force moves the denominator your ROI fraction divides by faster than peer-reviewed program evaluations update — a program measured across that boundary evaluates two populations and a coding change at once.

The verdict is uniform down the table: the gated pilot beats the open-ended contract in every row. Before renewing any intensive case-management agreement in 2026, get the comparison design, the per-book stratification, and the multi-year window into the contract itself — then enforce the kill threshold already committed above. Absence of disproof is not a business case.

Two candidate investments, one medical director, one plan year. A 25,000-member Medicare Advantage book carries 250 members above the 99th predicted-cost percentile and 3,000 inpatient discharges per year among members between the 50th and 80th percentiles. The tier ranking is argued elsewhere in this guide; what follows is the ledger a CFO actually signs, built on three lines in strict order — committed cost, benefit under the null, sensitivity granted last. Vendor proposals run that order backward, which is how seven-figure commitments survive review.

Concentration of spend is an accounting observation, not a clinical opportunity. The durable myth — that the costliest sliver of members is automatically where care management pays back — fails because a percentile rank describes where costs have already accumulated, while a discharge date describes when an intervention can still change their trajectory. The Camden null covered earlier tested the concentration logic directly on the highest utilizers and found nothing to harvest. The five rules below translate that lesson into enforcement language a medical director can actually apply.

Rule 1 does the heavy lifting because timing carries information a risk score cannot. Predicted-cost percentiles are backward-looking composites; the marginal signal for a preventable readmission comes from recency of utilization, which is why the enrollment clock starts at discharge, not at the quarterly stratification run. CPT 99495 and 99496 give the workflow a billing spine — payment ties to medication reconciliation and a face-to-face visit within days of discharge — so the ROI arithmetic runs on real revenue codes rather than modeled avoided costs. One edge case worth codifying: a member above the 80th percentile who discharges does not graduate into intensive management on the strength of that admission alone. The band boundary holds, because promoting on acuity is precisely how the super-utilizer trap reopens.

Camden gapEdge case it opensGate before believing any ROI claim
Single city (Newark), pre-COVID, readmission-poweredHousing-first or addiction-treatment add-ons, never randomizedRandomize the add-on component; pilot-only status
Pre/post trends, no matched controlsDouble-digit phantom year-one savings from regression to the meanPropensity-matched or stepped-wedge design written into the contract
One book is not anotherMedicaid: behavioral health, homelessness; MA: frailty, polypharmacyPer-book stratified readouts; reject pooled ROI
Panel-size noiseSeveral-readmission annual swings at 250 membersThree-year rolling window minimum
Moving 2026 denominatorsCMS-0057-F APIs (2026–2028), post-PHE baselines, v28 full phase-inRe-baseline percentiles annually; annotate evaluation windows

Rule 4 exists because a threshold set after results arrive is not a threshold — it is a negotiation. Committing before the first enrollment that the program sunsets unless 180-day all-cause readmissions fall by at least three absolute percentage points versus comparison at twelve months removes sponsor discretion exactly where it is most dangerous. Hold the full horizon even when month six looks strong; early gains in this population carry regression-to-the-mean exposure that only the complete period can wash out.

What the Camden Null Can't Tell You About Your 2026 — Which Risk Tier Pays Back? Churn

Worked Case

Rule 5 keeps the portfolio honest across model cycles. Every annual CMS-HCC recalibration moves members across percentile boundaries silently, and a roster built on the prior year's coefficients misprices break-even the moment the new model lands. Before renewing any care-management vendor contract, rebuild every tier roster and re-price each tier's break-even threshold against current fee schedules. The concrete move for the 2026 plan year: write all five rules into the RFP as pass/fail criteria, not scoring preferences — a vendor unwilling to accept a pre-committed kill gate has already told you what its own evidence looks like.

Arm A — intensive case management for the 250 super-utilizers at $1,000 per member per month for 12 months — commits $3.0M before a single outcome is measured. That is the only certain number in the arm, and it is certain in the wrong direction. Under the null consistent with the Camden result covered above, the readmission benefit is zero. Grant the sensitivity case anyway: a 1-percentage-point absolute reduction against roughly one index admission per member per year yields 2.5 avoided stays × $15,000 = $37,500, an ROI near 0.01:1. The generous case still loses seven figures; the point estimate loses the full $3.0M.

Arm B — discharge-triggered transitional care management under CPT 99495/99496, priced at $250 per episode on the 2026 Physician Fee Schedule basis plus two months of chronic care management at $62 per month — runs about $374 per episode all-in. At 40% uptake across the 3,000 eligible discharges, that is 1,200 episodes × ~$374 ≈ $450K committed. The structure differs in kind, not degree: Arm B buys discrete billable episodes triggered by a discharge event, so volume is capped by the discharge count and the only lever is uptake.

The benefit side: assume a 20% baseline readmission rate and a 3-point absolute reduction across the 1,200 treated episodes. That averts 36 stays × $15,000 ≈ $547K — an ROI of roughly 1.2:1 before counting a single star-rating point, HEDIS measure, or dollar of readmission-penalty exposure. Note the asymmetry: Arm A's upside is capped by the null, while Arm B's ledger omits its own tailwinds entirely. That asymmetry, not optimism, is what makes the rising-risk band the only fundable tier in the portfolio.

Close the ledger: the same budget buys either a certain ~$3M loss or a ~$100K surplus with quality-score upside attached. The myth that dies here is the oldest line in payer care management — that spending concentration proves addressability. The top tier concentrates cost; the Camden trial tested whether that concentration converts into avoided admissions and returned a null that has never been credibly overturned at scale. A 2026 CFO who funds Arm A anyway is buying narrative, not return.

Ledger lineArm A: top-1% intensiveArm B: rising-risk TCM
Population served250 members above the 99th percentile1,200 episodes (40% of 3,000 eligible discharges)
Annual commitment$3.0M ($1,000 PMPM × 12 months)≈ $450K (~$374 per episode)
Readmission assumption0 points (Camden null); 1 point as sensitivity3 points absolute on a 20% baseline
Avoided stays per year0 (2.5 in the sensitivity case)36
Benefit at $15,000 per stay$0 ($37,500 sensitivity)≈ $547K
ROI≈ 0.01:1 at best≈ 1.2:1 before quality upside
Funding verdictCertain seven-figure loss~$100K surplus plus star/HEDIS upside

Five Rules for Spending the 2026 Care-Management

Concentration of spend is an accounting observation, not a clinical opportunity. The durable myth — that the costliest sliver of members is automatically where care management pays back — fails because a percentile rank describes where costs have already accumulated, while a discharge date describes when an intervention can still change their trajectory. The Camden null covered earlier tested the concentration logic directly on the highest utilizers and found nothing to harvest. The five rules below translate that lesson into enforcement language a medical director can actually apply.

RuleFailure mode it blocksEnforcement test
1 — Discharge trigger beats risk scorePercentile-only targeting enrolls members whose readmissions are neither imminent nor modifiableEnroll 50th–80th percentile members into TCM within 48 hours of any inpatient discharge
2 — No open-ended enrollmentCensus padding with stable members dilutes per-member ROICap programs above $500 PMPM at 6 months; auto-discharge after 90 days with no admission or ED visit
3 — Matched control or nothingPre/post trends credit seasonality and secular drift to the programAccept only propensity-matched cohorts or stepped-wedge rollouts
4 — Kill gate written firstPost-hoc thresholds turn evaluation into negotiationSunset unless 180-day all-cause readmissions fall at least 3 absolute points versus comparison at 12 months
5 — Re-stratify each model cycleStale HCC rosters misprice tiers after recalibrationRebuild rosters and re-price break-even against current fee schedules before each vendor renewal

Rule 1 does the heavy lifting because timing carries information a risk score cannot. Predicted-cost percentiles are backward-looking composites; the marginal signal for a preventable readmission comes from recency of utilization, which is why the enrollment clock starts at discharge, not at the quarterly stratification run. CPT 99495 and 99496 give the workflow a billing spine — payment ties to medication reconciliation and a face-to-face visit within days of discharge — so the ROI arithmetic runs on real revenue codes rather than modeled avoided costs. One edge case worth codifying: a member above the 80th percentile who discharges does not graduate into intensive management on the strength of that admission alone. The band boundary holds, because promoting on acuity is precisely how the super-utilizer trap reopens.

Rules 2 and 3 police the two cheapest ways a program fakes success. Open-ended enrollment pads the census with stable members whose costs were never going anywhere, so anything priced above $500 per member per month gets a six-month cap with mandatory re-stratification, and anyone quiet for 90 days — no admission, no ED visit — auto-discharges. Measurement discipline matters just as much: according to the CMS Technical Expert Panel summary, CMS declined to finalize its proposed Physician Compare benchmarking methodology after public-comment concerns, committing instead to additional stakeholder work and evaluation of other programs' methodologies. If the federal benchmarking apparatus will not lock a methodology under comment pressure, a vendor's pre/post trend deck deserves categorical rejection regardless of brand; only a propensity-matched cohort or a stepped-wedge rollout separates program effect from a mild flu season.

Rule 4 exists because a threshold set after results arrive is not a threshold — it is a negotiation. Committing before the first enrollment that the program sunsets unless 180-day all-cause readmissions fall by at least three absolute percentage points versus comparison at twelve months removes sponsor discretion exactly where it is most dangerous. Hold the full horizon even when month six looks strong; early gains in this population carry regression-to-the-mean exposure that only the complete period can wash out.

Rule 5 keeps the portfolio honest across model cycles. Every annual CMS-HCC recalibration moves members across percentile boundaries silently, and a roster built on the prior year's coefficients misprices break-even the moment the new model lands. Before renewing any care-management vendor contract, rebuild every tier roster and re-price each tier's break-even threshold against current fee schedules. The concrete move for the 2026 plan year: write all five rules into the RFP as pass/fail criteria, not scoring preferences — a vendor unwilling to accept a pre-committed kill gate has already told you what its own evidence looks like.

What to do next

StepActionWhy it matters
1Build the eligible roster by querying 12 months of medical claims for members between the 50th and 80th predicted-cost percentiles with at least one inpatient stay — this exact band is the funding target.MEPS puts roughly a quarter of national health spending on the top 1%, but the Camden trial showed the costliest members' utilization barely moves; break-even lives in the mid-tier whose utilization is still directionally movable.
2Contract the intervention as discharge-triggered CPT 99495/99496 transitional care management — initiate the nurse or coach outreach off the inpatient discharge event itself, not a referral queue.Transitions coaching and follow-up outreach are the unglamorous mid-tier services that clear break-even, while super-utilizer programs hold the moral high ground and the budget line.
3Cap any proposed top-1% intensive case-management program at a 12-month gated pilot, designed like the Camden Coalition RCT: randomize the region's costliest patients, wrap them in nurses and community health workers at roughly $1,000 per member per month, against usual care.The NEJM 2020 result — 180-day readmissions of 62.3% treated versus 61.7% usual care across 800 patients — is the null your pilot must beat, not narrate around.
4Pre-commit the kill threshold in writing before enrollment opens: specify the readmission or total-cost delta below which the program terminates at month 12, with no extension clause.Camden's near-tie should have rewritten every payer's investment thesis overnight; five years on, super-utilizer budgets survive precisely because kill criteria were never pre-committed.
5Require external-population validation of every risk model used to assign tiers — demand the deployed-calibration figure, not the internal one, replicating the ORDER-DR check that saw ECE balloon from 0.049 on held-out APTOS to 0.160 on 1,744 external Messidor-2 images while QWK fell from 0.8960 to 0.6423.Internal calibration is a poor predictor of deployed calibration; elite discrimination rarely survives a change of population, and a drifting model reallocates care at random instead of allocating it.
6Audit any stratified benchmark or tiered carve-out for its net effect before trusting it — apply the lesson of the ESRD Treatment Choices analysis (JAMA, May 7, 2025; DOI 10.1001/jama.2025.5209), where stratified benchmarks shielded safety-net dialysis facilities yet left overall financial penalties higher than before.Tiered adjustments change who eats the penalty, not the size of the meal; verify the aggregate ledger, then route the savings into the 50th–80th percentile TCM cohort that actually pays back.

Frequently Asked Questions

What did the randomized Camden Coalition trial actually show at 180 days?

At 180 days, 62.3 percent of the intervention group had been readmitted versus 61.7 percent of controls — an adjusted relative risk of 0.99 (95% CI 0.88–1.11) — with no significant difference in readmissions or total costs.

If the top 1% of spenders drive a quarter of national health expenditures, why isn't targeting them automatically profitable?

Because only about a third of people in the top decile of spending are still in that decile the following year, so a roster built on last year's claims is mostly a list of people already reverting toward their own baseline.

How often do providers actually bill Medicare for transitional care management on eligible discharges?

Published billing audits find TCM billed on well under 10 percent of eligible discharges, even though Medicare pays separately for a CPT 99495/99496 fee per qualifying discharge plus a monthly chronic care management fee under CPT 99490.

Is there any head-to-head evidence that light-touch transition work beats intensive super-utilizer wraparound?

Yes — the Cochrane review led by Gonçalves-Bradley found structured, protocolized discharge planning modestly reduces readmissions with a relative risk around 0.92, while the intensive Camden wraparound model returned an adjusted relative risk of 0.99.

How badly does a diabetic retinopathy model's calibration degrade when deployed outside its home dataset?

ORDER-DR's referable-risk expected calibration error was 0.049 ± 0.001 on held-out APTOS test data but ballooned to 0.160 ± 0.008 on 1,744 external Messidor-2 fundus images, with quadratic weighted kappa falling from 0.8960 ± 0.0049 to 0.6423 ± 0.0364.

What savings rate does an intensive top-tier program need just to cover its own delivery costs?

With delivery costs running four figures per member per month at 15–30 caseloads, the break-even bar is roughly one $15,000 admission prevented per enrollee per year.

Quick answers

What did the January 2020 randomized evaluation of the Camden Coalition Core Model find at 180 days?62.3 percent of the intervention group had been readmitted versus 61.7 percent of controls — an adjusted relative risk of 0.99 (95% CI 0.88–1.11) — with no significant difference in readmissions or total costs.
According to AHRQ's Medical Expenditure Panel Survey, what share of national health expenditures do the top 1% of spenders account for, and how many remain in the top decile the following year?The top 1% of spenders account for roughly a quarter of national health expenditures, yet only about a third of people in the top decile of spending are still in that decile the following year.
What did the Cochrane review of discharge planning led by Gonçalves-Bradley and colleagues find?Structured, protocolized discharge planning modestly reduces readmissions, with a relative risk around 0.92.
How did ORDER-DR's referable-risk expected calibration error change between its held-out APTOS data and external Messidor-2 images?It was 0.049 ± 0.001 on held-out APTOS test data but 0.160 ± 0.008 on 1,744 external Messidor-2 fundus images.
Which member band does the payback live in, and what funding mechanism covers it?The payback lives below the 80th percentile, not above the 99th — members between the 50th and 80th predicted-cost percentiles with an inpatient stay in the past 12 months, funded through CPT 99495/99496 transitional care management.

Research Methodology & Editorial Standards

We begin by defining the specific objectives the reader needs to accomplish. Primary product documentation and authoritative secondary sources are assembled into a verified research corpus; drafting occurs only after this foundation is in place.

Every quantitative claim is subjected to dual-source verification. Any figure that cannot be independently corroborated is either qualified or omitted.

Published · Last reviewed · Owned by the Hcco editorial desk (About, Contact, Privacy).

Related answers