This cross-sectional study compares service-level utilization patterns between health care professional organizations in full-risk contracts and those in a traditional fee-for-service payment structure.
Key Points
Questions
Which individual services are used more vs less in risk-based contracts compared to fee-for-service contracts, and how might fee-for-service prices be reformed to encourage utilization patterns like those observed in risk-based contracts?
Findings
In this cross-sectional study examining 1271 services examined among Medicare Advantage plans, 74% were used less under full-risk contracts than in fee-for-service contracts, while 26% were used more. A simulated fee schedule reform demonstrated how these utilization patterns could guide changes to the physician fee schedule.
Meaning
Because risk-based contracts do not entail a uniform reduction in service use relative to fee-for-service contracts, service-specific utilization levels in risk-based contracts may serve as a benchmark to guide reforms to fee-for-service prices.
Abstract
Importance
Despite revived interest in reforming the physician fee schedule, it is unclear how policymakers should modify more than 1000 different physician service prices to promote high-value care. One systematic approach could leverage service-level utilization patterns in full-risk contracts, which have not been previously characterized.
Objective
To compare service-level utilization patterns between health care professional organizations in full-risk contracts and those in a traditional fee-for-service (FFS) payment structure, adjusted for detailed patient-level differences, and to simulate an FFS fee schedule reform that would encourage the service patterns observed in risk-based contracts.
Design, Setting, and Participants
This retrospective cross-sectional study used Humana Medicare Advantage claims, encounter, and administrative data from 2015 to 2019 for 585 487 beneficiaries attributed to full-risk contract, with 100% downside risk for total medical spending, and 1 153 455 beneficiaries in fee-for-service contract arrangements. Data were analyzed from September 2025 to June 2026.
Main Outcomes and Measures
The primary outcome was service-specific utilization (count per beneficiary-year), adjusted for beneficiary characteristics. A secondary analysis simulated a modified Medicare FFS schedule based on estimated service-level utilization differences and a supply elasticity drawn from prior literature.
Results
Among 1271 services examined, 940 (74%) had lower adjusted utilization in the risk cohort, and 331 (26%) had higher adjusted utilization. For 700 services (55%), adjusted utilization was at least 10% less in the full-risk cohort. For 199 services (16%), adjusted utilization was at least 10% greater. Low-level office visits and laboratory services were more common in risk-based settings, whereas higher-level office visits, hospital and emergency services, and rehabilitative therapies were less common. A simulated budget-neutral reform to the Medicare fee schedule, calibrated to shift utilization levels toward risk-based benchmarks, yielded a possible 6% increase in reimbursement to general practice physicians.
Conclusion and Relevance
In this cross-sectional study, the association between risk-based contracts and service utilization varied substantially across 1271 individual services. Service-specific utilization levels in risk-based contracts may serve as a benchmark to guide reforms to fee-for-service prices.
Introduction
A key challenge facing policymakers and insurers is the financial incentive under fee-for-service (FFS) payment for health care organizations and clinicians to grow the volume and intensity of health care services, regardless of their value to patients. A variety of tools have been used to manage this misalignment of financial incentives, such as value-based contracting, pay for performance, utilization management (eg, prior authorization), coverage restrictions, and patient cost sharing. These managed care tools have typically been layered onto an FFS framework.
Although the fee schedules underlying FFS payment have long been a subject of health policy research,1 fee-schedule reform has not been a major element of recent managed care efforts.2,3,4 Within the FFS framework, Medicare fee schedules loom large, determining prices paid by government insurers and strongly influencing the negotiated prices paid by private insurers.5 For physicians, whose professional services can be provided in outpatient or inpatient settings, the Medicare physician fee schedule (PFS) has been resource based since its introduction, with relative payment levels for services based on estimates of the relative cost of providing those services and other inputs. Recently, however, policymakers have expressed growing interest in reforming the PFS,6,7,8 echoing long-standing concerns from researchers regarding PFS misvaluation.2,3,9,10,11 Because higher prices tend to increase health care utilization by influencing the incentive to deliver a service, reforming the fee schedule can shape service utilization patterns and the value of care received.12 Despite this interest in fee-schedule reforms, it is unclear how to improve on a cost-based system of assigning relative prices.
Utilization patterns in full-risk contracts, such as global capitation, may offer guidance as to which services should be encouraged or discouraged in FFS reforms. A key aspect of risk-based models that differentiates them from other managed care approaches is that they encourage health care organizations to internalize incentives to defray overall spending; ideally, these models better align broad financial incentives with the value of health production. While risk-based contracts may create incentives to stint on care or select patients with favorable characteristics,13,14 these contracts have proliferated under sustained policymaker support,4 with generally favorable evidence on costs and quality.15,16 The broad incentives in these contracts contrast with more narrowly focused tools of utilization management or pay for performance, which target specific services or quality measures.
There is extensive literature estimating the effects of risk models.17,18,19 However, these studies typically measure utilization or spending changes across broad categories of services and often aim to identify causal impacts of risk models. In contrast, less attention has been paid to quantifying differences in granular service-level practice patterns for similar patients treated in risk-based contracts vs nonrisk contracts in a way that reflects any causal effects of risk models as well as organizational capabilities of health care organizations that voluntarily select to participate in them. Various policy strategies could be informed by these service-level practice pattern differences. When health care organizations internalize incentives to manage aggregate spending and adopt infrastructure to shift health care utilization, the resulting practice patterns may be informative regarding which health care services should be encouraged or discouraged. As policymakers consider reforming the resource-based relative value scale of physician prices, these comparisons of utilization across risk and nonrisk contracts may identify services that are overvalued or undervalued under the current cost-based scale.
In this study, we use data from Humana, which has a large national Medicare Advantage business, to compare service-level utilization patterns between health care organizations in full-risk contracts and those in a traditional FFS payment structure. Humana’s Medicare Advantage business is a uniquely useful setting for measuring differences in FFS and full-risk utilization because of the scale at which risk and FFS arrangements coexist, with otherwise similar product features. These analyses used adjustments for detailed patient-level differences. We hypothesized that utilization patterns would differ substantially across service types, with clinicians in risk-based contracts using fewer therapeutic services and more diagnostic services typically delivered in primary care settings.
Methods
Data Sources and Sample Population
We analyzed administrative enrollment, claims, and encounter data on a 100% sample of Medicare Advantage beneficiaries enrolled from 2015 to 2019 with the health insurer Humana. A unique aspect of the data is the detailed specifications of the payment contract under which a beneficiary received care.
We constructed the study sample by assigning beneficiaries to cohorts based on the contract type of beneficiaries’ attributed primary care clinicians and insurance product for each year of the study period. Beneficiary-years were excluded if beneficiaries changed cohorts during the year, were enrolled in hospice, or had incomplete claims or demographic data (see eTable 1 in Supplement 1 for details on exclusion criteria).
The research protocol was deemed exempt from review by the University of Pennsylvania’s institutional review board because this study involves secondary use of deidentified data. We followed the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) reporting guidelines.
Risk-Based Contracts
The primary analysis compared beneficiaries in 2 cohorts. The first, a nonrisk cohort, consisted of beneficiaries in preferred provider organization (PPO) plans whose primary care clinicians were in FFS payment arrangements without risk sharing. The second, a full-risk cohort, consisted of beneficiaries in health maintenance organization (HMO) plans whose primary care clinicians were in voluntary full-risk contracts that entailed delegation of financial risk for all medical spending, either through global prepaid capitation or retrospective settlement. Retrospective settlement included total cost-of-care arrangements, with organizations participating in 100% upside and downside financial risk within a defined corridor. Of note, other incentive programs like quality bonuses were generally unavailable to health care organizations with full-risk contracts during the study period; the exception was a small number of geographic markets where quality bonuses were universally offered to all health care organizations regardless of risk contract status.
Because risk-based contracts can take many forms, we focused on these cohorts at opposite ends of the spectrum of value-based contracts and restricted the full-risk cohort to HMO enrollees and the nonrisk cohort to PPO enrollees, aligning payment and plan type, to capture the greatest possible utilization differences. A third cohort, referred to as the partial-risk cohort, consisted of beneficiaries enrolled in HMO products whose primary care clinicians were paid FFS but shared in a portion of upside and downside risk. This partial-risk cohort was included in a supplemental comparison (eMethods and eFigure 1 in Supplement 1). We also defined an alternative nonrisk cohort as beneficiaries enrolled in HMOs whose primary care clinicians participated in non-risk FFS arrangements (eMethods and eFigure 3 in Supplement 1). Beneficiaries were attributed to primary care clinicians monthly by Humana as part of the insurer’s panel management, either through beneficiary self-selection or retrospectively based on claim patterns.
Measuring Service-Level Utilization
Utilization data were used to construct service-level utilization rates (service counts per beneficiary per year) for each Healthcare Common Procedure Coding System (HCPCS) code-level service. Data included medical claims and encounter data; encounter data are structured similarly to claims and are submitted by health care organizations under capitated arrangements. Both professional and outpatient facility claims/encounters were used to identify service utilization. For each year, we quantified utilization by service as the count of unique instances received by each beneficiary. We deduplicated claims where multiple bills for the same HCPCS code represented a single service (eMethods in Supplement 1 includes processing rules). We excluded beneficiaries whose health care organizations used delegated claims processing (ie, claims were not processed by Humana) and performed sensitivity tests and supplemental analyses to assess the impact of potentially different encounter completeness (eMethods and eFigure 1 in Supplement 1).
We measured service-level utilization for a broad set of medical services corresponding to most Part B spending. We sorted HCPCS services according to annual spending in the 2015-2019 Medicare Carrier Files, excluding physician-administered drugs. We selected the top codes accounting for 95% of spending, yielding 1271 codes, which corresponded to 81% of volume by service count.
Beneficiary Characteristics
Beneficiary characteristics were used to standardize utilization comparisons between cohorts. Sociodemographic and clinical characteristics for each beneficiary-year were drawn from plan enrollment data and Rx-Risk Index clinical indicators.20
Statistical Analysis
For each service, we estimated differences in utilization across the risk cohorts. The goal was not to estimate the causal impact of risk-based contracts on utilization. Rather, we quantified how practice patterns differed for similar patients across contract types, reflecting both the direct effects of risk-based contracts and the practice patterns of clinicians who select into them. Estimates were adjusted for year and sociodemographic, clinical, and geographic characteristics, including age, sex, race and ethnicity according to Centers for Medicare & Medicaid Services beneficiary coding (categorized as Black, White, or a third category, including American Indian or Alaksa Native, Asian, Hispanic, and not otherwise classified), dual-eligibility status, low-income subsidy eligibility, disability (based on original reason for Medicare eligibility), end-stage kidney disease status, and binary indicators corresponding to the individual Rx-Risk Index conditions.
Because we were estimating differences across a large number of services, we used an augmented inverse probability weighted (AIPW) estimator that is conceptually similar to propensity-based IPW approaches. The AIPW approach is considered doubly robust, in which a first IPW model is adjusted by a second outcome model, and is useful when biases from model specification errors are a concern due to many variables and covariates.21,22 In the first IPW model, utilization counts were weighted by estimated probabilities that each beneficiary was in each cohort, using a gradient-boosted decision-tree machine-learning model considering the covariates previously mentioned. To address residual biases, the second step refined the estimate using a model of expected service utilization conditional on covariates and contract type, again using gradient-boosted decision trees.
We used this approach to estimate expected utilization counts for each service code, as well as the percentage difference relative to nonrisk utilization. We applied a bayesian shrinkage adjustment to the relative utilization estimates, which shrinks estimates toward zero to reflect greater uncertainty about the percentage change for services with lower baseline utilization.
We summarized service-level utilization estimates in several ways. We ranked services by relative and absolute utilization difference between cohorts and used these rankings to visualize utilization differences across services and identify services with the largest differences. We summarized differences by clinical category using Restructured Berenson-Eggers Type-of-Service Classification System (RBCS) categories and subcategories, weighting by nonrisk service counts and, in certain presentations, average Medicare professional reimbursement (allowed amounts from 2015-2019).
Data were analyzed from September 2025 to June 2026. SAS, version 9.4 (SAS Institute), and R, version 4.5 (R Project for Statistical Computing), were used in analyses. See the eMethods in Supplement 1 for additional methodological details.
Illustrating Medicare FFS Schedule Changes
One potential application of these results is reforming FFS payment rates to incentivize utilization patterns similar to those under full-risk contracts. As a proof of concept, we developed a modified Medicare FFS schedule based on estimated service-level utilization differences and a supply elasticity drawn from literature.12 For tractability, this exercise required simplifying assumptions; importantly, we did not account for any cross-price elasticities (ie, the use of one service depending on the price of another). The goal was to generate budget-neutral fee adjustments that would yield changes in Medicare FFS utilization consistent with our estimates of utilization differences associated with full-risk contracts.
The modified FFS schedule was generated in 2 steps. First, we calculated changes to Medicare FFS payment rates for each service using the targeted utilization difference for each service and an elasticity estimate. We made the simplifying assumption that utilization is uniformly elastic to reimbursement changes. Conservatively, we used an elasticity of 1.71, among the higher estimates in prior literature, since higher elasticity implies that smaller payment changes are needed to achieve targeted utilization changes.12,23,24,25 Second, we applied a budget-neutral adjustment so that total spending would be unchanged after accounting for both payment rate and utilization changes. We summarized results by projecting revenue changes for each physician specialty. See the eMethods in Supplement 1 for FFS modification methodology details.
Sensitivity Tests
We performed several sensitivity tests (eMethods and eFigures 1-9 in Supplement 1). We included partial-risk contracts, under which health care organizations share in a portion of upside and downside risk within an FFS framework, to test whether there were incremental utilization differences associated with full-risk relative to partial-risk contracts, or whether estimates based on full-risk contracts may have been biased by utilization data incompleteness (eMethods and eFigure 1 in Supplement 1). We compared AIPW estimates with those obtained using ordinary least squares regression (eMethods and eFigure 2 in Supplement 1). To assess the impact of PPO vs HMO product differences, we used an alternative nonrisk HMO control group (eMethods and eFigure 3 in Supplement 1). To investigate the role of missing primary care clinician attribution, we excluded unengaged beneficiaries (eMethods and eFigure 4 in Supplement 1). To assess the impact of socioeconomic adjustment, we used an alternative specification without controls for dual eligibility and low-income subsidy eligibility (eMethods and eFigure 5 in Supplement 1). Given geographic composition differences between the risk and nonrisk cohorts, we both estimated an alternative specification without geographic controls (eMethods and eFigure 6 in Supplement 1) and estimated the model excluding beneficiaries in the South Atlantic region, which had a disproportionate concentration of full-risk contracts (eMethods and eFigure 7 in Supplement 1). We tested an alternative, simpler approach to deduplicating claims/encounter lines (eMethods and eFigure 8 in Supplement 1). Finally, we tested stand-alone–year models, rather than pooling 2015-2019 data (eMethods and eFigure 9 in Supplement 1). Similar patterns of utilization changes across all sensitivity tests demonstrated the robustness of the findings.
We also included illustrative FFS changes under alternative supply elasticity assumptions (eTable 7 in Supplement 2). These results demonstrate that the magnitude of FFS changes required to achieve targeted utilization is sensitive to the assumed supply elasticity.
Results
The sample population consisted of 2 210 293 beneficiary-years, of which 1 153 455 and 585 487 were in the nonrisk and full-risk cohorts, respectively (the remainder were in the partial-risk cohort). See Table 1 for sample characteristics.
Table 1. Characteristics of Cohorts, 2015-2019a.
| Characteristic | % | |||
|---|---|---|---|---|
| Nonrisk cohort | Full-risk cohort | Two-sided partial-risk cohort | Total | |
| No. of unique beneficiaries | 1 153 455 | 585 487 | 471 351 | 2 151 806 |
| No. of beneficiary-years | 2 883 155 | 1 600 581 | 1 206 065 | 5 689 801 |
| Age, mean (SD), y | 71.2 (9.7) | 72.0 (10.0) | 73.6 (9.6) | 71.9 (9.8) |
| Sex | ||||
| Female | 51.2 | 55.8 | 56.6 | 53.6 |
| Male | 48.8 | 44.2 | 43.4 | 46.4 |
| Race and ethnicityb | ||||
| Black | 12.1 | 16.1 | 18.6 | 14.6 |
| White | 85.0 | 74.2 | 72.6 | 79.3 |
| Other race and ethnicity | 3.0 | 9.7 | 8.8 | 6.1 |
| LIS eligible | 17.2 | 30.1 | 23.4 | 22.2 |
| Dually eligible | 10.3 | 23.0 | 18.2 | 15.6 |
| RxRisk Index | 4.52 | 5.38 | 5.18 | 4.90 |
| Disabilityc | 31.9 | 28.5 | 21.2 | 28.7 |
| ESKD | 0.5 | 0.6 | 0.6 | 0.6 |
| Region | ||||
| East North Central | 15.8 | 7.8 | 4.2 | 11.1 |
| East South Central | 14.3 | 3.1 | 15.3 | 11.4 |
| Middle Atlantic | 3.4 | 0.5 | 0.6 | 1.9 |
| Mountain | 5.5 | 1.6 | 8.2 | 5.0 |
| New England | 0.2 | 0.2 | 0.2 | 0.2 |
| Pacific | 1.0 | 0.9 | 2.4 | 1.3 |
| South Atlantic | 33.6 | 72.4 | 58.7 | 49.8 |
| West North Central | 6.9 | 0.2 | 8.1 | 5.3 |
| West South Central | 19.4 | 13.3 | 2.3 | 14.0 |
| Statistical area | ||||
| Metropolitan | 68.2 | 96.0 | 94.8 | 81.7 |
| Micropolitan | 21.5 | 3.4 | 4.1 | 12.7 |
| Rural | 10.3 | 0.5 | 1.1 | 5.6 |
Abbreviations: LIS, low-income subsidy; ESKD, end-stage kidney disease.
Data are from 2015-2019 Humana Medicare Advantage administrative data. The table presents the number of beneficiary-years and the characteristics of beneficiaries included in the primary model. The denominator for percentages is the number of beneficiary years in each cohort. The nonrisk and full-risk cohorts correspond to the primary comparison. The 2-sided partial-risk cohort was used in sensitivity analyses, rather than the primary analyses. All values reflect the application of the exclusion criteria described in eTable 1 in Supplement 1.
Race and ethnicity is reported according to Centers for Medicare & Medicaid Services beneficiary coding (categorized as Black, White, or a third category, including American Indian or Alaksa Native, Asian, Hispanic, and not otherwise classified).
Defined as having been originally eligible for Medicare for a reason other than age.
Comparisons of service utilization between full-risk and nonrisk cohorts showed substantial variation across individual services (Figure 1). Among the 1271 medical services analyzed, 940 (74%) had lower adjusted utilization in the risk cohort, of which 568 were statistically significantly lower after controlling for multiple tests using the Benjamini-Yekutieli procedure at a 5% false discovery rate. In contrast, 331 services (26%) had higher adjusted utilization, of which 141 were statistically significantly higher.
Figure 1. Line Graph of Risk vs Nonrisk Utilization Differences by Service Code, 2015-2019.

Data are from 2015-2019 Humana Medicare Advantage claim and encounter data. Percentage differences were estimated using bayesian shrinkage, and shaded areas reflect the 95% CIs. Services were ranked separately according to percentage difference and absolute difference (services per person-year). To facilitate data visualization, 8 outlier services with utilization differences higher than 200% were winsorized at 200%.
For many services, relative differences were substantial in magnitude. For 700 services (55%), adjusted utilization was at least 10% less in the full-risk cohort. For 199 services (16%), adjusted utilization was at least 10% greater. Because many services are used infrequently, there were fewer services with large absolute differences; 158 services (12%) differed by more than 1 service per 100 beneficiary-years.
Table 2 summarizes the clinical categories accounting for differences between the full-risk and nonrisk cohorts. Among services with higher utilization in the risk cohort, services in the office evaluation and management RBCS category accounted for 31% of the utilization difference, largely in part through HCPCS code 99213 (a 20-29–minute established patient office visit), which alone counted for 18% of the overall increase. Other large contributors to higher utilization included general laboratory services (24%), radiation oncology (10%), and eye evaluation and management (10%).
Table 2. Shares of Utilization Differences in Risk Cohort by Clinical Classificationa.
| Increases by RBCS subcategory | Share of utilization difference, % | Decreases by RBCS subcategory | Share of utilization difference, % |
|---|---|---|---|
| Office evaluation and management | 31 | Office evaluation and management | 12 |
| General laboratory | 24 | Physical, occupational, and speech therapy | 10 |
| Radiation oncology | 10 | Eye procedures | 7 |
| Eye evaluation and management | 10 | Ambulance | 7 |
| Musculoskeletal procedure | 3 | Skin procedures | 5 |
| Imaging, ultrasonography | 3 | Inpatient evaluation and management | 5 |
| Imaging, radiology | 2 | General laboratory | 5 |
| Vascular procedure | 2 | Emergency department evaluation and management | 4 |
| Gastrointestinal procedure | 2 | Home evaluation and management | 4 |
| Ambulance | 2 | Musculoskeletal procedure | 4 |
| Injection/infusion | 2 | Injection/infusion | 3 |
| Home evaluation and management | 1 | Nursing facility evaluation and management | 3 |
| Anesthesia | 1 | Cardiography | 2 |
| Dialysis | 1 | Imaging, ultrasonography | 2 |
| Cardiovascular procedure | 1 | Molecular testing | 2 |
| Other increases | 7 | Other decreasesb | 25 |
Abbreviation: RBCS, Restructured Berenson-Eggers Type-of-Service Classification System.
Data are from 2015-2019 Humana Medicare Advantage claim and encounter data, and average payments from 5% Medicare Fee-For-Service Limited Data Set carrier claims. Results represent each subcategory’s share of added or reduced utilization associated with the risk vs nonrisk utilization difference, with utilization of each service weighted by that service’s average Medicare fee-for-service payment. Clinical groups are defined using RBCS subcategories. Only the top 15 RBCS subcategories are shown individually; remaining subcategories are grouped as other increases or other decreases. Some RBCS subcategories appear in both Increases and Decreases, reflecting variation in utilization across services within the same subcategory; for office evaluation and management specifically, lower-level visits are used more in the risk cohort, while higher-level visits are used less.
The 5 largest categories within other decreases are nuclear imaging, standard radiography, other organ system procedures, anesthesia, and observation care services.
Services with lower utilization in the risk cohort were more evenly distributed across categories. Services in the office evaluation and management RBCS category made up most of the decreases (12%) here, largely owing to HCPCS code 99214 (longer-duration office visits), which accounted for 7% of the overall decrease. Other categories contributing to lower utilization were physical, occupational, and speech therapy (10%); eye procedures (7%); ambulance (7%); skin procedures (5%); and hospital evaluation and management (5%).
Even for clinically related services, there was often substantial variation in the direction or magnitude of utilization differences between risk and nonrisk cohorts. Figure 2 illustrates this pattern for testing services. Overall, the testing RBCS category had 8% higher utilization in the full-risk cohort relative to the nonrisk cohort. However, among testing services, utilization was lower for specialty test subcategories like pulmonary testing, cardiography, and neurologic testing. Individual services within subcategories also exhibited discordant patterns. For example, within the pulmonary testing subcategory, a comprehensive pulmonary function testing service (HCPCS code 94010) was used 32% more in the risk-based cohort, while a component of pulmonary function testing, the flow-volume loop (HCPCS code 94375), was used 37% less; this example highlights heterogeneity in utilization patterns, even for services that are quite similar. See eTables 2 through 6 in Supplement 1 for results for other highlighted individual services, as well as eTable 7 in Supplement 2 for the complete set of individual service results.
Figure 2. Bar Chart of Risk vs Nonrisk Utilization Differences by Clinical Category, Subcategory, and Service.

Data are from 2015-2019 Humana Medicare Advantage claim and encounter data. Clinical categories and subcategories are defined according to Restructured Berenson-Eggers Type-of-Service Classification System (RBCS) service codes. Category and subcategory differences reflect the aggregation of service-level estimated differences. Each service percentage difference was calculated using estimated utilization differences as the numerator and utilization levels in the nonrisk cohort as the denominator. Due to the large number of individual codes analyzed, aggregated estimates are provided across categories, within the illustrative category Test, which reflects diagnostic test services, and within an illustrative subcategory, Pulmonary. Negative estimates corresponding to bars to the left reflect lower utilization in risk contracts. These illustrative examples were chosen to illustrate both general patterns and also to highlight within-category and within-subcategory heterogeneity underlying the aggregate estimates. DME indicates durable medical equipment; PFT, pulmonary function test.
Figure 3 shows simulated estimates for how an alternative Medicare fee schedule would affect reimbursement across different clinical specialties. This simulation, calibrated to shift utilization levels toward risk-based benchmarks, would yield the greatest payment increases for obstetrics-gynecology (21%), other nonsurgical specialties (14%), and laboratory pathology (14%). Specialties with the largest decreases would be emergency medicine (−9%), nonphysician specialties (−9%), and hospital medicine (−5%). General practice physicians would experience an increase of 6%.
Figure 3. Bar Chart of Changes to Medicare Reimbursement Under an Alternative Physician Fee Schedule by Clinician Specialty.

Data are from 2015-2019 Humana Medicare Advantage claim and encounter data, the 2018 Medicare Physician Fee Schedule, and specialty-level service volume from the 5% Medicare Fee-For-Service Limited Data Set carrier claims. Each data point represents the simulated change in Medicare Part B revenue for each clinician specialty following implementation of a simulated, budget-neutral alternative fee schedule. The alternative fee schedule increases prices for services used more in risk-based contracts and reduces prices for services used less in risk-based contracts. See the eMethods in Supplement 1 for methodological details regarding the construction of the alternative fee schedule.
Discussion
In contrast to FFS payment, risk-based contracts provide incentives to reduce aggregate spending, ostensibly better aligning spending with health value of services. However, it has not been clear which specific services are used more or less in risk-based contracts. In this study of a large Medicare Advantage insurer, we observed substantial variation in relative utilization rates across the spectrum of services. While risk-based contracts were associated with lower utilization for most services, utilization increased for a large minority of services. Different clinical types of services accounted for lower utilization (eg, higher-level office visits, hospital or emergency room–based services, nonphysician services like ambulance, physical therapy, and home nursing) and higher utilization (eg, shorter-duration office visits and laboratory services). Services with higher utilization were generally consistent with the hypothesis that health care organizations under risk-based contracts provide more preventive and diagnostic services typically provided in primary care settings.
This work complements prior studies examining the role of risk-based contracts in shaping overall utilization and the use of specific services thought to be low value.26,27 Unlike those studies, however, the primary purpose of this analysis was not to inform policy regarding whether and how to expand risk-based contracts. Rather, our aim was to use practice patterns under risk-based contracts to inform FFS reform and other managed-care strategies broadly by highlighting which specific services might be overused or underused in the FFS setting. Although the present work represents an initial proof of concept, we believe that there are several possible implications of these findings.
First, these results provide uniquely granular evidence that value-based contracts do not entail a uniform reduction in service use. The wide variation in the direction and magnitude of service utilization observed in risk-based contracts, relative to FFS contracts, suggests that coarse policies to reduce utilization (eg, uniform cost sharing) are unlikely to produce the care patterns observed in value-based contracts.
Second, results from this analysis could be used as a benchmark. A simple tool could quantify service-level differences between observed utilization patterns for a group of Medicare beneficiaries and the utilization patterns that would be expected for similar beneficiaries under value-based care models. Such benchmarks could be useful for a broad set of managed-care stakeholders interested in identifying potential areas of medical overuse or underuse, such as researchers, policymakers, or insurers.
Third, results can inform physician fee schedule reforms. With important assumptions, we show how utilization comparisons between risk-based and FFS contracts translate into price reforms that would use service-level prices to shift utilization toward risk-based utilization benchmarks. This methodology, which is independent of historically used cost-based approaches, could be refined by further research quantifying service-level utilization responses to a service’s own price and to the price of related services. This preliminary fee schedule is not intended as a PFS alternative to be adopted outright; rather, it aims to illustrate how benchmarks based on clinicians’ service provision under risk-based contracts can be used to inform fee schedules.
Finally, this analysis demonstrates the feasibility of service-level analysis in payment policy evaluation. Rather than grouping services into coarse categories, we examined 1271 individual services, made possible by advances in machine learning.
Although there are drawbacks to fee schedule reforms, they can complement other efforts to promote high-value care and may offer advantages over other policy efforts. First, they build on an extensively used payment infrastructure that already influences clinical practice broadly. Second, unlike managed care initiatives requiring health care organizations to opt in, fee schedule reforms would apply universally in Medicare and would likely extend to commercial insurance as well.5 Finally, in contrast to some utilization management programs, fee schedule changes avoid administrative barriers to care while preserving the central role of physician and patient decision-making.
Limitations
This study has limitations. First, it is not certain that utilization patterns observed in full-risk contracts are superior to those under FFS. Although empirical research has generally found that risk-based contracts reduce utilization without harming patients, these changes may not be uniformly cost effective or harmless. Because this study is premised on the goal of aligning utilization with patterns observed under risk-based contracts, this premise deserves further investigation.15
Second, differential incompleteness in utilization data between risk-based and nonrisk cohorts could bias comparisons. However, we obtained similar results in a sensitivity analysis that included partial-risk contracts, which are likely to have similar data completeness as nonrisk contracts (eMethods and eFigure 1 in Supplement 1). Relatedly, categorization of health care organizations into risk cohorts was based only on Humana contracts, as we lacked data on contracts with other payers. However, we would expect to observe larger utilization differences when comparing organizations with uniform contract types across payers.
Third, simulated fee schedule modifications rely on simplifying assumptions regarding how utilization responds to price changes. Because utilization responsiveness likely varies across services and price-change magnitudes, the present simulations should be interpreted as a proof of concept, rather than a prescription for fee reforms. Nonetheless, the simulated fee schedule provides a useful illustration of payment rates that may align utilization with risk-based clinician practice patterns, reflecting clinicians’ perceived value of services.
Fourth, results from the study sample period may differ from present day practice patterns. However, we found marked persistence year to year (eMethods and eFigure 9 in Supplement 1) and selected this time period to avoid dynamic utilization changes during and after the COVID-19 pandemic based on data availability. Furthermore, we expect findings to extrapolate well given that relative value units do not change drastically year to year. Finally, although these results were not very sensitive to adjustment for patient characteristics (eFigures 5 and 6 in Supplement 1), it is possible that unmeasured patient characteristics were a source of confounding.
Conclusions
This cross-sectional study documents service-level utilization differences between full-risk and FFS health care organization payment contracts and demonstrates how these differences might be informative for reforming the physician fee schedule. Reforming the physician fee schedule represents an opportunity to leverage a rarely used tool for managed care and cost control in health care.
eMethods
eTable 1. Exclusion Criteria Waterfall
eTable 2. Top 20 Services with Higher Utilization in Risk (Services per Person-Year)
eTable 3. Top 20 Services with Lower Utilization in Risk (Services per Person-Year)
eTable 4. Top 20 Services with Higher Utilization in Risk (by Percentage)
eTable 5. Top 20 Services with Lower Utilization in Risk (by Percentage)
eTable 6. Service-Level Results for Potentially Low Value Services
eTable 7. Service-Level Results for All Services (See Supplement 2)
eFigure 1. Sensitivity Test Results: Comparison of Primary Results and Results Using an Alternative Two-Sided-Risk Control Group
eFigure 2. Sensitivity Test Results: Comparison of Primary Results and Results from an Alternative OLS Model
eFigure 3. Sensitivity Test Results: Comparison of Primary Results and Results Using an Alternative PPO and HMO Comparison Group
eFigure 4. Sensitivity Test Results: Comparison of Primary Results and Results Using an Alternative Non-Risk Control Group Excluding Unengaged Beneficiaries
eFigure 5. Sensitivity Test Results: Comparison of Primary Results and Results from an Alternative Model without Socioeconomic Adjustments
eFigure 6. Sensitivity Test Results: Comparison of Primary Results and Results from an Alternative Model without Geographic Adjustment
eFigure 7. Sensitivity Test Results: Comparison of Primary Results and Results with South Atlantic Region Excluded
eFigure 8. Sensitivity Test Results: Comparison of Primary Results and Results Using an Alternative Deduplication Method
eFigure 9. Sensitivity Test Results: Comparison of Primary Results and Results Using Standalone Years
eTable 7. Service-Level Results for All Services
Data Sharing Statement
References
- 1.Newhouse JP. Pricing the Priceless: a Health Care Conundrum. MIT Press; 2002. doi: 10.7551/mitpress/5534.001.0001 [DOI] [Google Scholar]
- 2.Berenson RA, Ginsburg PB. Improving the Medicare Physician Fee Schedule: make it part of value-based payment. Health Aff (Millwood). 2019;38(2):246-252. doi: 10.1377/hlthaff.2018.05411 [DOI] [PubMed] [Google Scholar]
- 3.Berenson RA, Hayes KJ. The road to value can’t be paved with a broken Medicare Physician Fee Schedule. Health Aff (Millwood). 2024;43(7):950-958. doi: 10.1377/hlthaff.2024.00299 [DOI] [PubMed] [Google Scholar]
- 4.APM measurement: progress of alternative payment models: 2024 methodology and results report. Health Care Payment Learning & Action Network . Accessed January 22, 2026. https://hcp-lan.org/wp-content/uploads/2025/08/2024-HCPLAN-Methodology-Report.pdf
- 5.Clemens J, Gottlieb JD. In the shadow of a giant: Medicare’s influence on private physician payments. J Polit Econ. 2017;125(1):1-39. doi: 10.1086/689772 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Whitehouse and Cassidy introduce legislation, release RFI on primary care provider payment reform. News release. Sheldon Whitehouse website . May 15, 2024. Accessed January 22, 2026. https://www.whitehouse.senate.gov/news/release/whitehouse-and-cassidy-introduce-legislation-release-rfi-on-primary-care-provider-payment-reform/
- 7.Diamond D. RFK Jr. weighs major changes to how Medicare pays physicians. The Washington Post . November 21, 2024. Accessed January 22, 2026. https://www.washingtonpost.com/health/2024/11/21/rfk-physician-payments/
- 8.Miller-Meeks’ Physician Fee Schedule Update and Improvements Act passes out of Energy and Commerce Committee. News release. Mariannette Miller-Meeks website . December 6, 2023. Accessed January 22, 2026. https://millermeeks.house.gov/media/press-releases/miller-meeks-physician-fee-schedule-update-and-improvements-act-passes-out
- 9.Skopec L, Berenson RA. Why the Medicare Physician Fee Schedule misvalues fee levels and how to fix it. Health Aff Sch. 2025;3(10):qxaf189. doi: 10.1093/haschl/qxaf189 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Chapter 1: reforming Physician Fee Schedule updates and improving the accuracy of relative payment rates (June 2025 report). The Medicare Payment Advisory Commission . June 12, 2025. Accessed June 30, 2026. https://www.medpac.gov/document/chapter-1-reforming-physician-fee-schedule-updates-and-improving-the-accuracy-of-relative-payment-rates-june-2025-report/
- 11.Chapter 3: rebalancing Medicare’s Physician Fee Schedule toward ambulatory evaluation and management services (June 2018 report). The Medicare Payment Advisory Commission . June 1, 2018. Accessed June 30, 2026. https://www.medpac.gov/document/http-www-medpac-gov-docs-default-source-reports-jun18_ch3_medpacreport_sec-pdf/
- 12.Clemens J, Gottlieb JD. Do physicians’ financial incentives affect medical treatment and patient health? Am Econ Rev. 2014;104(4):1320-1349. doi: 10.1257/aer.104.4.1320 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Ellis RP, McGuire TG. Provider behavior under prospective reimbursement: cost sharing and supply. J Health Econ. 1986;5(2):129-151. doi: 10.1016/0167-6296(86)90002-0 [DOI] [PubMed] [Google Scholar]
- 14.Frank RG, Glazer J, McGuire TG. Measuring adverse selection in managed health care. J Health Econ. 2000;19(6):829-854. doi: 10.1016/S0167-6296(00)00059-X [DOI] [PubMed] [Google Scholar]
- 15.Song Z. Taking account of accountable care. Health Serv Res. 2021;56(4):573-577. doi: 10.1111/1475-6773.13689 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Leao DLL, Cremers HP, van Veghel D, Pavlova M, Groot W. The impact of value-based payment models for networks of care and transmural care: a systematic literature review. Appl Health Econ Health Policy. 2023;21(3):441-466. doi: 10.1007/s40258-023-00790-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Muhlestein DB, Morrison SQ, Saunders RS, Bleser WK, McClellan MB, Winfield LD. Medicare accountable care spending patterns: shifting expenditures associated with savings. Am J Managed Care. February 12, 2018. Accessed April 26, 2026. https://www.ajmc.com/view/medicare-accountable-care-spending-patterns-shifting-expenditures-associated-with-savings
- 18.McWilliams JM, Gilstrap LG, Stevenson DG, Chernew ME, Huskamp HA, Grabowski DC. Changes in postacute care in the Medicare Shared Savings Program. JAMA Intern Med. 2017;177(4):518-526. doi: 10.1001/jamainternmed.2016.9115 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.McWilliams JM, Hatfield LA, Chernew ME, Landon BE, Schwartz AL. Early performance of accountable care organizations in Medicare. N Engl J Med. 2016;374(24):2357-2366. doi: 10.1056/NEJMsa1600142 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Widagdo I, Kerr M, Kalisch Ellett L, et al. Validity of the updated Rx-Risk Index as a disease identification and risk-adjustment tool for use in observational health studies. Clin Interv Aging. 2025;20:309-323. doi: 10.2147/CIA.S494145 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Robins JM, Rotnitzky A, Zhao LP. Estimation of regression coefficients when some regressors are not always observed. J Am Stat Assoc. 1994;89(427):846-866. doi: 10.1080/01621459.1994.10476818 [DOI] [Google Scholar]
- 22.Chernozhukov V, Chetverikov D, Demirer M, et al. Double/debiased machine learning for treatment and structural parameters. Econom J. 2018;21(1):C1-C68. doi: 10.1111/ectj.12097 [DOI] [Google Scholar]
- 23.Gottlieb JD, Polyakova M, Rinz K, Shiplett H, Udalova V. The earnings and labor supply of U.S. physicians. Q J Econ. 2025;140(2):1243-1298. doi: 10.1093/qje/qjaf001 [DOI] [Google Scholar]
- 24.Alexander D, Schnell M. The impacts of physician payments on patient access, use, and health. Am Econ J Appl Econ. 2024;16(3):142-177. doi: 10.1257/app.20210227 [DOI] [Google Scholar]
- 25.Cabral M, Carey C, Miller S. The impact of provider payments on health care utilization of low-income individuals: evidence from Medicare and Medicaid. Am Econ J Econ Policy. 2025;17(1):106-143. doi: 10.1257/pol.20220775 [DOI] [Google Scholar]
- 26.Boudreau E, Schwartz R, Schwartz AL, et al. Comparison of low-value services among Medicare Advantage and traditional Medicare beneficiaries. JAMA Health Forum. 2022;3(9):e222935. doi: 10.1001/jamahealthforum.2022.2935 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Schwartz AL, Kim S, Chhatre S, et al. Changes in health care utilization and low-value service use after risk-based contract adoption in Medicare Advantage. JAMA Intern Med. 2026;186(1):98-107. doi: 10.1001/jamainternmed.2025.5917 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
eMethods
eTable 1. Exclusion Criteria Waterfall
eTable 2. Top 20 Services with Higher Utilization in Risk (Services per Person-Year)
eTable 3. Top 20 Services with Lower Utilization in Risk (Services per Person-Year)
eTable 4. Top 20 Services with Higher Utilization in Risk (by Percentage)
eTable 5. Top 20 Services with Lower Utilization in Risk (by Percentage)
eTable 6. Service-Level Results for Potentially Low Value Services
eTable 7. Service-Level Results for All Services (See Supplement 2)
eFigure 1. Sensitivity Test Results: Comparison of Primary Results and Results Using an Alternative Two-Sided-Risk Control Group
eFigure 2. Sensitivity Test Results: Comparison of Primary Results and Results from an Alternative OLS Model
eFigure 3. Sensitivity Test Results: Comparison of Primary Results and Results Using an Alternative PPO and HMO Comparison Group
eFigure 4. Sensitivity Test Results: Comparison of Primary Results and Results Using an Alternative Non-Risk Control Group Excluding Unengaged Beneficiaries
eFigure 5. Sensitivity Test Results: Comparison of Primary Results and Results from an Alternative Model without Socioeconomic Adjustments
eFigure 6. Sensitivity Test Results: Comparison of Primary Results and Results from an Alternative Model without Geographic Adjustment
eFigure 7. Sensitivity Test Results: Comparison of Primary Results and Results with South Atlantic Region Excluded
eFigure 8. Sensitivity Test Results: Comparison of Primary Results and Results Using an Alternative Deduplication Method
eFigure 9. Sensitivity Test Results: Comparison of Primary Results and Results Using Standalone Years
eTable 7. Service-Level Results for All Services
Data Sharing Statement
