Value-based healthcare carries an explicit definition. Porter (2010) states it as the health outcomes achieved per dollar spent, assessed over the full cycle of care; the formulation originates in Porter and Teisberg (2006) and is operationalized as a payment and organizational agenda in Porter and Lee (2013), with the cost term specified by time-driven activity-based costing in Kaplan and Porter (2011).

The definition is a quotient. A quotient of outcome over cost is a measure of efficiency. The properties that distinguish an efficiency measure from a statement of value determine the behavior of everything constructed on top of it, and they are the subject of this note.

The prior definition of the term

Value was already a defined term in the production literature that value-based healthcare draws its organizational vocabulary from. Womack and Jones (1996) define it as determined by the ultimate customer, in terms of a specific product, meeting the customer's requirements, at a specific price and at a specific time; activity not contributing to it is classified as waste.

Three properties of that definition are not preserved under conversion to a ratio.

Specification precedes production. In the lean formulation value is an input to process design: it is stated first, and process content is then assessed against it. In the value-based formulation it is an output, computable only after the care cycle closes and costs are attributed. A quantity available only retrospectively cannot serve as a design constraint.

The unit is a single customer. Value is defined per customer, per product, per occasion. Aggregation across patients requires that outcomes be commensurable, and the outcomes of a single procedure performed on patients with different functional baselines and different objectives are not commensurable in the sense required for a common numerator. Porter (2010) acknowledges this through a three-tier outcome hierarchy, which orders outcome types but does not supply an exchange rate between them.

Cost is not a term. In the lean formulation cost is a consequence of process quality rather than a component of value. Deming (1986) states the relation directionally: quality improvement reduces cost through the removal of rework, scrap and delay, and productivity follows. A ratio does not express a directional relation. It returns the same value under an increase in the numerator and a proportional decrease in the denominator, and therefore cannot distinguish them.

Asymmetry between the terms

Where value is defined as O/C, the measured quantity increases under an increase in O or a decrease in C. The two operations are algebraically equivalent and differ materially in every implementation property.

The cost term is denominated in currency, observed at monthly or shorter intervals, subject to direct administrative control, and measured with small relative error. The outcome term is denominated in heterogeneous units, observed at intervals from 90 days to several years, determined jointly by clinical process, patient characteristics and stochastic factors, and measured with large relative error.

The consequence is a property of the estimator rather than of any organization's intent. For independent terms the relative variance of a ratio is approximately the sum of the relative variances of its terms, so O/C inherits the larger relative error, which is the numerator's. The observable quantity is therefore least responsive to effort where it is most closely coupled to clinical activity and most responsive where it is coupled to administrative activity. Campbell (1979) states the general result: the more a quantitative indicator is used for social decision-making, the more it is subject to corruption pressure, and the more it distorts the process it monitors.

Qualification of the outcome measure

Measurement systems intended for consequential decisions are conventionally qualified before use. The AIAG (2010) reference manual specifies the procedure for industrial measurement: repeatability and reproducibility are estimated, expressed against the tolerance, and converted to a number of distinct categories the system can resolve, with a minimum of five for a system used to make decisions on individual parts. ISO 5725-1 (1994) provides the corresponding trueness and precision framework. A system failing qualification is not used for acceptance.

Risk-adjusted outcome measures were incorporated into payment determination without an equivalent qualification step. Two consequences are computable from the sampling distribution alone.

Sampling spread. For a program performing 50 procedures annually against a true complication rate of 3%, the observed count is binomial. The central 95% interval of the observed rate runs from 0% to 8%; 21.8% of such programs record zero complications in a year and 6.3% record four or more, in the absence of any difference in process. Approximately 600 procedures are required to constrain the same interval to 1.7%–4.5%. Dimick, Welch and Birkmeyer (2004) report the corresponding empirical result for surgical mortality: for most procedures, hospital volumes are insufficient to detect even large differences in mortality rate.

Reliability. Provider profiling measures are characterized by the reliability coefficient R = σ²between / (σ²between + σ²within/n), the proportion of observed between-provider variance attributable to true between-provider variance; 0.7 is the conventional minimum for profiling use, and reliability adjustment by shrinkage toward the mean is the standard remedy (Dimick, Staiger and Birkmeyer, 2010). Taking the same 3% rate and assuming a true between-provider standard deviation of one percentage point:

Annual volume Reliability R Observed spread that is noise
50 0.147 85.3%
100 0.256 74.4%
200 0.407 59.3%
600 0.673 32.7%
2000 0.873 12.7%

Under these assumptions 679 procedures are required for R ≥ 0.70 and 2,619 for R ≥ 0.90. Rankings computed at volumes below the first threshold order sampling error, and they do so with sufficient year-to-year persistence to be interpreted as findings. Lilford and Pronovost (2010) set out the argument against the use of institutional mortality rates for performance judgment on these grounds.

Attribution and adjustment

Shewhart (1931) distinguishes variation from chance causes, inherent to the system, from variation with assignable causes. Deming (1986) derives the management consequence: intervention on common-cause variation treated as special-cause increases variation, and ranking individuals on outcomes their system determines is a specific case. Value-based payment ranks providers on outcome measures and attaches payment to the rank.

Risk adjustment is the acknowledgement of the attribution problem within the method. Its purpose is to model and remove the contribution of patient factors outside provider control (Iezzoni, 2013), which is a statement that a substantial share of observed outcome variation is not attributable to the entity being paid. Two properties follow. First, the adjustment model carries estimation error which is not propagated into the payment determination; a point estimate is treated as a fact. Second, the adjusted quantity is a function of recorded diagnoses, so recorded severity is a lever on measured value that is independent of care delivered. The effect is structural rather than a defect of a particular model, and is the case Campbell (1979) describes.

Correct elements of the diagnosis

The critique above concerns the definition and does not extend to the problem statement. Fee-for-service reimbursement compensates activity without reference to result. The full care cycle is the appropriate unit of analysis and corresponds to the value stream in the production literature: the sequence of steps acting on the object of interest, traced across organizational boundaries. Systematic collection of patient-reported outcomes did not previously exist at scale. The integrated practice unit described in Porter and Lee (2013) is a recognizable statement of flow.

These are consistent with the lean formulation. The divergence occurs at the definition, where the object the patient receives is specified as a quotient with cost in the denominator, and payment is then attached to the quotient.

Practical notes

Specify the outcome with the patient before the episode. The functional result sought is patient-specific and is obtained by elicitation. It is an input to care design rather than a score assigned retrospectively.

Address the process; treat outcome change as the dependent variable. The mechanisms available for outcome improvement are standard work, flow, mistake- proofing and short feedback intervals. A payment attached to an outcome specifies a target without specifying a method.

Establish control before ranking. Plot the provider series against control limits, or plot providers against volume on a funnel plot with exact binomial limits (Spiegelhalter, 2005), before treating a difference as real. At the volumes typical of the unit being profiled, most observed differences fall inside the limits.

Qualify the measure before attaching payment to it. Report the reliability coefficient at the volumes involved, the estimation error of the risk-adjustment model, and the minimum detectable difference. A measure that cannot separate two providers is not a basis for paying them differently.

References

AIAG (2010). Measurement Systems Analysis, 4th Edition. Automotive Industry Action Group.

Campbell, D. T. (1979). Assessing the impact of planned social change. Evaluation and Program Planning 2(1), 67–90.

Deming, W. E. (1986). Out of the Crisis. MIT Center for Advanced Engineering Study.

Dimick, J. B., Staiger, D. O., & Birkmeyer, J. D. (2010). Ranking hospitals on surgical mortality: the importance of reliability adjustment. Health Services Research 45(6 Pt 1), 1614–1629.

Dimick, J. B., Welch, H. G., & Birkmeyer, J. D. (2004). Surgical mortality as an indicator of hospital quality: the problem with small sample size. JAMA 292(7), 847–851.

Iezzoni, L. I. (Ed.) (2013). Risk Adjustment for Measuring Health Care Outcomes, 4th Edition. Health Administration Press.

ISO 5725-1 (1994). Accuracy (trueness and precision) of measurement methods and results — Part 1: General principles and definitions. International Organization for Standardization.

Kaplan, R. S., & Porter, M. E. (2011). How to solve the cost crisis in health care. Harvard Business Review 89(9), September 2011.

Lilford, R., & Pronovost, P. (2010). Using hospital mortality rates to judge hospital performance: a bad idea that just won't go away. BMJ 340, c2016.

Porter, M. E. (2010). What is value in health care? New England Journal of Medicine 363(26), 2477–2481.

Porter, M. E., & Lee, T. H. (2013). The strategy that will fix health care. Harvard Business Review 91(10), October 2013.

Porter, M. E., & Teisberg, E. O. (2006). Redefining Health Care: Creating Value-Based Competition on Results. Harvard Business School Press.

Shewhart, W. A. (1931). Economic Control of Quality of Manufactured Product. D. Van Nostrand Company.

Spiegelhalter, D. J. (2005). Funnel plots for comparing institutional performance. Statistics in Medicine 24(8), 1185–1202.

Womack, J. P., & Jones, D. T. (1996). Lean Thinking: Banish Waste and Create Wealth in Your Corporation. Simon & Schuster.