Abstract
The Global Burden of Disease (GBD) study and its central metric, the disability-adjusted life year (DALY), are the invisible infrastructure beneath nearly every claim about global health priorities, including the claims that built the Movement for Global Mental Health. This article turns from the diseases GBD ranks to the instrument doing the ranking. I trace the genealogy of DALYs from Christopher Murray and Alan Lopez's original 1990 study for the World Bank through the metric's institutional relocation to the Institute for Health Metrics and Evaluation in Seattle, funded principally by the Gates Foundation. I examine how disability weights, the numbers that let blindness, depression, and a fractured femur be compared on one scale, are derived from population surveys whose respondents are not representative of the populations whose suffering they are used to rank. I compare what this apparatus does across three disease domains, infectious disease, noncommunicable disease, and mental disorder, showing that its fit with the underlying phenomenon varies sharply by domain. I then argue that GBD's additive architecture faces its sharpest challenge not from bad data but from multimorbidity, the co-occurrence of multiple conditions in one body, which the metric's own comorbidity corrections can only approximate by assuming a statistical independence between conditions that the epidemiology of multimorbidity contradicts.
Counting the World's Suffering
In 2019, the Global Burden of Disease study estimated that ischaemic heart disease was the world's leading cause of death, that low back pain was the leading cause of years lived with disability, and that depressive disorders ranked among the top twenty-five causes of overall disease burden across every region of the planet (GBD 2019 Diseases and Injuries Collaborators 2020). These figures circulate with an air of simple fact. A newspaper cites them without qualification. A health ministry builds a national strategy around them. A funding body sets research priorities by them. The Lancet Commission on Global Mental Health and Sustainable Development opens its case for urgent action with exactly this kind of figure: depressive disorders as a leading cause of disability, ranked and quantified with an apparent precision that invites no further question (Patel et al. 2018).
This article asks the question that gets skipped whenever a burden figure is cited: what kind of instrument produces a number that lets blindness, back pain, depression, and a road traffic injury sit on the same scale, ranked against one another as though suffering were a single, fungible substance that only varies in quantity? The Global Burden of Disease (GBD) study and its central metric, the disability-adjusted life year (DALY), are the answer, and they are also the least examined part of the entire global health apparatus. A related analysis has already traced how contested epidemiological claims, uncertain economic modelling, and a biomedicalised notion of treatment gaps combined to build the Movement for Global Mental Health (Ecks 2021). This article goes one level further down, into the metric that made the comparison of mental and physical suffering possible in the first place.
The stakes of getting this right extend well beyond a methodological quibble. That earlier analysis traced one of the central paradoxes global mental health policy keeps producing: that reported rates of depression continue rising in exactly the countries that have spent the most on diagnosing and treating it (Ecks 2021). Part of the explanation for that paradox lies in what gets counted as depression, and how. But part of it also lies in the instrument doing the counting, which has never been built to distinguish a genuine rise in an underlying condition from a rise in the sensitivity, reach, or definitional breadth of the surveillance apparatus measuring it. An instrument this consequential deserves the same scrutiny usually reserved for the diseases it ranks.
The argument proceeds in four moves. I first trace the genealogy of DALYs, from a technical exercise commissioned by the World Bank in the early 1990s to a metric now produced by a single well-funded research institute whose findings shape health policy for most of the world's population. I then open up the black box of disability weighting, the process by which a number gets attached to a health state, and ask whose judgements those numbers actually encode. Third, I compare what the resulting apparatus does across three quite different kinds of disease, showing that its fit with the underlying reality is uneven rather than uniform. Finally, I turn to multimorbidity, the simultaneous presence of several conditions in a single person, which I argue is not a data problem GBD can eventually fix but an architectural limit built into what an additive metric can represent. The article closes, as global health metrics themselves rarely do, in the register of the clinic and the household, where illnesses do not arrive one at a time to be added up but several at once, entangled, and unsequenced by any DALY table.
One further point of framing is worth making before the genealogy begins. DALYs function, in the policy documents that cite them, as something close to a universal solvent: a single unit capable of dissolving the categorical differences between an infectious disease with a known pathogen, a chronic condition with a slow and variable course, and a psychiatric diagnosis resting on symptom report alone, so that all three can be poured into the same comparative table. This solvent property is precisely what makes DALYs so useful to a funder or a ministry that must, in the end, choose where a finite budget goes. It is also precisely what makes the metric worth examining critically, because a solvent that dissolves categorical difference does not eliminate that difference. It only makes it harder to see in the resulting figure.
The Making of an Instrument
The Global Burden of Disease study did not emerge from within epidemiology as a response to an internal disciplinary need. It was commissioned. In 1990, the World Bank contracted Christopher Murray, then a young physician-economist at Harvard, and Alan Lopez, an epidemiologist at the World Health Organization, to produce a comprehensive, comparable accounting of the world's disease burden, one that could be used to set investment priorities across the Bank's health lending portfolio (Murray and Lopez 1996). The resulting study appeared as the technical backbone of the Bank's 1993 World Development Report, Investing in Health, the same report that marked the moment global health policy began speaking of a single, calculable global disease burden rather than a patchwork of locally described health problems (World Bank 1993).
The scale of the original undertaking is easy to underestimate in retrospect. Murray and Lopez's 1990 study assembled estimates for more than one hundred disease and injury categories across eight world regions, drawing on whatever mortality registration, hospital record, and survey data existed for each, filling enormous gaps with modelled estimates where direct data were absent (Murray and Lopez 1996). This was, by any standard, a remarkable act of synthesis, and it answered a genuine and previously unanswerable question: not simply how many people die of a given cause, which vital registration systems could sometimes already show for wealthier countries, but how much healthy life a population loses in total, once premature death and nonfatal disability are combined into one figure. Before DALYs, a health ministry comparing malaria against arthritis had no common unit in which to make the comparison. Malaria killed visibly; arthritis mostly did not. DALYs made the comparison possible for the first time, which is exactly why the metric proved so quickly indispensable to donors and ministries needing to justify where money would go.
The instrument Murray and Lopez built to answer the Bank's question was the disability-adjusted life year. A DALY combines two previously separate ways of counting health loss: years of life lost to premature death, and years lived with disability, weighted by the severity of the disabling condition (Murray 1994). One DALY represents one lost year of healthy life. Summed across a population and across every disease category, DALYs allow an analyst to say, for the first time, that a particular country lost a particular number of healthy years to malaria, and a particular number to depression, and to rank the two against each other on a single scale. This was one of the most consequential metrics ever invented in public health, because it did something no prior mortality statistic could do: it let disability, previously invisible to national accounting because it does not kill, compete on equal terms with death for policy attention and resources.
The metric's early success was inseparable from its institutional backing. A World Bank-commissioned study, adopted soon afterward by the World Health Organization for its own World Health Report 2001, carried the authority of two of the most powerful institutions in global health simultaneously (WHO 2001). That combination of technical ambition and institutional weight is what allowed a single number, the DALY, to become the default currency in which global health priorities are now expressed and contested, a currency whose origins in one bank's investment-planning exercise are rarely mentioned in the reports that now cite it as though it were simply a fact about the world.
It is worth being precise about what was genuinely new here, since health economists had already been combining mortality and morbidity into single figures for years before Murray and Lopez began their work. Quality-adjusted life years, developed within health technology assessment to help wealthy-country health systems decide which treatments to fund, and years of potential life lost, a simpler metric used in some national vital statistics offices, both anticipated pieces of the DALY's logic. What Murray and Lopez added was scale and comparability across an entire planet's worth of disease categories at once, built from the start as a tool for ranking causes of ill health against one another globally rather than for evaluating a single treatment within a single, resource-rich health system. A metric designed for the comparatively narrow task of health technology appraisal in wealthy countries was, in effect, generalised into an instrument for ranking the world's suffering, a considerably larger claim than the one its methodological ancestors had ever made.
Funding Metrics, Setting Priorities
A metric's institutional home shapes what gets measured as surely as its mathematical architecture does, and GBD's institutional home has changed considerably since 1990. In 2007, Christopher Murray left Harvard to found the Institute for Health Metrics and Evaluation (IHME) at the University of Washington, with a founding grant of pattern-setting size from the Bill and Melinda Gates Foundation, renewed and expanded across the subsequent GBD study cycles IHME has produced roughly every few years since (McGoey 2015). This relocation deserves more scrutiny than it typically receives in policy documents that cite GBD figures as neutral scientific output.
A study once commissioned by a multilateral development bank, answerable in however limited a fashion to its member governments, is now produced by a single research institute whose principal funder is a private foundation governed by its own board and accountable, in the end, to no electorate. Linsey McGoey's study of the Gates Foundation's broader role in global health describes a recurring pattern in which the Foundation's funding does not merely support existing research priorities but actively reshapes which questions get asked, which diseases attract sustained measurement infrastructure, and which methodological approaches, favouring standardised, comparable, exportable metrics over locally specific health information systems, come to dominate an entire field (McGoey 2015). GBD, as the single most heavily resourced and most widely cited metric in all of global health, is the clearest possible instance of this pattern rather than an exception to it.
Jeremy Shiffman's work on the politics of global health priority-setting offers a useful frame for what is at stake. Shiffman argues that claims to scientific or technical authority in global health routinely double as claims to institutional power, and that the actors most successful at setting global health agendas are rarely those with the strongest evidence in some abstract sense, but those best positioned institutionally to have their framing of a problem, and their metric for measuring it, accepted as the default (Shiffman 2014). GBD did not simply describe an existing global disease burden that had always been there, waiting to be counted correctly. It created a new, singular, comparable object, the global burden of disease, that had not existed as an object of policy attention in that form before Murray and Lopez built the instrument capable of producing it. Having built that instrument, and having relocated its production to a single, well-resourced, Gates-funded institute, its architects and funders are now positioned to define which health problems count as global priorities for every subsequent funding cycle.
None of this requires attributing bad faith to anyone involved. IHME's staff include some of the most technically accomplished epidemiologists working today, and the Gates Foundation has directed enormous resources toward genuine, measurable improvements in global health outcomes over the past two decades. The point is structural rather than personal: an instrument this consequential, feeding directly into how billions of dollars of donor and domestic health spending gets allocated, is produced by an institution whose governance is considerably narrower than the population whose suffering it purports to rank. That narrowness leaves a trace in every subsequent choice this article goes on to examine, from whose intuitions get surveyed for disability weights to which risk factors receive sustained measurement attention.
Weighing Suffering
The mortality half of the DALY calculation is comparatively straightforward. Deaths can, with reasonable reliability in countries with functioning civil registration systems, be counted and attributed to a cause. The disability half is a different undertaking entirely, and it is here that GBD's most consequential and least examined choices are made. To calculate years lived with disability, every one of the hundreds of health states GBD tracks, from mild anaemia to quadriplegia to major depressive episode, must first be assigned a disability weight: a number between zero, representing full health, and one, representing a state judged equivalent to death, indicating how much healthy life that condition is deemed to subtract.
Where do these numbers come from? Early GBD studies relied on the judgements of small groups of health experts, asked to rank conditions against one another. This approach was widely criticised for encoding the intuitions of a narrow professional class as though they were universal facts about suffering, and subsequent GBD cycles moved toward population surveys instead, using methods borrowed from health economics: paired comparison, in which respondents are shown descriptions of two health states and asked which they judge worse, and variants of the time trade-off and person trade-off techniques, in which respondents are asked to weigh years of life against degrees of disability directly (Salomon et al. 2012). The GBD 2010 disability weights study, still the methodological basis for much of the weighting scheme in use today, combined household surveys conducted in a handful of countries, including Bangladesh, Indonesia, Peru, and Tanzania, with a much larger open-access internet survey conducted in English, Spanish, and a small number of other languages (Salomon et al. 2012).
The internet survey did most of the numerical heavy lifting, simply because it could reach far more respondents than any household survey. But an open, unincentivised internet survey does not reach a representative sample of humanity. It reaches whoever has internet access, whoever speaks one of a small number of survey languages, and whoever chooses to spend their own time answering hypothetical questions about the relative severity of blindness versus infertility versus intellectual disability. Respondents to this kind of instrument have skewed, by every available account, toward the wealthier, more educated, more English-speaking populations for whom internet access and unpaid survey participation were both more available. The resulting disability weights, then applied uniformly to every country in the GBD study regardless of that country's own health beliefs, its own social organisation of disability, or its own relative valuation of, say, blindness against chronic pain, encode a set of judgements disproportionately drawn from populations quite unlike the populations whose disease burden the weights are subsequently used to calculate.
This is not merely a sampling problem correctable by a larger or more diverse survey next time, though GBD has made efforts in that direction across successive study cycles. It is a conceptual problem about what a disability weight is even claiming to represent. The premise behind a universal weighting scheme is that the loss associated with a given health state is, at some level, culturally and contextually invariant, that blindness subtracts a comparable amount of wellbeing whether the person going blind lives in rural Tanzania or suburban Ohio. Amartya Sen raised exactly this difficulty in relation to health metrics more broadly, arguing that interpersonal and cross-cultural comparisons of wellbeing are far more fraught than standard health economics acknowledges, because the same objective health state can be experienced, valued, and adapted to in profoundly different ways depending on what a person and their community have come to expect as normal (Sen 2002). A degree of disability that would be devastating against one set of background expectations may be lived as entirely manageable against another, not because the underlying impairment differs but because what counts as a livable life differs. A universal disability weight cannot represent this variation. It can only average over it, and the average, once calculated, disappears back into a number that looks exactly as precise for Tanzania as it does for Ohio.
The instability this produces is not merely theoretical. Successive rounds of disability weight measurement, conducted for the GBD 2010 study and revised again for later cycles, have produced meaningfully different weights for the same nominal health states, even though the underlying conditions being weighted had not themselves changed (Salomon et al. 2012). Some of this movement reflects genuine methodological improvement. But some of it reflects nothing more than a different, still unrepresentative sample of respondents answering the same hypothetical questions slightly differently, a source of variation with no equivalent in a measurement like a blood pressure reading, which does not change simply because a different, equally untrained observer happens to be holding the cuff. A metric whose central weights shift between study cycles for reasons that have nothing to do with the world it is measuring is not drifting toward greater precision so much as revealing how much of its apparent precision was constructed rather than discovered in the first place.
What the Funnel Does to Different Diseases
A single instrument applied uniformly across every disease category does not produce uniform distortion. Comparing what the GBD apparatus does to three different kinds of disease, infectious disease, noncommunicable disease, and mental disorder, shows that its fit with the underlying phenomenon varies sharply depending on what kind of thing is being measured, and that this variation tracks something more precise than a general, evenly distributed margin of error.
Infectious Disease: The Best Case
Infectious disease is, relatively speaking, GBD's best case. Malaria, tuberculosis, and most vaccine-preventable illnesses have clear, laboratory-confirmable diagnostic criteria: a blood smear, a sputum culture, a polymerase chain reaction test either finds the pathogen or it does not, in a way that leaves comparatively little room for the kind of clinical judgement that governs a psychiatric diagnosis. Decades of dedicated disease control programmes, built originally around single diseases rather than around comparative burden accounting, left behind surveillance infrastructures, national reference laboratories, sentinel reporting sites, that GBD could draw on rather than build from nothing. Cases can be counted with reasonable confidence, at least where these surveillance systems function and are adequately resourced, and prevalence figures, however imperfect at the margins, refer to something with a comparatively stable and externally verifiable existence independent of who is doing the counting.
This is precisely the class of disease GBD's underlying architecture, built around discrete, identifiable, single-cause conditions with an unambiguous case definition, fits best. It is not a coincidence that the disease categories least contested in GBD's output, and least often the subject of methodological critique in the epidemiological literature, are overwhelmingly infectious diseases with laboratory confirmation available. The instrument and the object being measured were, in an important sense, made for each other.
Noncommunicable Disease: Measurement Meets Industry
Noncommunicable disease presents a more mixed picture, and one shaped as much by politics as by epidemiology. Cardiovascular disease and diabetes have reasonably stable diagnostic criteria and can, with adequate health system infrastructure, be measured directly through blood pressure readings, blood glucose tests, lipid panels, and clinical records rather than through self-report or clinical impression alone. In this respect they sit closer to infectious disease than to mental disorder on the measurability spectrum.
But what counts as a risk factor contributing to these diseases, and how much weight that risk factor receives in GBD's parallel accounting of risk-attributable burden, a separate but closely linked strand of the GBD apparatus, is shaped by which industries have the resources and the incentive to contest the underlying science. The sugar, alcohol, and ultra-processed food industries have each, at various points and in various jurisdictions, funded research, lobbied international standard-setting bodies, and disputed methodological choices explicitly aimed at narrowing how their products are counted as contributors to noncommunicable disease burden, a dynamic documented extensively in the broader public health literature on corporate influence over health metrics and inherited in full by GBD's risk factor modelling, even though it is rarely discussed in GBD's own technical documentation (Academy of Medical Sciences 2018). A burden estimate for obesity-related disease, or for alcohol-attributable liver disease, is therefore shaped not only by measurement difficulty in the ordinary epidemiological sense but by whose competing measurement gets funded, published in journals GBD's modellers draw on, and ultimately accepted into the study's enormous input database. The disease itself may be as measurable, in principle, as tuberculosis. The burden attributed to any particular cause of it is considerably less settled.
Mental Disorder: The Hardest Case
Mental disorder is GBD's hardest case, and the reasons run deeper than data quality alone, though data quality is certainly part of the problem, as a detailed examination of the epidemiology underlying the Movement for Global Mental Health has already shown (Ecks 2021). A psychiatric diagnosis has no laboratory confirmation, no biomarker accepted widely enough to serve as a diagnostic gold standard, and depends entirely on a clinician's or a survey instrument's judgement about whether a cluster of reported symptoms crosses a severity and duration threshold that was itself set by professional consensus rather than by any external, independently verifiable measurement. Prevalence estimates for major depressive disorder vary by a factor of more than thirty between the highest- and lowest-scoring countries in cross-national surveys using the same nominal diagnostic instrument, a range that cannot plausibly reflect genuine underlying variation in a single, biologically uniform condition applied consistently across such different settings (Bromet et al. 2011).
GBD's disability weighting compounds this problem rather than correcting for it. A health state as clinically and experientially variable as depression, ranging from a brief, situationally intelligible period of grief or demoralisation to a chronic, incapacitating condition lasting years, is collapsed into a small number of discrete severity categories, mild, moderate, severe, each assigned a single disability weight derived from the same population surveys already shown, in the preceding section, to be unrepresentative of the populations the weights are subsequently applied to. The instrument's core assumption, that the underlying phenomenon can be diagnosed with reasonable consistency across settings and then weighted with a single universal number, is precisely the assumption mental disorder is least able to satisfy.
The comparison across these three domains yields a claim stronger than the general observation that all quantification is imperfect. GBD's fit with reality is not evenly bad, or evenly good, across the diseases it ranks against one another. It tracks, with reasonable precision, how far a given disease category sits from the conditions the metric's underlying architecture was originally built to handle: a discrete, externally verifiable, single-cause condition with an unambiguous case definition and minimal contestation over its risk factors. Infectious disease sits close to that architecture. Mental disorder sits far from it, and noncommunicable disease sits somewhere unevenly in between, depending on how contested its risk factors happen to be in a given jurisdiction. Applying the same numerical apparatus to all three and then ranking the resulting figures against one another on a single global scale, as every GBD report and every policy document built on it necessarily does, treats this variation in fit as though it did not exist, precisely because a DALY, once calculated, carries no visible marker of how confidently it was calculated.
The Multimorbidity Problem
The preceding section identified a problem GBD could, in principle, address through better data: improved surveillance, more representative disability weight surveys, more careful handling of industry-funded risk factor research. This section identifies a different kind of problem, one built into the architecture of the DALY metric itself and not resolvable by better data alone.
DALYs are additive. The entire comparative power of the metric, its ability to say that a country's total disease burden equals the sum of the burden attributable to each disease category, depends on this additivity. Burden from malaria, burden from depression, burden from road traffic injury: add them together, and the total represents the country's overall loss of healthy life. This works cleanly for a hypothetical population in which each person suffers from, at most, one condition at a time. It works considerably less cleanly for actual populations, in which multimorbidity, the co-occurrence of two or more chronic conditions in a single person, is not a marginal complication but close to the norm in older populations everywhere, and increasingly common at younger ages in both high-income and low-and-middle-income countries (Barnett et al. 2012).
The epidemiological literature on multimorbidity, much of it developed independently of and largely uncited by the GBD apparatus, makes an empirical claim of direct relevance here: conditions that co-occur in the same body are frequently not independent of one another. Depression and diabetes co-occur far more often than chance would predict, and each appears to worsen the clinical course and the subjective experience of the other, not merely sit alongside it. Chronic pain, depression, and cardiovascular disease cluster together in patterns that recur across very different health systems and cultural settings, suggesting a shared physiological and social patterning rather than three independent draws from three separate disease lotteries (Academy of Medical Sciences 2018). A patient describing their situation to a clinician, or to an anthropologist, rarely describes three separate conditions requiring three separate additions to a national tally. They describe one entangled, embodied predicament, in which the exhaustion of chronic pain, the hopelessness of depression, and the dietary and mobility constraints of heart disease are not separable inputs but a single, mutually reinforcing condition of being unwell.
GBD is not unaware of this problem, and its methodologists have built a comorbidity correction into the calculation of years lived with disability, precisely to avoid the absurd outcome of a single person's disability weight summing to more than one, which would imply a state worse than death simply by virtue of having several moderate conditions rather than one severe one. But the correction, as described in the study's own methods papers, works by combining the disability weights of co-occurring conditions using a formula that treats their disabling effects as approximately statistically independent, discounting the combined weight below the simple sum of the parts, but without modelling the interaction between conditions that the multimorbidity literature identifies as central to how these conditions are actually lived (Vos et al. 2012). Independence is a mathematically convenient assumption. It is also, for the depression-diabetes pairing, the chronic pain and mental illness cluster, and most of the other multimorbidity patterns documented in the literature, an assumption the evidence directly contradicts. The correction solves GBD's internal arithmetic problem, that disability weights should not exceed one, without solving the substantive problem, that the conditions being combined were never independent draws to begin with.
A simplified illustration makes the mismatch concrete. Suppose a national dataset separately estimates a disability weight of 0.4 for moderate depression and 0.3 for moderate chronic pain. Treated as independent, GBD's comorbidity correction would combine the two multiplicatively rather than by simple addition, yielding a combined weight below 0.7, the naive sum, on the reasoning that a person is not simply worse off by the arithmetic total of two separate afflictions. This is a sensible mathematical safeguard against a real problem, disability weights summing past one. But the multimorbidity literature's central empirical finding is that depression and chronic pain are not two independent draws whose combination merely needs discounting for overlap. Each measurably worsens the clinical course of the other: depression lowers pain thresholds and reduces engagement with pain management, while chronic pain is one of the more robust predictors of subsequent depressive episodes across very different populations (Mercer et al. 2018). The true combined burden of the pairing may plausibly exceed, not fall below, the naive sum of the two separate weights, precisely the opposite of what an independence-based discount assumes. A correction built to prevent overcounting may, for exactly the multimorbidity clusters most common in ageing and economically stressed populations, be undercounting instead.
The deeper issue this raises goes beyond a fixable technical detail. An additive metric assumes that the whole is the sum of the parts, discounted for overlap. Multimorbidity, as it is actually lived, suggests that the whole is frequently something else again: not less than the sum of the parts, nor even a simple discount on that sum, but a qualitatively different condition that emerges from the interaction of several ailments in one embodied, socially situated life. No adjustment to the independence assumption, however sophisticated, can capture this, because the problem is not in the size of the discount but in the premise that disability can be decomposed into separable, addable components in the first place. This is why multimorbidity poses a sharper challenge to GBD than data quality does. Better surveillance can, eventually, produce better prevalence estimates for any single condition. No amount of additional data can make an inherently entangled condition decompose cleanly into separately measurable, independently combinable parts.
A further dimension of multimorbidity that GBD's architecture cannot register at all is treatment burden itself, the cumulative demand that managing several conditions simultaneously places on a person's time, attention, and material resources, independent of the disability directly attributable to any one of the conditions. A person managing diabetes, depression, and hypertension may face multiple clinic visits, multiple medication regimens with their own side effects and interactions, and the ordinary cognitive load of tracking several treatment plans that were designed separately by clinicians who may never have consulted one another. This burden is not disability in the sense GBD measures, since it is generated by the health system's own fragmented response to multimorbidity rather than by the underlying conditions directly. But it is a real and, for many patients, a substantial component of what living with several conditions actually costs, and it disappears entirely from a metric built to sum disease-specific weights rather than to account for what coordinating care across several diseases at once demands of the person doing the coordinating.
Local Arithmetics of Suffering
GBD tables and Lancet Commission reports describe disease burden from an altitude at which individual patients disappear into aggregate statistics. Fieldwork on how people actually manage co-occurring illness offers a view from ground level, and it does not resemble the additive architecture of a DALY table.
In earlier fieldwork among Rural Medical Practitioners in West Bengal, patients rarely arrived describing a single, isolated complaint that could be checked against a single diagnostic category and referred to a single line of a national disease tally (Ecks and Basu 2009, 2014; Ecks 2013). A patient managing joint pain, poor sleep, digestive complaints, and what a psychiatric instrument would likely code as a depressive symptom cluster did not experience these as four separate conditions competing for four separate treatment decisions. They experienced, and described, a single, worsening state of not being well, in which each complaint fed the others: pain disrupted sleep, poor sleep worsened mood, low mood reduced appetite for the physical activity that might have eased the joint pain, and economic precarity, itself often the presenting concern before any bodily symptom was mentioned, ran through and beneath the whole cluster rather than sitting alongside it as a separate risk factor to be tallied.
This pattern is not specific to West Bengal, and a growing body of published research on multimorbidity in low- and middle-income countries documents very similar clustering across widely different settings, from urban South Africa to rural China to periurban India, generally finding that multimorbidity is at least as prevalent in poorer populations as in wealthier ones, and often more so, contrary to an intuition, built into much of the burden-of-disease literature's implicit developmental narrative, that multiple chronic conditions are primarily a rich-country problem of ageing populations (Prince et al. 2015). Where GBD's architecture assumes that a national disease profile can be built up disease by disease and then summed, patients and the practitioners who treat them, formal and informal alike, are working the arithmetic in the opposite direction: starting from an entangled, lived totality and only reluctantly, if ever, decomposing it into the separate diagnostic categories that a health system, an insurance scheme, or a burden-of-disease survey requires before it can be counted at all.
This has practical consequences for exactly the kind of question GBD figures are used to answer: where should scarce health resources go. A treatment gap calculated disease by disease, as MGMH calculates the gap for depression specifically, implicitly assumes that closing the depression gap and closing the diabetes gap are two separable investments with two separable returns (Chisholm et al. 2016). Where the two conditions are, in the patient's actual embodied experience, mutually reinforcing aspects of one predicament, an intervention addressing only one may show a disappointing return precisely because it left the reinforcing loop, and therefore much of the actual suffering, untouched. The RMPs in West Bengal, without any formal training in either psychiatry or multimorbidity epidemiology, arrived independently at a version of this insight: several of them, when asked directly about depression, immediately widened the frame to encompass economic precarity, bodily complaint, and low mood together, treating a search for the single relevant category as somewhat beside the point (Ecks and Basu 2014). A biomedical health system organised around discrete diagnostic codes, and a global metric organised around discrete, addable disease categories, are both poorly equipped to recognise what these practitioners were describing without being asked to.
This convergence between lay practitioners' working knowledge and the formal multimorbidity literature is worth pausing on, because it runs against the usual direction of travel in global health, where formal epidemiology is assumed to discover what informal practice has missed. Here the RMPs, working entirely outside any biomedical training programme, had already built their diagnostic and therapeutic practice around a premise, that bodily complaint, low mood, and economic precarity are usually facets of one condition rather than three separate ones, that a multi-billion-dollar global metrics apparatus has only recently, partially, and still marginally begun to accommodate. Whatever else this suggests, it undercuts any assumption that the direction of learning in global health necessarily runs from formal metrics toward local practice rather than, at least sometimes, the other way around.
Some correctives already exist in embryonic form. A small but growing methodological literature has begun proposing multimorbidity-weighted burden estimates that model clusters of co-occurring conditions directly, rather than combining independently estimated single-disease weights after the fact, and primary care research in both wealthy and lower-income settings has developed patient-reported measures that ask directly about the cumulative, entangled burden of living with several conditions at once rather than scoring each condition separately (Boyd and Fortin 2010; Xu, Mishra, and Jones 2017). None of these approaches has yet been adopted at the scale or with the institutional weight of the standard GBD architecture, and it is not obvious that a metric built for global comparison could ever fully absorb the locally variable, individually particular way multimorbidity clusters actually form. But their existence at least demonstrates that the additive assumption critiqued throughout this article is a choice built into GBD's specific architecture, not an unavoidable feature of quantification as such.
None of this is an argument for abandoning quantification, any more than the critique of MGMH's epidemiology developed elsewhere is an argument for abandoning psychiatric care (Ecks 2021). Counting matters. Resources are finite, and some basis for comparing where they are most needed has to exist. The argument is narrower and, I think, more useful than a blanket objection to metrics: an instrument built to add up discrete, independent, one-at-a-time conditions will keep producing numbers that look authoritative while quietly mismeasuring the growing share of the world's suffering that does not arrive one condition at a time. Multimorbidity is not a gap in GBD's data. It is a gap in what an additive metric can be made to say, however good the data feeding into it become.
The Global Burden of Disease study will continue to rank depression against malaria, back pain against blindness, because policymakers need some basis for prioritisation and GBD is, for all its limits, the most comprehensive attempt available. But every figure it produces carries, invisibly, the history traced in this article: a metric commissioned to answer a development bank's question, weighted by surveys that overrepresent a narrow slice of humanity, applied with uneven fit across radically different kinds of disease, and built on an additive logic that a great many of the people it counts have already, in the way they live and describe their own illness, quietly refused.