Abstract
During the nineteenth century, European states and scientific institutions accumulated unprecedented quantities of numerical information about births, deaths, crimes, marriages, illnesses and occupations. Ian Hacking showed how this flood of printed numbers transformed both probability theory and prevailing conceptions of society, making statistical regularities visible in domains where individual events had always appeared irreducibly contingent. This article argues that the ontological transformation underlying that history has remained insufficiently described. Population statistics does not mathematically analyse recursive living processes directly. Before any statistical operation can occur, selected aspects of those processes must be observed, classified, recorded and deposited into durable symbolic form. The resulting datum is nonrecursive: once recorded, it no longer changes in response to the life from which it was produced. Statistical datasets therefore consist not of recursive lives but of accumulated nonrecursive multisymbolic traces of past recursive processes. Large-number regularities arise within these accumulated deposits, and their stability must not be mistaken for an emergent nonrecursivity of the population itself, which remains recursive throughout. Statistical inference rests upon a temporal and ontological transformation from living process to symbolic trace, a transformation with its own conditions, costs and characteristic errors. Population statistics is best understood, on the account developed here, as the mathematical analysis of accumulated nonrecursive symbolic sediments of recursive life.
1. Hacking's Puzzle: How Did Society Become Statistical?
1.1 Begin with the astonishment
Nineteenth-century observers confronted a striking phenomenon in exactly the domains of life that seemed least suited to mathematical regularity. Suicide, crime, marriage, birth, death and illness are, each of them, acts and events bound up with the most private and unpredictable dimensions of a person's existence. No one could say in advance whether a particular individual would take their own life in the coming year. Yet annual suicide rates, tabulated across a city or a nation, could remain remarkably stable from one year to the next. No one could say which particular individuals would commit which particular crimes. Yet the aggregate number and kinds of recorded offences exhibited a regularity that seemed to defy the sheer unpredictability of the individual acts composing it. This is the astonishment around which the historian and philosopher Ian Hacking organised his landmark account of the emergence of statistical thinking in the nineteenth century, The Taming of Chance.
1.2 Hacking's avalanche of numbers
Hacking's account deserves a generous reading before this article introduces its own reframing of it. His central historical claim is not that statistical regularity became visible because someone invented a superior branch of mathematics. It is that the regularity became visible because states and institutions first acquired an unprecedented capacity to count. Civil registration systems recorded births, deaths and marriages as a matter of administrative routine. Censuses counted entire populations at regular intervals. Police forces logged offences according to standardised categories. Hospitals and asylums kept medical records organised around emerging diagnostic classifications. Bureaucracies multiplied the categories through which a life could be entered into a record, and printed reports circulated the resulting tables to an expanding public of officials, reformers and scientists. The nineteenth century generated an entirely new infrastructure through which human events, previously known only locally and anecdotally, became numerically comparable across regions, across institutions and across years. This expansion of counting capacity, which Hacking memorably described as an avalanche of printed numbers, is one of his deepest and most durable contributions to the history and philosophy of science.
1.3 Statistical law
The conceptual novelty this avalanche of numbers eventually produced deserves to be stated with some precision. Classical conceptions of natural law had generally meant deterministic law: given a set of initial conditions, a law of nature determined a single, necessary outcome, and any departure from that outcome indicated either an error of measurement or an incompleteness in the law itself. Statistical law introduced an entirely different relation between the particular and the general. A statistical law permits individual events to remain genuinely uncertain, unpredictable case by case, while the aggregate frequency with which those events occur across a large collection of cases exhibits a stable and mathematically describable pattern. This was a transformation not merely of statistical method but of what nature, society and normality could even be taken to mean, since it became possible, for the first time, to speak of lawlike regularity at the level of the aggregate without any corresponding claim to lawlike determination at the level of the individual case.
1.4 What exactly became statistical?
A question follows immediately from this history, and it is the question this article exists to answer properly. What exactly became statistical over the course of the nineteenth century? Did society itself undergo some kind of transformation? Did human behaviour become, in some sense, less recursive than it had previously been, more determined, more machine-like, more available to calculation because something about human beings themselves had changed? Did populations, considered collectively, acquire a new emergent ontology unavailable to any of the individuals composing them?
The answer to all three questions is no. The people counted in nineteenth-century tables of suicide, crime and marriage remained exactly what people had always been: living, responsive beings whose next action was shaped by their own prior states, by their interpretation of what was happening around them, and by the responses of others. Nothing about human recursivity diminished during the period Hacking studies. What changed, and changed on a scale without historical precedent, was the production of an entirely new symbolic layer sitting between recursive human life and the mathematics eventually applied to it. The decisive historical event was not simply the discovery that large populations display regularities. It was the systematic sedimentation of living processes into standardised, durable, nonrecursive traces, on a scale and with an institutional apparatus no earlier period had ever possessed. That reframing is the burden of the rest of this article.
2. The Ontological Distinction Hacking Needs
2.1 Recursive processes
Making this reframing precise requires only a small amount of conceptual apparatus, introduced here in the minimal form the argument needs. A process is recursive when its subsequent course is affected by how it is engaged: when, in other words, it answers back. A human being engaged in ordinary social life is recursive in an unusually rich sense, responding at once to their own previous condition, to their interpretation of what has just happened to them, to the behaviour of the people around them, to those people's responses to their own behaviour in turn, to classifications and expectations imposed upon them by institutions, and to predictions made about them that they may come to know. Living social life, examined this closely, turns out to be saturated with responsiveness at every level, a being's coordination with its own ongoing state folding continuously into its coordination with the responses of others.
2.2 Nonrecursive symbolic deposits
Human beings also possess a capacity, without equivalent anywhere else in the rest of the animal kingdom in anything like the same degree, to detach a distinction from the occasion of its drawing and deposit it into a durable symbolic form that persists across absence, delay, distance and death. Call this capacity multisymbolization. A written answer on a survey form, a recorded age, a diagnosis code entered into a medical file, a death certificate, a crime category logged in a police report, an entry made during a census, are all instances of exactly this capacity at work. Once produced, each of these is nonrecursive. The figure '42' written in an age column does not subsequently age. The response 'strongly agree' circled on a questionnaire does not revise itself when the respondent later changes their mind. The word 'widowed' entered under marital status does not alter itself when the respondent remarries the following year. This point cannot be stated too plainly, because the entire argument of this article depends on holding it firmly in view: a recorded distinction, once deposited, behaves nothing like the living process it was drawn from.
2.3 The datum and the person belong to different ontological types
These two observations together yield the central proposition of this section, and one of the central propositions of the article as a whole. A statistical datum is not a miniature piece of a recursive process, preserved intact and simply made smaller. It is a nonrecursive symbolic deposit, produced through an act of engagement with that process but ontologically distinct from it. This claim is considerably stronger than the more familiar observation that data are representations of reality, because 'representation' leaves the underlying ontology comfortably vague, suggesting a picture or a copy that stands in some loose correspondence with its object while remaining silent about what kind of thing the copy actually is. 'Deposit' identifies the transformation precisely. A deposit is what a recursive process leaves behind once an act of measurement, classification or recording has occurred, and what it leaves behind belongs to a different ontological category from the process that produced it.
2.4 Statistics begins after deposition
No statistical calculation can begin until a trace of this kind already exists, and the point deserves to be stated as a strict temporal sequence rather than left as a general impression. Before an answer is given, in the moment a question is being formulated and understood, the process is recursive throughout. During the encounter in which a survey is administered, an interview conducted, or an examination performed, the process remains recursive throughout, since interviewer and respondent, or clinician and patient, continue to shape one another's next move for as long as the encounter lasts. Only once a response has been coded, entered, or otherwise fixed into symbolic form does the process cross into nonrecursion. Statistics begins at that third moment, and not a moment before it. This temporal architecture, recursive encounter followed by nonrecursive deposit followed by mathematical operation, is the basic structure the rest of this article develops in progressively greater detail.
3. The Temporal Ontology of the Datum
3.1 Every datum is already past
An implication of the sequence just established is worth drawing out fully, because it runs against a common and largely unexamined assumption about what data are. Statistics never has direct access to a living present. Even data described as real-time are, on inspection, already traces of something that has occurred. A heart-rate monitor reports a bodily event that has already taken place by the time its reading appears on a screen, however brief the interval. A stock price tick records a transaction that has already been completed. A GPS coordinate records the outcome of a measurement already performed. A survey response records an answer that has already been given. The delay separating the living event from its symbolic registration may be measured in milliseconds rather than months, but the ontological sequence is exactly the same in both cases: recursive event first, nonrecursive trace second. Data, considered as such, are constitutively retrospective, however immediate their production may feel.
3.2 The recursive present cannot be frozen
Consider a concrete case. A census enumerator records a person at ten minutes past ten in the morning: age thirty-seven, occupation teacher, marital status married, residence Edinburgh. The living person whose details have just been entered does not pause at that moment. They continue thinking, moving, responding to whoever speaks to them next, ageing by the second, and remain entirely capable of changing occupation, changing address, or changing marital status at any point thereafter. The datum entered into the census form does not follow any of this. It remains exactly as deposited, fixed at the point of recording, indifferent to everything the living person goes on to do. The census record and the person it describes diverge from the instant of recording onward, and they never converge again.
3.3 Statistical simultaneity is manufactured
A further complication deserves attention here, because it reveals something about statistical practice not usually made explicit. A census speaks of 'the population in 2026' as though the entire population had been measured at a single instant. In practice, of course, no such simultaneous measurement ever occurs. Enumeration proceeds across weeks or months, household by household, region by region, and the resulting entries are drawn from an enormous number of asynchronous encounters, each occurring at its own particular moment. Statistical practice constructs a common temporal plane onto which these asynchronous observations are subsequently projected, treating entries gathered across an extended period as though they described a single coherent snapshot. This construction is neither an error nor a deception. It is a deliberate and often indispensable simplification. But it deserves to be named for what it is: population statistics manufactures synchronicity out of what was, in its production, an inherently diachronic sequence of separate recursive encounters.
3.4 Snapshot ontology
A statistical dataset, considered in light of the preceding three subsections, is therefore not a population. It is a structured collection of snapshots, each one a nonrecursive trace of a moment already past by the time it enters the record. Even longitudinal datasets, which follow the same individuals across repeated waves of measurement, consist of sequences of such snapshots rather than of any continuous access to the transitions occurring between them. The movement from one wave's recorded state to the next wave's recorded state can be inferred, modelled and estimated with considerable sophistication, but it is not directly present within the entries themselves, which register only the states arrived at, not the living process of arriving at them. A powerful formulation follows from this. Statistics stores states. Recursion consists in transitions. A statistical dataset can compare deposited states, and it can build increasingly refined models of the transitions presumed to connect them, but no accumulation of stored states, however densely sampled, thereby comes to possess the living transition itself. That transition belongs entirely to the recursive process the dataset was drawn from, and it is never handed over to the dataset along with the trace the process left behind.
4. Measurement Is Poiesis
4.1 Poiesis as deposition
The transformation traced in the preceding section has a name within the broader theoretical framework this article draws on. Poiesis is the capacity to deposit recursive coordination into a durable, relatively nonrecursive form, so that later coordination need not reconstruct the original event from nothing. Statistical measurement is among the clearest and most consequential instances of this capacity at work anywhere in modern life. A question is formulated and asked. A response occurs, itself the product of a recursive encounter between questioner and respondent. A distinction relevant to that encounter is selected from among everything that could, in principle, have been recorded. That distinction is then deposited into a durable symbolic form, a coded answer, an entry in a register, a mark on a form, capable of persisting long after the encounter that produced it has ended and everyone who took part in it has moved on.
4.2 The measurement event itself is recursive
This account should not be allowed to slide into an overly mechanical picture in which data simply fall off the world into a waiting spreadsheet. Measurement, at every point, involves decisions that a living investigator has to make, often under conditions of considerable ambiguity: what to measure in the first place, where the boundaries of a category begin and end, when exactly measurement is to occur, which unit of measurement will be used, what counts as a single case rather than several, what counts as a missing value rather than a genuine zero, what counts as a death for the purposes of a mortality register, what counts as unemployment for the purposes of a labour survey, what counts as a household for the purposes of a census. In social statistics especially, these decisions frequently unfold through intensive interrecursive exchange: an enumerator negotiating an answer with a respondent uncertain how to categorise their own situation, a clinician arriving at a diagnosis through extended dialogue with a patient, a police officer determining how an incident should be logged, a teacher assessing a student's performance, an administrator processing an applicant's claim. The measurement event, in every one of these cases, is itself thoroughly recursive, however nonrecursive its eventual output turns out to be.
4.3 Then the trapdoor closes
Once the recursive encounter has run its course, the resulting datum presents itself to everyone who subsequently encounters it as a stable, self-evident fact. Unemployed equals one. Depression equals yes. Ethnicity equals a particular coded category. Household income equals a particular figure. The recursive labour that produced this deposit, all the negotiation, judgement and interpretation that went into settling on a particular coded value rather than any of the several others that might have been assigned, disappears behind the symbolic form the deposit takes. Ordinary symbolic practice hides its own production history as a matter of course, a functional feature of symbols generally rather than a peculiarity of statistics, since symbols that constantly announced how they came to mean what they mean would be unusable for the everyday purposes symbols exist to serve. But this convenience has a cost specific to statistical practice. The deposit looks immediate. Its production history becomes invisible to anyone encountering the figure without having taken part in producing it, and this invisibility is precisely what allows a coded value to be mistaken, later, for a direct report of the world rather than the settled outcome of a prior recursive negotiation.
4.4 This is not necessarily distortion
An important methodological discipline needs to be observed at this point, because the argument developed so far could easily be mistaken for an attack on measurement itself, and it is not. Deposition is not inherently a form of distortion. Without it there could be no census, no epidemiology, no clinical trial, no system of national accounts, no demography, no public health surveillance capable of tracking an epidemic across a population too large for any single observer to survey directly. Poiesis, applied to measurement, produces enormous and valuable epistemic capacities that no purely recursive, face-to-face mode of knowing could ever achieve at comparable scale. The issue this article is building toward is not deposition as such. It is the far more specific error of forgetting that deposition occurred, of treating a nonrecursive trace as though it had arrived from the world unmediated by any prior act of selection, judgement and recursive negotiation. That forgetting, rather than measurement itself, is what the later sections of this article are concerned to diagnose.
5. What Large Numbers Actually Do
5.1 Statistics accumulates deposits, not lives
The argument can now be extended from the single datum to the large dataset, and the extension needs to proceed carefully, because this is exactly the point at which an earlier and mistaken formulation tends to creep back in. Suppose a dataset contains forty observations. What has accumulated at that point is not forty recursive beings considered as recursive beings. It is forty entries: forty nonrecursive symbolic deposits, each the settled trace of a particular recursive encounter that has already concluded. At four thousand observations, four thousand such deposits have accumulated. At four million, four million have. In every case, however large the number grows, what the mathematics operates upon remains, without exception, the accumulated deposits rather than the living processes those deposits were drawn from.
5.2 Why large numbers matter
This is not to deny that the number of observations matters enormously for statistical practice, only to specify correctly what it is that changes as that number grows. As the number of recorded observations increases, the distribution of the recorded values may become increasingly stable, in the specific sense that any single additional observation exerts a progressively smaller influence on summary statistics such as the mean or the variance. A single outlying response can shift the average of ten observations considerably; the same outlying response barely moves the average of ten million. This stabilising effect is real, well understood mathematically, and central to why large samples are generally preferred to small ones. But it is, throughout, a property of the accumulated traces themselves, not an ontological transformation occurring in the population those traces were drawn from.
5.3 Abandon the mystical threshold
Statistical training conventionally teaches a series of rules of thumb, a sample of thirty here, a sample of forty there, past which certain approximations are said to become reliable. These thresholds are genuinely useful pragmatic heuristics, developed and refined through long practical experience with particular classes of estimation problem. But they should not be mistaken for evidence of any deeper ontological transition occurring at the threshold itself. The point at which a given approximation becomes trustworthy differs according to the shape of the underlying distribution, the variance of the quantity being measured, the degree of dependence among observations, the size of the effect being sought, the reliability of the measurement instrument, the design of the sampling procedure, and the specific quantity the analysis is attempting to estimate. No universal number exists past which a dataset suddenly becomes, in some deeper sense, statistical. The rules of thumb are pragmatic conveniences, not markers of an underlying phase transition in the nature of what is being studied.
5.4 What stabilizes
Precision matters here, and it is worth stating the claim as narrowly as it should be stated. What stabilises, as observations accumulate, is the distribution of recorded traces. Not the population the traces were drawn from. Not the living process those traces record. Not the recursivity of the beings whose engagement produced the traces in the first place. A statistical object can exhibit steadily increasing stability as observations accumulate, while stability and recursivity remain, throughout, two entirely independent properties, exactly as they are for any nonrecursive counterpart considered elsewhere: a mountain does not become more or less nonrecursive depending on how many times it has been surveyed, and neither does a population become more or less recursive depending on how many entries about it a dataset happens to contain.
5.5 Why this distinction matters
Without holding this distinction firmly in place, it becomes very easy to slide into imagining the population itself as a kind of quasi-object that gradually takes on the properties of its own statistical representation as sample size grows, becoming, in some vague sense, more orderly, more settled, more thing-like the larger the dataset describing it becomes. This article resists that slide at every point. A distribution is a property of a dataset assembled under a particular, specified classification regime. It can track something of enormous importance about the population it was drawn from, and frequently does. But it is never ontologically identical to that population, however large the dataset grows and however stable the resulting distribution becomes.
6. Statistical Assumptions Are Conditions on Trace Production
The argument developed across the preceding sections has direct and, this article will argue, previously underappreciated consequences for how the standard assumptions underlying statistical inference should be understood. Independence, identical distribution, exchangeability and stationarity are ordinarily presented as formal properties of a probability model, conditions to be checked, where possible, against the dataset itself. The account developed here requires a different framing. These assumptions are not merely properties internal to a dataset. In empirical use, they are hypotheses about the recursive histories through which the dataset's nonrecursive traces were produced, and no amount of scrutiny confined to the dataset alone can ever fully verify or refute them.
6.1 Independence
Statistical independence, within a formal probability model, is a precisely defined property: two variables are independent when knowledge of one carries no information about the other. Applied to an actual body of empirical data, assuming independence amounts to treating each recorded observation as though it carries no information relevant to any other observation beyond what the model already specifies. The question this article's framework raises is correspondingly precise: what recursive relations, among the living processes that generated these traces, have been suppressed by the decision to treat the resulting observations as independent? Suppose survey respondents have discussed the questionnaire among themselves beforehand. Suppose pupils sitting a standardised test share a single classroom, a single teacher and a single set of test-taking conditions. Suppose members of the same household share exposures to the same environmental or economic conditions being measured. In every one of these cases, the resulting deposits remain nonrecursive, exactly as any deposit is. But the recursive processes that produced them were interrecursively linked in ways the assumption of independence, applied without further scrutiny, quietly erases. Ignoring that linkage produces a specific and well documented form of misfit between the statistical model and the process it claims to describe.
6.2 Identically distributed observations
A closely related assumption treats the observations in a dataset as though they were drawn from a single, stable underlying distribution, generated by processes sufficiently comparable to one another to be pooled together for analysis. Historical change can undermine this assumption in ways that are easy to overlook precisely because the resulting data continue to look formally comparable. A diagnosis of depression recorded in 1980 and a diagnosis of depression recorded in 2026 may employ entirely different classification criteria, different diagnostic instruments and different clinical thresholds, even though both are entered into a dataset under the identical label. The symbolic deposits look interchangeable. The recursive regimes that generated them, the clinical practices, professional norms and diagnostic categories in use at each moment, were not.
6.3 Exchangeability
Exchangeability raises the point with particular clarity. Under certain formal conditions, a probability model treats the order in which observations were collected as irrelevant to the conclusions drawn from them. Ontologically, this move brackets the entire history through which each individual trace was produced, treating a sequence of recursive encounters as though their sequencing carried no information worth preserving. This bracketing can be entirely appropriate for some purposes and thoroughly misleading for others. Where the process generating successive observations involves contagion, learning, cascading financial behaviour, or institutional reform, sequencing is not an incidental feature of the data that can be safely discarded. It is constitutive of the very process being studied, and treating observations as exchangeable in such cases suppresses temporal structure that the recursive process itself depended on.
6.4 Stationarity
Stationarity assumes that the statistical properties relevant to an analysis, means, variances, correlations, remain stable across the period the data span. The issue this framework raises is, once again, not that the recorded data somehow become recursive over time. It is that successive traces, collected at different points across an extended period, may have been generated by recursive processes that were themselves changing throughout that period, in ways the formally stable series of recorded values does not, on its own, disclose. A statistical series can therefore appear stationary at the level of its recorded traces while concealing a generative history that was anything but stable.
6.5 The new general principle
A single general principle now emerges from these four cases together, and it deserves to stand as one of this article's central claims. Statistical assumptions are not merely properties of datasets, to be checked, confirmed or rejected by inspecting the dataset in isolation. In their empirical application, they are hypotheses about the recursive histories through which nonrecursive traces were produced, hypotheses that reach back beyond the dataset itself into the living processes and encounters that generated it. This explains something that has long troubled careful statisticians without always being stated in quite these terms: checking model assumptions solely from within the dataset can never be fully sufficient, because the dataset by its nature retains no direct trace of the recursive relations, historical shifts, sequencing dependencies and generative changes that produced it. The generative process matters, and it matters in a way no amount of purely internal diagnostic testing can ever completely substitute for.
7. Hacking's Avalanche of Numbers Reinterpreted as Sedimentation
7.1 The nineteenth-century revolution was poietic
With this ontology established, it becomes possible to return to Hacking's history and offer it a more precise theoretical grounding than it has previously received. The nineteenth-century state developed an unprecedented capacity to transform lived events into durable, standardised traces, at a scale and with an institutional reach no earlier period had possessed. Birth became a birth certificate. Death became a death certificate. Crime became a recorded offence, filed under a standardised category. Marriage became a civil registry entry. Illness became a diagnostic classification entered into a hospital or asylum record. Occupation became an enumerated class within a census schedule. The nineteenth-century explosion of printed numerical information Hacking documents was, described in the terms this article has developed, an explosion of multisymbolic deposits, an expansion in a society's poietic capacity to sediment recursive human life into durable, comparable, mathematically tractable form.
7.2 The state needed equivalence
None of this deposition could serve statistical purposes without a further, prior achievement: comparability. Statistics requires that a death recorded in Glasgow be capable of entering the same column, on the same table, as a death recorded in Manchester. It requires that a crime committed in January be classifiable alongside a broadly similar crime committed in November, under a shared category that treats the two as equivalent for the purposes of the count, whatever differences of circumstance separated them. This requirement for standardised equivalence classes is not incidental to the statistical enterprise. It is a precondition for it, and it explains why the nineteenth-century avalanche of numbers depended on an enormous prior labour of classification, definition and bureaucratic standardisation that rarely appears in the resulting tables themselves, however much of the tables' eventual coherence it made possible.
7.3 Hacking's normality
This reframing also allows Hacking's account of the emergence of the statistically normal person to be restated with greater ontological precision. Once traces have been standardised and accumulated at sufficient scale, distributions become visible for the first time: a mean can be calculated, a variance estimated, a normal curve fitted, deviations measured, outliers identified. The normal person, in the specifically statistical sense this history produced, emerges only after large-scale symbolic sedimentation has assembled a mathematically comparable population out of what had previously existed only as scattered, locally known, incommensurable individual cases. The normal, on this account, is not discovered directly within recursive life, the way a mountain is discovered within a landscape. It is discovered within the distribution of standardised traces that recursive life has, through an extensive apparatus of measurement and classification, been made to leave behind.
7.4 Neither arbitrary construction nor transparent discovery
A crude constructivism should be resisted here just as firmly as a naive realism. A nineteenth-century mortality table discloses something real and consequential: people actually died, in the numbers and at the ages the table records, and the resulting mortality curve reflects genuine facts about the conditions those people lived and died under. But the statistical object that discloses this reality, the curve itself, the mean age at death, the comparison across regions or occupations, requires a classificatory and depositional infrastructure that did not exist before the nineteenth century built it, and that infrastructure shapes what the resulting figures can and cannot show. The regularity disclosed by a statistical table can therefore be entirely real without the statistical object disclosing it being ontologically immediate, a formulation that occupies a considerably more useful middle position than either the claim that statistics simply invents the regularities it reports or the claim that statistics simply reads regularities directly off an unmediated world.
8. Hacking's Looping Effects: When Statistical Deposits Re-enter Recursive Life
8.1 Statistical products can become recursively relevant
A statistic, once produced, is a nonrecursive object in exactly the sense this article has been developing throughout. But living beings can encounter that object, and their encounter with it is itself a fully recursive event, capable of altering their subsequent behaviour in ways the original recursive process that generated the underlying data never anticipated. A published risk score, an IQ result, a prevalence estimate for a psychiatric condition, an opinion poll, a ranking, a crime map showing which neighbourhoods report the highest offence rates: each of these is a nonrecursive symbolic deposit that living beings can, and routinely do, encounter, interpret and respond to.
8.2 The sequence becomes cyclical
The sequence this article has traced from recursive process through to statistical output can therefore be extended a further step, and the extension reveals something important about how statistical knowledge actually operates in society. A past recursive process generates nonrecursive data. Statistical analysis is performed on that data, producing a nonrecursive result. Living beings encounter that result, and their encounter constitutes a further recursive uptake, altering how they subsequently behave. That altered behaviour, in turn, generates new deposits, which subsequent statistical analysis will, in due course, operate upon. This is not recursion occurring inside statistics itself. Statistics remains, throughout this entire cycle, a nonrecursive mathematical operation performed on nonrecursive traces. It is recursion occurring around statistics, in the living uptake of statistical results by the very population those results describe, and the distinction between these two is essential to keep clear if the cyclical structure just described is not to be mistaken for a breakdown of the categorical boundary this article has defended throughout.
8.3 Hacking's classifications of people
Hacking's later work on what he called looping effects describes precisely this phenomenon: classifications applied to people can, over time, change the people classified, as those classified come to understand themselves through the classification, organise around it, resist it, or adjust their behaviour in light of it. The framework developed in this article sharpens the mechanism behind this observation considerably. The classification itself, as a symbolic form, remains nonrecursive throughout; a diagnostic category does not alter itself in response to being applied. The person classified is recursive throughout, in the full sense this article has developed. The looping Hacking documents occurs not because the classification somehow becomes recursive, but because the classification acquires recursive relevance within the classified person's own subsequent coordination, becoming part of how they interpret themselves, are treated by others, and organise their future action. The nonrecursive deposit does not change. What changes is the recursive life that has taken the deposit up.
8.4 The statistical system may alter future data
A further consequence follows, and it creates one of the deepest and most persistent problems facing longitudinal social statistics. Once a classification has begun to alter the behaviour of the people it classifies, later observations are generated by a recursive field that differs, sometimes substantially, from the recursive field that generated earlier observations under the nominally identical variable. A rate of diagnosed unemployment, a measured prevalence of a psychiatric condition, or a recorded crime rate at one point in time and the apparently identical variable measured a decade later may look formally comparable while resting on entirely different histories of production, the later measurement partly shaped by the population's own prior encounter with earlier measurements of the same kind. The act of statistically knowing a population can, through exactly this loop, become part of how that population subsequently behaves, and any analysis that treats the earlier and later measurements as straightforwardly comparable risks mistaking a change in the underlying recursive field for a change in the quantity nominally being tracked.
9. Four Cases: What Exactly Is the Statistical Object?
9.1 Mortality
Four brief case studies illustrate how far a statistical object can sit from the recursive process it is drawn from, arranged here from the case where the distance is smallest to the case where it is largest. Life, throughout its course, is recursive. Death, when it occurs, is an event a death certificate sediments into a small number of selected facts: date, cause, age, place. Mortality statistics operate exclusively on these certificates, and the regularities they disclose, mortality curves stable enough to found the entire actuarial industry, can be extraordinarily robust. But what such statistics analyse is a standardised set of records of death, not life itself, and mortality is, among the cases considered here, the easiest to sediment symbolically, because the event being recorded, death, is comparatively simple to bound and to date.
9.2 Unemployment
Unemployment presents a considerably harder case. What exactly counts as unemployed. Someone actively seeking work. Someone available to begin work immediately. Someone who has worked zero paid hours in a reference period. Someone currently registered for a particular category of benefit. Each of these criteria draws the category's boundary in a different place, and different national statistical agencies, applying different definitions, can report markedly different unemployment figures for populations experiencing broadly similar economic conditions. A person's actual economic coordination, their search for work, their discouragement, their informal earning arrangements, their caring responsibilities, is vastly richer than any of these categories, and what a seemingly simple published number actually reports already depends on a considerable prior act of classification, settled long before any individual is counted.
9.3 Psychiatric prevalence
Psychiatric prevalence is a particularly instructive case for the framework this article has developed. Distress, whatever form it takes, is lived recursively through a person's body, their relationships, their surroundings, and the symbolic categories available to them for making sense of what they are going through. A diagnostic instrument selects a limited set of questions from among everything that could, in principle, be asked. Responses to those questions become scores. Scores are compared against a threshold. Cases above the threshold are counted. Prevalence is calculated across the counted cases. At every stage in this sequence, the resulting statistical object moves further from the recursive distress it originated in, while typically acquiring greater mathematical tractability at each remove. None of this makes the resulting prevalence estimate useless; epidemiological surveillance of this kind has produced valuable public health knowledge. It specifies, rather than undermines, exactly what the statistic is and is not able to show.
9.4 Elections
Votes present an unusually clean case at the opposite extreme from psychiatric prevalence. Whatever recursive deliberation, social influence and personal history produced a given voter's decision, the ballot converts that entire recursive process into an exceptionally narrow nonrecursive symbol: a mark beside one candidate's name, or another's, or a spoiled paper counted separately from both. This is why elections can produce mathematically exact outcomes, a vote count correct to the last ballot, from political processes that were, in their production, extraordinarily rich and recursive. A vote count can be perfectly, mathematically correct while explaining almost nothing about the recursive processes, the arguments encountered, the loyalties inherited, the last-minute doubts resolved one way or another, that produced the votes it counts.
10. The Ontological Error of Quantitative Social Science
10.1 The error is not quantification
The argument developed across this article should not be mistaken for another entry in the long-running genre of essays arguing that numbers cannot capture life. Numbers capture exactly what has been deposited into them, precisely and often with great mathematical elegance, and for a great many research questions that is exactly what is required. A well designed survey, a carefully constructed mortality table, a properly specified econometric model, can answer real and important questions with a rigour no purely qualitative account could match. The argument of this article is not an objection to quantification as such.
10.2 The error is trace and process substitution
The characteristic error this article has been building toward throughout is considerably more specific than a general suspicion of numbers. It consists in treating a mathematically precise description of a set of symbolic traces as though it were an exhaustive description of the recursive process those traces were generated from. A test score is treated as though it simply were a person's intelligence. A diagnosis is treated as though it simply were the disorder it names. Income is treated as though it simply were a person's economic life. A vote is treated as though it simply were a voter's political orientation. A survey response is treated as though it simply were a belief. A crime record is treated as though it simply were the crime. Gross domestic product is treated as though it simply were the economy. In each case, the error is not the measurement itself but a subsequent, usually unstated, act of ontological substitution, in which a deposit stands in silently for the entire process it was drawn from.
10.3 Precision can increase while fit decreases
An especially important paradox follows from the argument developed in earlier sections of this article and deserves to be stated on its own. The more thoroughly a symbolic deposit has been standardised, the easier it typically becomes to subject it to precise mathematical treatment. But standardisation achieves this precision, in case after case, by suppressing exactly the parts of the generative process that resisted standardisation in the first place, the ambiguous cases, the contested classifications, the negotiated judgement calls that measurement always requires and rarely records. Mathematical precision and ontological completeness can therefore move in opposite directions rather than together, a possibility easy to overlook when precision is, quite reasonably, taken as a mark of quality in quantitative research.
10.4 The cleaner the data, the more work has already been done
A related point closes this section. A clean dataset, one free of missing values, coded consistently, temporally aligned, categorically standardised, is not unmediated reality delivered directly to the analyst. It is the endpoint of an extensive prior labour: classification decisions, exclusion decisions, coding decisions, decisions about what counts as missing and what counts as a genuine value, decisions about how to align observations gathered at different times into a common temporal frame, decisions about which categories should be treated as equivalent for the purposes of analysis. Cleanliness of this kind is an achievement of poiesis, the product of a great deal of prior recursive labour rendered invisible by its own success, not a property the data simply arrived with.
11. Toward an Ontology of Statistical Evidence
The argument developed across this article can now be synthesised into a general model applicable to any statistical claim, and this section sets that model out explicitly. Every statistical claim can be decomposed into five layers, arranged in the sequence this article has traced from the outset.
The first layer is the recursive source process: what was actually happening, in its full recursive richness, before any observation was made of it. The second layer is the deposition event: how some aspect of that process was selected, observed, measured, classified or recorded, and by whom, under what institutional and interpersonal conditions. The third layer is the nonrecursive trace itself: what exactly was deposited, in what form, according to what categories, and with what left out. The fourth layer is the statistical operation: what mathematical transformation was performed on the accumulated traces, under what assumptions about their independence, comparability, exchangeability and stability across time. The fifth layer is recursive uptake: what happens when living beings encounter the statistical result, and how their encounter with it feeds back into their subsequent behaviour and, potentially, into the traces later analyses will draw upon.
This five-layer decomposition supplies a reusable protocol for evaluating any statistical claim encountered in research or in public discussion. What is the recursive source the claim ultimately rests on? What does the underlying datum actually preserve of that source, and what does it necessarily suppress? What equivalence relations had to be assumed in order to permit the traces to be aggregated at all? What temporal interval separates the living process from the recorded trace, and does that interval matter for the claim being made? What recursive relations among the cases studied, shared exposure, mutual influence, common institutional context, have been set aside by the statistical model's assumptions? And finally, how is the resulting output likely to re-enter living coordination once it circulates, and what further looping effects might that re-entry produce?
Applied consistently, this protocol does not tell an analyst which statistics to trust and which to discard. It tells them what question to ask before trusting any statistic at all, not whether the mathematics has been performed correctly, which is usually the easiest part of the process to check, but what has already happened to a recursive process before the mathematics was ever permitted to touch it.
12. Conclusion: Statistics Never Sees the Living Present
Hacking showed that modernity produced an unprecedented world of statistical objects, an avalanche of printed numbers that transformed how states, scientists and ordinary citizens came to understand society, normality and risk. The argument developed in this article has attempted to specify the ontology underlying that transformation more precisely than the historical record alone can supply.
Statistics does not reach directly into recursive life and translate it mathematically. A transformation occurs first, and every claim this article has made follows from taking that transformation seriously. Living processes are encountered by an observer, a clinician, an enumerator, an institution. Selected aspects of those processes are articulated and recorded. These articulations are deposited into durable multisymbolic traces. The traces accumulate, often at enormous scale. Mathematics then operates upon the accumulated traces, not upon the living processes that gave rise to them. The stability so often observed in the resulting distributions belongs to the accumulated symbolic record, not to any emergent nonrecursivity supposedly achieved by the population itself, which remains recursive throughout, from the first observation to the last.
Nonrecursivity does not emerge from recursive populations through aggregation. Recursive populations remain recursive, regardless of how many observations a dataset happens to contain. Statistics becomes possible not because aggregation transforms recursive life into something else, but because recursive processes can leave nonrecursive symbolic traces behind them, traces an entire modern apparatus of registration, classification and computation has been built to collect, standardise and analyse.
Population statistics is therefore not the mathematics of recursive life itself. It is the mathematics of what recursive life has already left behind. Statistics never sees recursive life in the present tense. It sees the nonrecursive symbolic past that living processes continuously deposit behind them.