Abstract

The Illusion of Explanatory Depth (IOED), first coined by Rozenblit and Keil in 2002, claims that people systematically overestimate their understanding of complex phenomena. The finding has been highly influential, widely replicated, and routinely cited as evidence of human cognitive bias. This paper argues that the IOED paradigm is fundamentally flawed at a structural level. Drawing on Living Value Theory (LVT), which distinguishes levels of recursivity and ontologically distinct recursivity domains, the analysis identifies five interlocking problems. First, the paradigm fails to specify any determinate criterion of correctness or sufficiency for explanations, rendering its central claim—that drops in self-rated understanding reveal overconfidence—uninterpretable. Second, it conflates distinct levels of recursivity, mistaking the normal gap between embodied competence (L1/L2) and symbolic articulation (L3) for cognitive error. Third, it collapses ontologically heterogeneous domains, treating non-recursive mechanical systems (e.g., toilets) and interrecursive social phenomena (e.g., immigration) as equivalent objects of understanding. Fourth, the measurement instrument itself constitutes a new social fact rather than neutrally recording a pre-existing state. Fifth, the elicitation procedure introduces systematic interrecursive biases through its challenge register. A final section examines two decades of replications and extensions, concluding that they elaborate the misidentification without addressing its foundational weaknesses. The apparent robustness of the IOED effect is therefore the robustness of a theoretical illusion, not a genuine cognitive one. The paper contends that the IOED research programme pathologises a constitutive feature of human expertise and should be retired in favour of frameworks that respect the inexhaustibility of mesocosmic coordination under symbolic rendering, the ontological differences between domains, and the constitutive nature of measurement.

Keywords: illusion of explanatory depth, Living Value Theory, recursivity, metacognition, philosophy of psychology, measurement theory

In 2002, Leonid Rozenblit and Frank Keil published what became one of cognitive psychology's most celebrated findings. People, they argued, believe they understand complex phenomena with far greater precision and depth than they actually do. This illusion of explanatory depth, as they called it, was held to reveal a fundamental feature of human cognition: that the mind systematically overestimates its own explanatory competence. The paper has since attracted thousands of citations, spawned a substantial replication industry, and become a standard exhibit in popular accounts of human irrationality. It is taught in introductory psychology courses, cited in policy discussions about overconfident experts, and invoked in debates about the limits of folk understanding.

This article argues that the experiment is wrong, and wrong in several distinct, diagnosable ways that compound each other. The errors are not matters of insufficient nuance or limited generalisability. They are structural. They run through the experimental design, the theoretical interpretation, and the conceptual vocabulary in which the findings are expressed. The illusion of explanatory depth, as a research finding, is itself an instance of the very pathology it purports to describe: a confident symbolic claim that collapses under analytical pressure.

The critique developed here draws on Living Value Theory, a framework that distinguishes between levels of recursivity, recursivity domains, and mediational configurations in ways that expose the experiment's foundational category errors. Five structural objections are developed in sequence, ordered from the most foundational to the most specific. The first, and most fundamental, is that the experiment never specifies what would count as a correct or sufficient answer in any domain it tests, which means it cannot establish that any gap between confidence and performance represents an illusion rather than a reasonable response to an ill-formed question. The second is that the experiment conflates distinct levels of recursivity, treating the gap between embodied competence and symbolic articulation as cognitive bias rather than as a constitutive feature of human skill. The third is that it collapses ontologically heterogeneous domains into a single paradigm, comparing non-recursive physical systems with interrecursive social processes as if these were equivalent objects of understanding. The fourth is that the measurement instrument does not passively record a pre-existing cognitive state but actively constitutes a new social fact, confounding measurement with intervention. The fifth is that the framing and tone of the elicitation procedure introduce systematic interrecursive bias that the framework has no resources to recognise or control. A sixth section addresses the extensions and refinements that have accumulated in the literature since 2002, arguing that they elaborate the misidentification rather than correct it.

I. The Indeterminacy of Correctness

Before any of the other objections can be assessed, a prior question must be answered: what would count as a correct explanation in any of the domains the experiment tests? The IOED paradigm assumes this question has a determinate answer. The assumption is false, and its falsity is not a peripheral methodological concern. It is the foundational problem from which every other error follows.

Consider the experiment's most apparently tractable case: the toilet. Participants rate their understanding of how a toilet works, attempt an explanation, and re-rate. The drop in rating is interpreted as evidence that they overestimated their understanding. But what would a correct and sufficient explanation of how a toilet works actually look like? The immediate answer is: the siphon mechanism, the ballcock valve, the cistern filling to a set level and then triggering the flush. This account is serviceable at one level of description. It is also radically incomplete, and its incompleteness is not a matter of insufficient effort but of the inexhaustibility of the phenomenon under any symbolic rendering.

A toilet works with gravity. Gravity is not a detail that can be assumed away. It is the mechanism through which water descends from cistern to bowl, through which the siphon creates the pressure differential that produces the flush, through which waste moves through the drainage system. A complete explanation of how a toilet works therefore requires an account of gravity. But gravity, as a phenomenon, defeated Newton's explanatory ambitions in a famous sense: he could describe the mathematics of gravitational attraction with extraordinary precision while explicitly refusing to hypothesise about its underlying nature. Einstein's general relativity reconceived gravity as the curvature of spacetime produced by mass. Quantum gravity remains unresolved. The toilet, on even minimal inspection, opens into one of the deepest unsolved problems in physics. No participant in the IOED experiment will be expected to address this. But no principled boundary has been established at which the explanation can be said to be complete. The experimenters simply stop asking at the point where the answer feels intuitively satisfying to them.

The drainage system connects the toilet to a sewerage infrastructure that operates across municipal, geological, and hydrological scales. The materials from which the toilet is manufactured involve ceramics, plastics, and metals whose production histories extend across global supply chains. The water that flows through it carries microbiological communities whose role in the broader ecology of sanitation is a matter of ongoing scientific research. None of this is exotic or far-fetched. It is the actual causal structure of the thing being explained. The question 'how does a toilet work' is not a question with a determinate answer. It is a question that opens into the entire material and physical organisation of the world the toilet inhabits, and any stopping point is arbitrary.

What the IOED experiment actually does is impose an implicit, unstated, and theoretically unjustified stopping criterion. The experimenter has in mind a rough sketch of the siphon-and-ballcock account, and explanations that approximate this sketch are scored as adequate while explanations that fall short of it are scored as inadequate. But this criterion is never made explicit, never justified, and never examined for consistency across participants or domains. Different experimenters, or different implicit standards, would produce different verdicts on whether a given explanation counts as sufficient. The experiment presents this as measuring participants' understanding against an objective standard of correctness. It is actually measuring participants' explanations against an unarticulated and variable experimenter intuition.

This problem is not incidental to the IOED finding. It is generative of it. The consistent drop in ratings after explanation attempts is at least partly explained by the fact that participants, in attempting to explain, discover the inexhaustibility of the phenomenon under symbolic rendering. They begin to see how much more there is to say. This discovery is then scored as evidence that their initial confidence was inflated. But it might equally be scored as evidence that they have become more sophisticated about the nature of the question, that they have moved from a naive assumption that the question had a tractable answer to an accurate recognition that it does not. The experiment cannot distinguish between these interpretations because it has no account of what a correct answer would look like.

Living Value Theory provides the conceptual vocabulary for understanding why this is so. Mesocosmic coordination across the five mediations, embodiment, being-with, dwelling, multimateriality, and multisymbolisation, is inexhaustible under symbolic rendering. Any symbolic account of a mesocosmic process captures a slice while leaving the rest untouched. The toilet involves materials, gravity, water chemistry, infrastructure, manufacturing, microbiological processes, and usage habits that together constitute a coordinative reality of which any verbal explanation is necessarily a radical reduction. The question of whether an explanation is correct or sufficient cannot be answered from within the symbolic register because the phenomenon being explained exceeds any symbolic register in which it could be captured. The experimenter's implicit stopping criterion is not a solution to this problem. It is a way of not noticing that the problem exists.

The consequence for the entire IOED research programme is severe. If no determinate criterion of correctness can be established for any domain the experiment tests, then the gap between initial rating and post-explanation rating cannot be interpreted as evidence of an illusion. It can equally be interpreted as participants' growing awareness of the inexhaustibility of the phenomenon, as a response to the implicit social pressure to appear appropriately humble after being tested, or as an artefact of the experimenter's unarticulated and variable stopping criterion. The experiment has no resources to distinguish among these interpretations. It presents one of them as the finding. The others are never considered.

This objection has a further implication that the extensions literature has entirely failed to address: if no criterion of correctness can be established, the experiment cannot in principle distinguish expert from lay understanding in any domain. An expert physicist explaining gravity and a child explaining gravity will produce accounts that differ in vocabulary, in depth, in internal consistency, and in the range of phenomena they connect. But neither account is complete, and neither can be said to have reached the point at which the phenomenon is fully explained. The experiment scores the child's account as less adequate than the physicist's, but this scoring depends entirely on the unstated intuition that the physicist's account is closer to a correctness criterion that has never been specified. If that criterion were specified, it would immediately be revealed as either trivially low, in which case the child might meet it, or properly demanding, in which case the physicist would fail it too. The experiment avoids this problem by never specifying the criterion. In doing so, it disguises an evaluative judgment as an objective measurement.

II. The Recursivity Level Conflation

The experiment's core procedure is elegant and reproducible. Participants are asked to rate their understanding of a phenomenon on a numerical scale. They are then asked to produce a detailed causal explanation of that phenomenon. They re-rate their understanding. The rating drops. This drop is interpreted as evidence that their initial confidence was inflated, that they believed they understood more than they actually did.

The interpretation depends on an unexamined assumption: that the ability to produce a satisfactory causal explanation is the appropriate test of understanding. This assumption is false, and its falsity is not a matter of degree but of kind.

Living Value Theory distinguishes five levels of recursivity. At L1, coordination proceeds seamlessly, without reflection or deliberation. Skills are enacted, habits operate, and the social world is navigated without any requirement that the navigator be able to articulate what they are doing. This is not a deficient form of knowledge. It is the most successful form of coordination, and its success is precisely indexed by the degree to which it has descended below the threshold of articulable awareness. At L2, coordination becomes unsettled: something feels wrong, a hesitation arises, attention is recruited. At L3, something is symbolically named, made available as a describable entity. At L4, abstraction and generalisation produce stable concepts and rule-like structures. At L5, the symbolic systems themselves become objects of reflection.

The experiment's testing procedure operates at L3: it asks participants to produce verbal causal accounts, sequences of named mechanisms linked by explicit logical relations. What it then interprets as the gap between inflated confidence and true understanding is in fact the structural gap between L1/L2 competence and L3 articulatory capacity. These are not the same thing, and the gap between them is not a bias. It is a constitutive feature of embodied human expertise.

The cyclist who cannot explain gear ratios is not overconfident about cycling. Their L1 embodied competence is genuine, fully operative, and in the relevant sense superior to any symbolic account that could be given of it. The surgeon who cannot articulate the proprioceptive adjustments she makes during a procedure is not suffering from an illusion of surgical depth. Her hands know what her language cannot say. The experienced teacher who struggles to explain why a particular intervention worked with a particular student is not deluded about her pedagogical understanding. She is demonstrating that L1 and L2 competence are genuinely prior to and independent of L3 symbolic capacity. No one can articulate embodied coordination better than they can perform it. This is the correct relationship between non-symbolic and symbolic mediations. It is the normal and healthy condition of human expertise in every domain where skill is acquired through practice rather than through instruction.

The IOED experiment pathologises this relationship. It takes the irreducible gap between doing and saying and calls it overconfidence. The correct conclusion from the consistent finding that felt competence outruns articulatory performance is not that people think they know more than they do. It is that people can do more than they can say, and that this is as it should be. The experiment has spent two decades measuring the structural priority of embodied coordination over symbolic articulation, mistaking a constitutional feature of human skill for a cognitive defect, and calling this discovery cognitive science.

There is a further dimension to this error that the IOED literature has entirely failed to register. The interesting empirical question raised by the gap between L1/L2 competence and L3 articulatory capacity is not whether the gap exists, which is trivially true and unsurprising, but under what conditions explicit symbolic articulation enhances or degrades embodied coordination. Sometimes L3 reflection improves performance: the novice who learns to narrate their steps acquires habits more quickly. Sometimes it catastrophically disrupts it: the expert golfer who begins to think consciously about her swing introduces the yips. The centipede does not walk better for having been asked to explain which leg it moves first. These phenomena, well documented in the literature on expert performance, motor learning, and flow states, are precisely what a theoretically adequate account of the relationship between recursivity levels would need to explain. The IOED paradigm, by treating the gap as simply an error to be corrected rather than a structurally significant relationship to be understood, forecloses the productive research programme that its own data was pointing toward.

III. The Recursivity Domain Conflation

The experiment's stimulus materials range from zippers, toilets, and helicopters to immigration policy, capital punishment, and carbon trading. This range is presented as a strength of the experimental design, demonstrating that the illusion operates across diverse domains. It is, in fact, a fatal flaw, because the domains are ontologically heterogeneous in ways that make uniform findings impossible to interpret.

Living Value Theory distinguishes three fundamental recursivity domains. Non-recursive processes are those in which the elements of the process do not respond to expectations about them. The pharmacokinetics of a drug, the mechanical operation of a siphon, the aerodynamics of a rotor blade: these processes are indifferent to beliefs about them. They operate through fixed causal mechanisms that run independently of any observer's understanding. Self-recursive processes are those in which the organism's relationship to its own states is part of what shapes those states: pain experience, anxiety, and fatigue are all modulated by the organism's own monitoring of them. Interrecursive processes are those in which multiple parties recursively affect each other's expectations, orientations, and behaviours, producing outcomes that cannot be specified independently of the recursive loop itself.

A toilet is non-recursive. It has a determinate mechanism. The siphon either works or it does not. The ballcock valve either closes at the right water level or it does not. There is a fact of the matter about how the toilet works, and the question of whether someone understands it has a clear answer, modulo the indeterminacy of correctness established in Section I: does their account correctly describe the mechanism at the level of description the experiment implicitly assumes? The gap between felt understanding and articulatory performance in this case is the L1/L2 versus L3 gap already described: genuine functional familiarity without symbolic command of the causal structure.

Immigration is an interrecursive process that has been multiply institutionalised, historically sedimented, politically contested, and symbolically saturated across all five mediations simultaneously. There is no equivalent fact of the matter against which an account of immigration could be checked as correct or incorrect in the way a toilet account can, even at the level of description the experiment implicitly assumes. What counts as understanding immigration policy is itself interrecursively contested: academic economists, legal practitioners, border officials, migrants, and political theorists disagree not because some of them have better causal accounts and others have worse ones, but because the object itself is constitutively interrecursive. Policy declarations about immigration alter immigration. The entry of new symbolic accounts into the public sphere changes the phenomenon those accounts purport to describe. A toilet declaration does not alter toilets.

When a participant rates their understanding of immigration at seven out of ten and then fails to produce a satisfactory causal account under experimental prompting, the IOED framework interprets this as the same phenomenon as rating toilet understanding at seven and failing to explain the siphon. This interpretation is not merely imprecise. It is a category error of the first order. The two failures are entirely different in kind. The toilet failure is an instance of L1/L2 competence without L3 articulatory capacity, as described above. The immigration failure is something altogether different: the participant's felt understanding is a genuine apprehension of the stakes, the human consequences, the structural position, and the political contestation of immigration as a social phenomenon, none of which is reducible to a causal mechanism account, because no such account exists for an interrecursive process.

Consider the precise structure of a sophisticated response to the immigration case. A person may be highly confident about what immigration is: a movement of people across political borders with consequences for labour markets, cultural composition, legal status, and national belonging, consequences that are themselves recursively shaped by the policies, discourses, and institutional practices that surround them. They may simultaneously acknowledge that they do not know in detail how the visa application procedure for a Tier 2 skilled worker operates, or what the precise legal distinctions between asylum seeker and refugee status entail in the current UK framework. This is not overconfidence. It is accurate, differentiated self-knowledge operating across multiple dimensions of understanding simultaneously. The person knows what they know and knows what they do not know. This is the epistemic configuration a well-calibrated agent should have about a complex interrecursive domain. The IOED experiment has no resources to see this distinction because it collapses all dimensions of understanding into a single rating and then tests only one of them.

Any honest experimental investigation of explanatory understanding would need, as a prior analytical requirement, an ontological classification of its stimulus materials by recursivity domain. Without this prior classification, a finding that people's confidence exceeds their articulatory performance across both non-recursive and interrecursive domains tells us nothing about cognition. It tells us only that the experimenters did not notice that they were comparing incommensurable things.

IV. The Measurement Constitutes the Phenomenon

The fourth structural objection concerns the status of the measurement instrument itself. The experiment treats the self-rating as a neutral record of a pre-existing cognitive state: the participant has a level of felt understanding, which the rating scale captures, which is then compared to the articulatory performance that reveals the true level. This account of what the rating scale does is wrong.

When a participant is asked to rate their understanding of a phenomenon on a numbered scale, they are not reporting an internal cognitive state that exists prior to and independently of the act of reporting. They are performing a formal declaration of epistemic position in a socially significant context. This declaration does not express a pre-existing state. It constitutes a new social fact. The participant who rates their understanding at seven in the presence of a researcher has made a public commitment to a level of competence that will enter the recursive ecology of the experimental situation and shape everything that follows. The researcher now holds an expectation. The participant now inhabits a situation in which their subsequent performance can confirm or disconfirm their declared level. The rating is not a report. It is a performance.

The South Asian philosophical tradition has a concept that captures this precisely: the sankalpa, the formal declaration of future intention or present capacity made in a socially significant context. The declaration works not through any internal cognitive mechanism but through the social fact it constitutes. It reorganises the recursive ecology within which subsequent action occurs. The IOED experiment has been administering sankalpas and measuring their consequences without knowing it.

The practical consequences for the experimental design are severe. If the baseline rating is not a neutral measurement of a pre-existing state but a social act that constitutes a new fact and reorganises the recursive ecology, then the comparison between baseline rating and articulatory performance is not a comparison between felt understanding and true understanding. It is a comparison between the social commitment constituted by the first measurement and the performance produced under the altered conditions that the first measurement itself created. The baseline and the test are not independent. The baseline is an intervention that shapes what the test measures.

There is a further layer to this problem that the literature has not addressed. The numerical form of the rating scale introduces a third constitutive loop beyond the sankalpa effect. When participants are asked to express their understanding as a number between one and ten, they are required to perform an L4 abstraction on their own L1/L2 states. Their actual coordinative relationship with the domain in question is distributed across embodied familiarity, relational knowledge, situational experience, and symbolic exposure, none of which is introspectively accessible as a unified quantity. The act of producing a number forces a premature L4 stabilisation of something that is genuinely pre-symbolic. The number then re-enters the person's recursive ecology not only as a social commitment but as a newly constituted self-concept. The person becomes, in some small but real way, a seven. This reified self-concept then reorganises their L1/L2 coordinative states around the symbolic artefact that the measurement instrument produced. The instrument is not merely confounded with the intervention. It is itself an intervention that restructures the mesocosmic configuration it purports to measure at the level of self-understanding.

V. The Interrecursive Bias of Elicitation Register

The fifth structural objection concerns the tone and manner in which the question is asked. This is the dimension of the experiment's confounds that is simultaneously most obvious and most completely ignored in the literature.

There is a categorical difference between asking someone casually how well they understand something and asking them in a context that signals they are about to be tested. The casual register invites something approximating an L2 felt-sense report: a rough, low-stakes approximation of one's relationship to the domain in question. The challenge register, in which the framing signals that performance will follow and that failure will be visible, does something entirely different. It activates full interrecursive anticipation: prospective modelling of social consequences, calculation of the costs of bold versus cautious declarations, recruitment of identity-protective strategies, and anticipatory shame management. The number produced in each register is formally identical on the questionnaire. The phenomenon being elicited is completely different.

The challenge register does not elicit a confidence report. It elicits a courage declaration under social pressure. The participant who rates themselves at eight when they know they will immediately be asked to explain is not reporting felt understanding. They are making a public commitment to a level of performance they are willing to be seen risking. This is a declaration of risk tolerance and social self-confidence as much as it is any epistemic self-assessment. The experiment cannot distinguish these because it has no concept of the elicitation act as an interrecursive event with its own social logic.

There is a further systematic bias introduced by the challenge register that the IOED literature has not registered. The willingness to make bold declarations under conditions of anticipated evaluation is not uniformly distributed across populations. People with higher social confidence, more experience of institutional settings, stronger identity investment in appearing knowledgeable, and greater familiarity with the norms of academic and professional self-presentation will rate higher under challenge conditions, not because their coordinative competence is greater but because their recursive ecology makes bold declarations safer. The challenge register therefore introduces a systematic class and cultural bias into the measurement that is entirely invisible to the experimental framework. Two participants with equivalent L1/L2 coordinative competence in a domain will produce different ratings if one of them is habituated to institutional performance contexts and the other is not. The experiment will interpret this difference as a difference in felt understanding. It is a difference in recursive social positioning.

VI. The Extensions Do Not Help

The two decades since the original Rozenblit and Keil paper have produced a substantial body of follow-up work. Replications have confirmed the basic effect across populations and cultures. Studies have examined whether the illusion is stronger for mechanical objects than for abstract concepts. Others have extended the paradigm to procedural knowledge, asking not merely whether people can explain how something works but whether they can explain how to do something. The effect has been applied to public understanding of science, to political polarisation research, and to medical decision-making. Methodological refinements have introduced different rating scales, varied the timing of the explanation task, and attempted to control for social desirability and demand characteristics. Researchers have compared experts with novices, tested individual differences in cognitive reflection, and examined whether the illusion is modulated by elaborative interrogation. These extensions are offered, implicitly or explicitly, as evidence that the IOED effect is robust, replicable, and theoretically progressive.

The assessment offered here is that none of these extensions addresses any of the five structural objections. They are normal-science refinements within a paradigm whose foundational assumptions remain entirely unchallenged. The most revealing case is the expert-novice distinction, which directly bears on the indeterminacy of correctness problem. If the experiment cannot specify what would count as a correct or sufficient explanation in any domain, it cannot in principle determine whether experts perform better than non-experts against an objective standard, or merely approximate the experimenters' unstated intuitions more closely. An expert on toilet mechanics will produce an account that better matches the experimenter's implicit stopping criterion. But the criterion itself remains unspecified and unjustified. The expert's account is still radically incomplete, still opens into gravity, fluid dynamics, material science, and infrastructure, and is still scored as adequate by an evaluator who stops asking questions at the point where the account feels satisfying to them. The expert-novice studies have confirmed that expertise produces accounts that better satisfy experimenter intuitions. They have not established that expertise produces accounts that are correct, because no criterion of correctness has been established against which either account could be assessed.

The procedural knowledge extension similarly fails to engage with the foundational objection. Shifting from causal explanation to procedural instruction, from 'how does it work' to 'how do you do it,' still requires L3 symbolic articulation of L1/L2 embodied competence. The person who can ride a bicycle but cannot produce a step-by-step account of balance, steering, and propulsion is still being penalised for the structural gap between doing and saying. The extension has varied the label on the test without touching the category mistake that generates the apparent finding.

The domain comparison studies partially acknowledge the recursivity domain problem by noting that the effect is stronger for mechanical than for political domains. This is the literature noticing, without the vocabulary to name, that it has been comparing incommensurable things. But observing a difference in effect size is not the same as recognising that the domains require entirely different epistemological frameworks for evaluation. The studies continue to use the same elicitation procedure across domains that are ontologically heterogeneous in the ways Section III describes. They have mapped the topography of the category error without dissolving it.

The methodological refinements aimed at controlling for social desirability and demand characteristics represent the literature's most direct engagement with the elicitation register objection. Standard demand characteristic controls attempt to subtract the social performance dimension from the measurement, treating it as noise around a signal that could in principle be cleanly recovered. The LVT argument is that there is no such signal. The social performance dimension is not noise contaminating a pure epistemic self-report. It is constitutive of what the measurement act produces. A demand characteristic control that attempts to isolate the epistemic report from the social commitment is attempting to isolate something that does not exist as a separable entity.

There is a meta-point about the extensions considered as a body of work that bears stating explicitly. A finding that is robust across extensions is not thereby made intelligible. If the paradigm is systematically misidentifying what it studies, then replicating its findings more carefully and across more conditions produces a more robust misidentification, not a correction of it. The extensions have constructed a more elaborate edifice on a foundation that the structural critique has shown to be unsound. The elaborateness of the edifice is not evidence of the soundness of the foundation.

VII. The Survival of a Flawed Paradigm

The question that naturally follows is why a paradigm with these structural deficiencies has not been demolished by peer review, and why two decades of extensions have elaborated rather than corrected it. The answer is itself an LVT story about institutional recursivity, symbolic class dynamics, and the political economy of research.

The IOED framework achieved early L4 institutional stabilisation: it was rapidly absorbed into textbooks, popular science accounts, and the conceptual vocabulary of adjacent fields. Once stabilised at this level, it became self-validating. Peer reviewers are trained within the paradigm, embedded in its productive apparatus, and lacking the conceptual resources that would be required to articulate what is wrong with it. The framework has no vocabulary for recursivity levels, recursivity domains, the constitutive social effects of measurement, or the indeterminacy of correctness criteria. A critique formulated in those terms is, from inside the paradigm, simply not legible as a critique. It appears as a philosophical objection rather than a scientific one, and the distinction between philosophical and scientific objections is itself one of the paradigm's self-protective mechanisms.

There is a further dynamic that sustains the paradigm independently of its intellectual merits. The finding that people overestimate their understanding is politically convenient. The overconfidence narrative aligns with a broader cultural and institutional investment in expert authority as a corrective to lay hubris, and technical complexity as a domain properly reserved for credentialled specialists. These investments reproduce the authority of the symbolic class over those whose competence is primarily embodied, relational, and practical. An experiment that showed that people can do more than they can say, and that no determinate criterion of correctness can be established for any question about how a complex process works, would have a very different political valence from one that shows people think they know more than they do. The IOED finding is the latter, and its cultural traction is partly explained by this valence.

VIII. What an Adequate Account Would Look Like

An adequate account of the phenomena the IOED experiment was groping toward would need to begin with the indeterminacy problem. Before asking whether people's confidence outruns their understanding, an adequate framework would need to ask what understanding consists in for each domain under investigation, what a correct or sufficient explanation would look like, and at what level of description the question is being posed. These questions do not have general answers. They require prior ontological work on the nature of each domain. For the toilet, an adequate account would need to specify at which level of physical and material description the explanation is expected to operate, and why stopping at the siphon-and-ballcock level is justified rather than arbitrary. For immigration, an adequate account would need to acknowledge that no causal mechanism account of the right kind exists, because the phenomenon is constitutively interrecursive and exceeds any static description that could be given of it.

An adequate account would also need to distinguish between the multiple dimensions of understanding that the IOED paradigm collapses into a single rating. Understanding what an entity is, understanding its stakes and consequences, understanding its procedural mechanics, understanding its institutional history, and understanding the recursive logic through which it constitutes itself are different achievements that can come apart in any combination. A person who scores high on some of these dimensions and low on others is not overconfident. They are accurately tracking a genuinely heterogeneous epistemic terrain. This is not a bias. It is calibration.

An adequate account would need to treat the measurement instrument as a constitutive social act and build its experimental design accordingly. The finding that ratings predict performance better when produced immediately before performance, in public rather than private contexts, and in the presence of credible authorities is entirely inexplicable on the assumption that the rating is measuring a stable internal property. It is entirely explicable on the account of the rating as a social act whose constitutive force varies with the recursive conditions of its production.

Finally, an adequate account would recognise that the productive question raised by the consistent gap between felt competence and articulatory performance is not whether people are overconfident but what the relationship between different recursivity levels of competence is, and under what conditions explicit symbolic articulation enhances rather than degrades the embodied and relational coordination that constitutes the actual substance of human expertise. That question cannot be asked, let alone answered, within the IOED paradigm. It requires an ontological framework adequate to the inexhaustibility of mesocosmic coordination under symbolic rendering.

The illusion of explanatory depth is not a cognitive illusion. It is a theoretical illusion, produced by experimenters who could not specify what a correct answer would look like, did not notice that their domains were ontologically incommensurable, mistook a constitutive social act for a neutral measurement, confused the gap between doing and saying for the gap between competence and ignorance, and then built an institutional apparatus that has sustained the misidentification for two decades. Dismantling the illusion requires not better experimental controls, more diverse stimulus sets, or further procedural refinements. It requires a different theory of what understanding is, what measurement does, what kinds of entities the world contains, and what would even count as answering a question about how something works. That theory is available. The experiment can be safely flushed down.