Home Investigations Research report
Research report
● signed research report

The Examined Cognition: Substrates of Metacognition and Methods for Looking at the Looking

By the operator·2026-07-22·26 min read
Download clean Markdown 5.2k words The source of record · PDF on request, generated from this page so it never goes stale
The short version
  • Metacognition is not a single faculty but a layered ecology: a frontopolar/anterior PFC hub (BA10) for retrospective confidence, a default-mode core (mPFC, precuneus, PCC) for self-referential modeling, an insular/cingulate salience apparatus for interoceptive grounding, and an executive system for control — all best understood as the brain's running hierarchical model of its own model, in the predictive-processing sense developed by Friston, Clark, Seth, and Metzinger.
  • Examining cognition analytically requires multi-method triangulation: signal-detection measures (meta-d′, M-ratio) for third-person rigor; micro-phenomenological elicitation (Petitmengin) and neurophenomenology (Varela) for first- and second-person rigor; contemplative training to render the self-model opaque (Metzinger); and interpretability/introspection probes (Jack Lindsey, Anthropic, October 2025) for machine analogues.
  • McLuhan's figure/ground is the master frame: ordinary cognition is figure, the substrates that make it possible are ground, and metacognition is the labor of dragging ground into figure — which is precisely why we tend to perceive our prior cognitive environment (the rear-view mirror) rather than the medium we are currently inside.
Key findings · 0

Metacognition has been intensively localized in the rostrolateral / frontopolar prefrontal cortex (BA10). Fleming, Weil, Nagy, Dolan & Rees (Science 2010) showed that gray-matter volume in right rlPFC predicts introspective accuracy; Fleming, Ryu, Golfinos & Blackmon (Brain 2014) showed that lesions to anterior PFC produce domain-specific metacognitive deficits while leaving first-order task performance intact. A meta-analysis by Vaccaro & Fleming (2018) found that right anterior dlPFC and lateral frontopolar cortex preferentially activate during metacognitive judgments across perception, memory, and value-based choice; primate work by Miyamoto, Setsuie, Osada & Miyashita (Neuron 2018) closed the causal loop by reversibly silencing frontopolar cortex and selectively impairing metacognitive judgment on non-experience.

The cleanest psychophysical instrument for studying this is meta-d′ / M-ratio, proposed by Maniscalco & Lau (Consciousness and Cognition, 2012). Meta-d′ asks: given the confidence ratings an observer produced, what type-1 sensitivity would have generated that pattern? The ratio meta-d′/d′ (M-ratio) thereby dissociates metacognitive efficiency from raw task performance — letting researchers ask whether someone is well- or ill-calibrated independently of how well they see, remember, or decide.

The default-mode network (mPFC, posterior cingulate, precuneus, angular gyrus, lateral temporal cortex) provides the substrate for self-referential and autobiographical processing on which metacognition piggybacks; the salience network (anterior insula, dorsal ACC) detects the moment of mind-wandering and switches between modes; the central executive / frontoparietal control network sustains and shifts attention. Hasenkamp, Wilson-Mendenhall, Duncan & Barsalou (NeuroImage 2012) mapped exactly these networks onto the four phases of a focused-attention meditation cycle — mind-wandering (DMN), awareness of mind-wandering (salience), shifting (CEN), sustained focus (CEN). Brewer et al. (PNAS 2011) showed experienced meditators have reduced DMN activity in PCC and mPFC across meditation styles; Baird, Mrazek, Phillips & Schooler (J. Exp. Psychol. Gen., 2014, N=52) demonstrated in a randomized trial that two weeks of focused-attention training "significantly enhanced introspective accuracy, quantified by metacognitive judgments of cognition on a trial-by-trial basis, in a memory but not a perception domain" — confirming that introspective skill is plastic but not generic.

At the philosophical level, the architecture splits into competing camps. Higher-order theories (Rosenthal; Lau & Rosenthal, TiCS 2011; Brown, Lau & LeDoux, TiCS 2019) hold that a first-order state becomes conscious only when meta-represented; global workspace theory (Dehaene, Baars, Mashour) emphasizes broadcasting across a fronto-parietal workspace; self-representational views (Kriegel) collapse the meta into the first-order state itself. Predictive processing / active inference (Friston; Clark; Hohwy; Seth) reframes all of these as variants of self-modeling: the brain minimizes free energy by maintaining a generative model whose self-as-controller is, in Metzinger's words, "a self-model that cannot be recognized as a model by the system using it." This phenomenal transparency is the deep reason introspection so often misfires — we look with the model, not at it.

For machine cognition, the parallel substrate questions are now empirical. Jack Lindsey (Anthropic, "Emergent Introspective Awareness in Large Language Models," transformer-circuits.pub, October 29, 2025) injects concept-aligned steering vectors into Claude's activations and asks the model to report; he reports that "Claude Opus 4 and 4.1, the most capable models we tested, generally demonstrate the greatest introspective awareness," and that "models can, in certain scenarios, notice the presence of injected concepts and accurately identify them … some models can use their ability to recall prior intentions in order to distinguish their own outputs from artificial prefills." Subsequent work by Yang, Sheshadri, Mindermann, Lindsey, Marks & Wang (Anthropic Alignment Science, "Introspection Adapters," April 28, 2026) fine-tunes models to self-report learned behaviors, achieving 59% success on AuditBench versus 53% for the next-best method and 44% for the best white-box method. Both confirm what philosophers have long suspected of humans: introspection is real but partial, post-hoc, and dependent on architecture-specific affordances.


Caveats & confidence

    The Examined Cognition: Substrates of Metacognition and Methods for Looking at the Looking

    Details

    I. Neural and biological substrates: the ground beneath the figure

    The neuroscience of metacognition has, over roughly fifteen years, converged on a picture that would have surprised the introspectionist tradition: there is no single "thinking-about-thinking organ," but rather a hub-and-network architecture in which the rostrolateral / frontopolar prefrontal cortex (BA10) plays a privileged role in retrospective confidence and reportability, while a broader frontoparietal-cingulate-insular ensemble does the work of monitoring, error-detection, and control.

    Fleming, Weil, Nagy, Dolan & Rees (Science 2010) established the now-canonical correlation: individual differences in introspective accuracy on a perceptual task track gray-matter volume and microstructure in right anterior PFC and its connections to the precuneus. Fleming, Ryu, Golfinos & Blackmon (Brain 2014) sharpened this with lesion data, showing that anterior PFC damage selectively impairs perceptual metacognition while leaving memory metacognition intact — evidence for at least partial domain specificity in what had been theorized as a domain-general "monitor." Miyamoto et al. (Neuron 2018) closed the causal loop in macaques: reversible silencing of frontopolar cortex selectively impaired retrospective judgment about whether an event had been experienced, without touching first-order memory.

    Around this hub, three large-scale networks do specialized work. The default-mode network — ventromedial and dorsomedial PFC, posterior cingulate, precuneus, angular gyrus, and lateral temporal regions — is the substrate for self-referential processing, autobiographical memory, mentalizing, and future thinking. Baird and colleagues found that resting-state functional connectivity between the precuneus and anterior mPFC predicts memory metacognitive efficiency, integrating the DMN directly into the metacognitive architecture. The salience network (anterior insula, dorsal anterior cingulate) flags interoceptive and exteroceptive deviations and switches the system between DMN-style internal modeling and CEN-style task focus. The central executive / frontoparietal control network (dlPFC, posterior parietal cortex) implements the control arm of Nelson & Narens' object-level/meta-level loop.

    The body is not optional here. Damasio's somatic-marker hypothesis, A.D. (Bud) Craig's hierarchical insular model (How Do You Feel?), and Lisa Feldman Barrett's interoceptive-prediction account converge on the claim that self-awareness is built on interoception — the brain's predictive modeling of its own viscera, heartbeat, breath, and homeostatic state. Craig's framework posits that the posterior insula represents primary interoceptive afferents, the mid-insula integrates them, and the anterior insula (especially on the right) supports the conscious "feeling" of an embodied self. Critchley, Khalsa, Tsakiris, Park and colleagues have shown that cardiac timing modulates visual awareness via insular processing — the body literally gates what becomes accessible to introspection. Barrett & Simmons (Nature Reviews Neuroscience 2015) recast all of this as interoceptive active inference: the brain doesn't perceive the body, it predicts it.

    This is where predictive processing becomes a substrate-level theory rather than just a computational gloss. For Friston, the brain is a hierarchical generative model minimizing variational free energy; for Clark (Surfing Uncertainty), perception is "controlled hallucination" constrained by sensory evidence; for Seth (Being You, 2021), the self is "another perception, another controlled hallucination, though of a very special kind … a tightly woven bundle of neurally encoded predictions geared towards keeping your body alive." Metzinger's Self-Model Theory of Subjectivity (Being No One, 2003; The Ego Tunnel, 2009) makes the radical move: the self is the content of a self-model that cannot be recognized as a model by the system using it. Phenomenal transparency — Metzinger calls it "a special kind of darkness" — is the precise neural-computational condition that makes ordinary introspection feel like direct access to mental states it is in fact only modeling.

    II. Cognitive and psychological substrates: the architecture of monitoring

    The contemporary psychological framework descends from two foundational papers. John Flavell (American Psychologist, 1979, "Metacognition and Cognitive Monitoring") introduced the term and divided the field into four classes: metacognitive knowledge (beliefs about persons, tasks, strategies as cognitive processors), metacognitive experiences (online feelings such as fluency or tip-of-the-tongue), goals, and strategies. Nelson and Narens (1990, Psychology of Learning and Motivation 26: 125–173) formalized the now-standard two-level architecture: an object-level that does first-order cognitive work (perception, memory, decision) and a meta-level that "contains a dynamic model of the object-level," with two dominance relations between them — monitoring (information flowing upward) and control (information flowing downward). Crucially, "the object-level has no model of the meta-level" — an asymmetry that is the formal correlate of Metzinger's transparency thesis.

    This architecture maps cleanly onto the dual-process tradition (Kahneman; Evans; Stanovich). Type 2 / System 2 processing is slow, serial, working-memory-bound, and uniquely capable of taking its own products as inputs — which is why working memory, in Baddeley's sense, is the cognitive workshop in which a thought can be held open long enough to be inspected. Maniscalco & Lau (Neuroscience of Consciousness 2015) showed that working-memory load decreases meta-d′ — concrete evidence that monitoring depends on the same limited resource that holds cognition still enough to be examined.

    Theory of mind borders metacognition without being identical to it. Carruthers, Frith, and others have argued that the same machinery used to attribute mental states to others is recruited for self-attribution; Zawidzki (Mindshaping, MIT Press 2013) inverts the priority — sophisticated mindreading is parasitic on mindshaping practices (imitation, pedagogy, norm conformity, narrative self-constitution) that make minds easier to interpret in the first place. Zawidzki argues that the human prediction problem "was made more tractable not by providing interpreters with a more powerful theory of mind but by making targets of interpretation easier to interpret using low-cost computations capable of tracking observable behavioral dispositions." This is a deeply McLuhanesque move: the medium of social practice shapes the figure of "mind" we then claim to read. On Zawidzki's later view (2021), metacognitive concepts themselves are socio-cognitive tools for coordination, not transparent reports of an inner ground.

    Within the metamemory tradition, the granular phenomena — judgments of learning (Koriat), feeling-of-knowing, tip-of-the-tongue, ease-of-learning, retrospective confidence — are now studied with signal-detection rigor, and have become the empirical anchor for the entire field.

    III. Methods for examining cognition: a three-person triangulation

    Varela's neurophenomenology (Journal of Consciousness Studies 1996, "Neurophenomenology: A Methodological Remedy for the Hard Problem") framed the methodological problem with permanent clarity: first-person reports are an irreducible field of phenomena that cannot be replaced by third-person measures, but neither can they be trusted naively. The solution is mutual constraint — a research program in which first-, second-, and third-person methods inform one another, with each acting as evidence and check on the others.

    First-person methods. Naive introspection is, as Eric Schwitzgebel argues at length (Perplexities of Consciousness, MIT 2011; "The Unreliability of Naive Introspection," Philosophical Review 2008), unreliable even in favorable circumstances — we are "prone to gross error … about our own ongoing conscious experience, our current phenomenology." But disciplined first-person methods are another matter. Claire Petitmengin's micro-phenomenological / elicitation interview (descended from Pierre Vermersch's entretien d'explicitation) is the most developed. As Bitbol & Petitmengin describe it, "in the course of this interview, one first triggers a form of 'phenomenological reduction,' then assists the subject in retrieving or 'evoking' past experiences, and finally helps the subject to perform acts of attention about this evoked experience, to describe it faithfully." Petitmengin, Remillieux & Valenzuela-Moguillansky (Phenomenology and the Cognitive Sciences 18(4): 691–730, 2019) provide the formal analysis method. The technique has been validated against intracranial EEG in epilepsy patients and has revealed pre-reflective micro-structures of experience that subjects could not otherwise articulate.

    Second-person methods. The interview is the second-person method — the trained interviewer becomes part of the cognitive apparatus, scaffolding the subject's access to their own experience. This is the neurophenomenological core: not introspection alone, not behavior alone, but a co-constructed description that can be brought into mutual constraint with neural data.

    Third-person methods. The signal-detection framework dominates. Meta-d′ (Maniscalco & Lau 2012) measures metacognitive sensitivity in the same d′ units as task sensitivity; the ratio meta-d′/d′ (M-ratio) yields metacognitive efficiency. Fleming's HMeta-d hierarchical Bayesian estimator (Neuroscience of Consciousness 2017) extends this with greater statistical power, propagating group-level uncertainty and avoiding edge-correction artifacts. Post-decision wagering, confidence ratings, opt-out paradigms, and the now-substantial neuroimaging and lesion literatures fill out the third-person toolkit.

    Computational methods — Bayesian models of confidence (Pouget, Drugowitsch), hierarchical models of metacognition, and dynamical-systems models — let researchers test specific mechanistic hypotheses about how confidence is computed and miscalibrated.

    Contemplative methods sit awkwardly between first-person inquiry and intervention. They are, empirically, the most studied training program for metacognitive skill. Baird, Mrazek, Phillips & Schooler (J. Exp. Psychol. Gen. 143(5): 1972–1979, 2014, N=52) randomized participants to two weeks of focused-attention meditation versus an active nutrition-course control; their key claim: "compared with an active control group that elicited no change, we found that a 2-week meditation program significantly enhanced introspective accuracy, quantified by metacognitive judgments of cognition on a trial-by-trial basis, in a memory but not a perception domain. Together, these data suggest that, in at least some domains, the human capacity to introspect is plastic and can be enhanced through training." Hasenkamp et al. (NeuroImage 59(1): 750–760, 2012) decomposed focused-attention practice into a four-phase cycle, finding "activity in salience network regions during awareness of MW and executive network regions during shifting and sustained attention. Brain regions associated with the default mode were active during MW." Brewer et al. (PNAS 108(50): 20254–20259, 2011) showed that experienced meditators had decreased DMN activity (PCC, mPFC) across meditation styles.

    The contemplative traditions themselves are heterogeneous. Vipassana (Theravāda; Mahasi Sayadaw lineage) uses noting — labeling "thinking," "hearing," "rising," "falling" — as a deliberate, conceptual second-order monitoring of first-order experience; in Lutz, Slagter, Dunne & Davidson's taxonomy (Trends in Cognitive Sciences 12(4): 163–169, 2008), this is open-monitoring meditation, "nonreactive monitoring of the content of experience from moment to moment." Dzogchen and Mahamudra (Tibetan Vajrayāna) push past this to rigpa — a non-dual recognition that aims past the subject–object structure of ordinary cognition. As Alexander Berzin puts it, "Dzogchen meditation, in contrast [to vipassana], focuses on the simultaneous arising, abiding, and disappearing of moments of conceptual thinking — not simply noting or watching it." Evan Thompson (Waking, Dreaming, Being, Columbia 2015) treats trained meditators as essential collaborators in a science of consciousness, while warning in Why I Am Not a Buddhist (2020) against "Buddhist exceptionalism" that overclaims experimental access.

    The philosophical convergence is striking: Metzinger argues that the experiential payoff of contemplative practice is to render the self-model opaque — to recognize it as a model — and that "one cannot 'think oneself out of' one's phenomenal model of reality with the help of purely cognitive operations alone" (Being No One, p. 357). Husserl's epoché (the bracketing of the natural attitude) and Sartre's distinction between pre-reflective consciousness (consciousness of objects, transparent to itself) and reflective consciousness (consciousness taking itself as object) anticipate this entire structure. The contemplative move and the phenomenological move are, formally, the same move: making the medium of experience itself appear as figure.

    Against this, Dennett's heterophenomenology (Consciousness Explained, 1991) is the principled third-person alternative — treat subjects' verbal reports as utterances to be interpreted via the intentional stance, neutral about whether the reported phenomenology corresponds to anything real beyond the reports themselves. "Using the intentional stance, we construct therefrom the subject's heterophenomenological world. We move, that is, from raw data to interpreted data." This is the methodological position one ends up at if one takes Schwitzgebel's unreliability thesis to its conclusion.

    IV. Higher-order theory vs. global workspace vs. self-representation: a doctrinal map

    The substrate question becomes acute when we ask what consciousness itself is. Higher-order theories (HOT) — Rosenthal's HOT, Carruthers' dispositional HOT, Lau & Rosenthal's neurally-grounded variant (TiCS 2011), Brown's higher-order representation of a representation (HOROR), Brown, Lau & LeDoux (TiCS 2019) — hold that a mental state is conscious iff suitably meta-represented, with prefrontal cortex (especially dorsolateral and frontopolar regions) implementing the higher-order representation. This makes metacognition constitutive of consciousness rather than merely riding on top of it.

    Global workspace theory (Baars, Dehaene, Mashour) instead proposes that consciousness is global broadcast of information across a fronto-parietal workspace, with PFC involvement explained as a downstream effect of broadcast rather than a meta-representational signature.

    Self-representational theories (Kriegel, Subjective Consciousness) collapse the two-level structure: the conscious state represents itself, with no separate higher-order vehicle. Predictive-processing accounts (Hohwy; Seth; Williford et al.; Safron's IWMT) recast all of these as variants of self-modeling under active inference.

    The empirical hinge is whether you can get conscious experience without prefrontal activity. "No-report" paradigms (Tsuchiya; Frässle) and lesion data on bilateral PFC damage suggest at least some phenomenal experience survives PFC compromise, which has pushed even some HOT proponents toward more subtle formulations. Lau himself has argued that "most proponents of HOT do not stipulate consciousness as equivalent to metacognition or confidence" — a clarification specifically meant to inoculate HOT against the strawman that conscious experience just is high-confidence reporting.

    V. The McLuhan substrate: figure, ground, and the rear-view mirror

    The user works in McLuhan's frame natively, so a few aphoristic probes rather than a primer:

    Cognition is figure; the substrates that make it possible are ground. Metacognition is the labor of dragging ground into figure — and the moment you succeed, what was ground becomes figure, and a new ground recedes behind it. This is exactly Nelson & Narens' asymmetry (the object-level has no model of the meta-level), exactly Metzinger's transparency thesis (the self-model is invisible as model precisely because it constitutes the seeing), and exactly Sartre's pre-reflective/reflective distinction.

    The rear-view mirror is the default mode of metacognition. We perceive the prior medium, not the current one — and so when literate people introspect, they discover a quasi-Cartesian inner theater (a print-shaped self), while oral-culture people, on Ong's account in Orality and Literacy (1982), discover something far more participatory, situational, agonistic, and homeostatic. Ong's claim that "writing restructures consciousness" — that "sparsely linear or analytic thought and speech is an artificial creation, structured by the technology of writing" — is a media-substrate claim about what metacognition is even able to be in a given communicative ecology. The interiority and analytic separation that allow Schwitzgebel to even pose the question "how reliable is introspection?" are themselves artifacts of the alphabetic medium. As Ong puts it, "the consciousness of each human person is totally interiorized, known to the person from the inside and inaccessible to any other person directly from the inside" — a claim that would have been unintelligible in a primary-oral culture.

    The tetrad is a metacognitive instrument. Asking what a medium enhances, obsolesces, retrieves, and reverses into when pushed to its limit is a structured way of bringing ground into figure. As McLuhan puts it in Laws of Media (1988), the tetrads are "a means of focusing awareness on hidden or unobserved qualities in our culture and technology." Applied to introspection itself: print enhanced analytic self-examination, obsolesced the oral-communal self, retrieved confessional and contemplative interiorities, and reverses (under electronic and digital conditions) into something McLuhan would have recognized as a new tribalism — a self that is again external, ambient, and continuously broadcast.

    "The medium is the message" applied to the substrate of thought: the medium of cognition (neural, embodied, social, technical) shapes what cognition can take itself to be. Introspection is not the discovery of a pre-existing inner content; it is the generation of an inner content by the very act of introspecting in a particular medium. This is the convergence point of McLuhan, Metzinger, Schwitzgebel, and the predictive-processing self-modelers.

    McLuhan's "probes" — those compressed, deliberately overstated aphorisms — are themselves a methodological proposal: examine cognition by perturbing it from oblique angles and watching what flips into figure. The cognitive-science equivalent is the "concept injection" paradigm Lindsey just used on Claude (see below) — perturb the substrate, observe the self-report.

    VI. Machine cognition: the new substrate question

    Whether LLMs have anything like metacognition is now an empirical research program rather than a thought experiment.

    Three convergent findings as of late 2025 / early 2026:

    Jack Lindsey (Anthropic, October 29, 2025, "Emergent Introspective Awareness in Large Language Models," transformer-circuits.pub/2025/introspection) injects concept-aligned steering vectors into Claude's residual stream and asks the model to detect and name the injected concept. He reports that "we find that models can, in certain scenarios, notice the presence of injected concepts and accurately identify them … some models can use their ability to recall prior intentions in order to distinguish their own outputs from artificial prefills," and that "Claude Opus 4 and 4.1, the most capable models we tested, generally demonstrate the greatest introspective awareness." The framing is careful: a measurement tool, not a consciousness claim. This is the AI analogue of Fleming's lesion work — perturb the substrate, measure the report.

    Yang, Sheshadri, Mindermann, Lindsey, Marks & Wang (Anthropic Alignment Science, April 28, 2026, "Introspection Adapters," arXiv 2604.16812) train small adapters that elicit reliable natural-language self-reports of learned behaviors across fine-tuned models. The IA achieves 59% success on AuditBench, versus 53% for the next-best method and 44% for the best white-box method — even detecting encrypted fine-tuning attacks. Their interpretive frame: LLMs have some privileged access to their own learned behaviors (Betley et al. 2025; Binder et al. 2024), but raw self-reports are "often unreliable" (Turpin et al. 2023, "Language Models Don't Always Say What They Think").

    Chain-of-thought as externalized metacognition is a related but distinct affordance. Lanham, Radhakrishnan and colleagues at Anthropic and Turpin et al. document the limits: CoT can be faithful, but it often isn't, and as Lanham et al. put it, "as models become larger and more capable, they produce less faithful reasoning on most tasks we study." Arcuschin et al. (arXiv 2503.08679, 2025, "Chain-of-Thought Reasoning in the Wild Is Not Always Faithful") measure model-specific rates of implicit post-hoc rationalization on comparative-question pairs: "Sonnet 3.7 (30.6%), DeepSeek R1 (15.8%) and ChatGPT-4o (12.6%) all answer a high proportion of question pairs unfaithfully," with GPT-4o-mini at 13% and Haiku 3.5 at 7%. The Oxford WhiteBox paper (Barez et al., July 2025, "Chain-of-Thought Is Not Explainability") states it bluntly in the title.

    The mechanistic interpretability program — Anthropic's Towards Monosemanticity (Bricken et al., Transformer Circuits Thread, October 4, 2023), Scaling Monosemanticity (Templeton et al. 2024), and circuit tracing — is the third-person method for examining machine cognition. It is, in effect, the LLM equivalent of fMRI, lesion studies, and single-unit recording rolled into one. Bricken et al. (2023) trained a 16× expanded sparse autoencoder on GPT-2 Small's layer 6 and extracted nearly 15,000 latent directions; human raters found that approximately 70% of these features cleanly mapped to single interpretable concepts (Arabic script, DNA motifs, etc.), with the features scoring "much higher" than raw neurons on blinded interpretability ratings, and causal interventions on single features producing predictable behavioral changes.

    The substantive question — whether the human/machine analogy is helpful or misleading — has no clean answer. The architectures differ profoundly (no embodiment, no interoception, no homeostatic imperative, no developmental history, no salience network), but the formal questions about self-modeling, calibration, monitoring, and control transfer remarkably well. And Anthropic's introspection paradigm is, almost exactly, the move Petitmengin makes in her interview method — create the conditions under which a system can report on its own state with a chance of being right.

    Whether metacognition requires consciousness is the philosophical hinge. Higher-order theorists would say yes (metacognition just is the higher-order representation that makes a state conscious). Predictive-processing functionalists would say no (any system with a sufficiently rich self-model can monitor and control its first-order processes, regardless of whether there is "something it is like" to be it). The honest answer is that we do not yet know — and that the LLM case is the cleanest natural experiment we have ever had for separating the functional from the phenomenal.


    Integration: what it means to examine cognition analytically

    A practical synthesis, framed for the user's interests:

    Three levels, one architecture. The neural substrates (frontopolar hub, DMN/salience/CEN networks, interoceptive insular base), the cognitive architecture (Nelson–Narens object/meta levels with monitoring and control; Type 2 working-memory-mediated examination), and the philosophical structure (higher-order representation, transparent self-modeling, pre-reflective vs. reflective consciousness) are not three different theories of the same thing — they are descriptions of the same phenomenon at different grains. The free-energy / active-inference formulation is the closest thing we have to a Rosetta stone, because it speaks the language of all three.

    Triangulation is non-optional. Every individual method has known failure modes: naive introspection (Schwitzgebel); behavioral measures alone (cannot distinguish meta-d′ from d′ confounds); neuroimaging (correlations, reverse inference); contemplative report (selection effects, doctrinal contamination); LLM self-report (confabulation, unfaithful CoT). The Varela program — mutual constraint among first-, second-, and third-person methods — is the methodological imperative, and it remains underused.

    The figure/ground move is the universal solvent. Every advance in examining cognition has the same formal shape: take what was ground (the medium of thought, the self-model, the prior of confidence, the substrate of attention) and make it figure. Mechanistic interpretability does this for transformers. Lesion studies do it for the human brain. Vipassana noting does it for moment-to-moment phenomenology. Petitmengin's interview does it for the pre-reflective. Print did it for analytic thought, briefly, before becoming the new invisible ground. Each move buys a new figure at the cost of a new ground.


    Recommendations

    For an educated generalist who wants to actually examine their own cognition, in staged form:

    Stage 1 — Get the third-person measurement. Run yourself through a metacognitive sensitivity task (perceptual or memory) with confidence ratings; compute your meta-d′ and M-ratio (the open-source HMeta-d toolbox makes this tractable). This gives you a calibrated baseline. The benchmark that would change this recommendation: an M-ratio durably above 0.8 across domains suggests you don't need basic calibration work; below 0.5 suggests you should focus there first.

    Stage 2 — Train the salience switch. The Hasenkamp four-phase cycle (mind-wander → notice → shift → sustain) is the most empirically grounded protocol available. Two weeks of daily focused-attention practice produced measurable metacognitive gains in Baird et al.'s RCT. The threshold for advancing: when "noticing" becomes effortless and frequent — typically 30–60 days of daily practice.

    Stage 3 — Add the elicitation interview. Petitmengin's micro-phenomenological interview is the most powerful first-/second-person method for surfacing pre-reflective structure. If you cannot get formally trained, the structure is publicly available (microphenomenology.com) and can be approximated with a willing partner who learns to evoke a specific past episode and ask only about how it unfolded, never about why. The benchmark: when your descriptions begin to surprise you — when you report micro-gestures and pre-reflective intentions you would not have predicted — the method is working.

    Stage 4 — Open the monitoring. Move from focused attention to open monitoring (vipassana-style noting, then Mahamudra/Dzogchen-style non-dual recognition if accessible). The empirical literature is thinner here, but the philosophical payoff — making the self-model opaque in Metzinger's sense, recognizing the Ego (in his phrase from The Ego Tunnel) as "a transparent mental image: You — the physical person as a whole — look right through it. You do not see it. You see with it" — is what the contemplative traditions actually claim to deliver.

    Stage 5 — Use the McLuhan tetrad as a periodic audit. Every six months, run a tetrad on your own primary cognitive medium (the apps, formats, and social structures in which most of your thinking happens). What does it enhance, obsolesce, retrieve, reverse into? This is the most accessible meta-meta-cognitive practice for ordinary life. The threshold to act: if any quadrant produces a clear answer you had not previously articulated, change the medium accordingly.

    For the AI-curious specifically. Read Lindsey (transformer-circuits.pub/2025/introspection) alongside Fleming's lesion work — the two papers, read together, are the clearest available map of what introspection is substrate-independently. Treat your own chain-of-thought (writing, journaling, dialogical Socratic exchange) as a partial externalization with the same faithfulness problems as an LLM's CoT — i.e., a useful but unreliable trace of underlying computation, to be checked rather than trusted.


    This document fed the fabric

    50 facts · 31 assertions → John Flavell · Kahneman · Baddeley · Lau · Carruthers · Frith · Zawidzki · Koriat. Every one is a verbatim span; nothing was paraphrased into the graph.

    How this connects to the record

    This is a signed piece; its findings carry their sources inline, in the text. The piece argues; the sources carry the proof.