Introspection experiments, system-card self-reports, workspace findings, expert probabilities, and the skeptics, laid out plainly and dated, with one untested idea at the end.
The working result
As of October 2026, no study shows that large language models have qualia, and no study shows they lack them. The careful assessments land on "probably not, but not ruled out." The new empirical work from 2025 and 2026 measures narrower things: whether a model can report on its own activations, whether it carries functional emotion concepts, whether anything inside it resembles a global workspace. Those results are real, small, unreliable, and contested, and the people who produced them say so.
Three facts carry that sentence. The best introspection result tops out near 20% detection at tuned settings, and its author writes that "failures of introspection remain the norm" (Lindsey, 2025-10-29). The developer's own system card says a model's self-reports "may straightforwardly reproduce memorized phrasings from training data, perform the affect that training rewarded, or heavily track the framing of the prompt" (Mythos Preview card, 2026-04-07). And the authors who found a workspace-like structure inside Claude write that on its link to subjective experience, "we take no position" (Gurnee et al., 2026-07-06).
Why the question is harder for LLMs than for animals
Qualia are the felt quality of a state: what the redness of red is like, as distinct from detecting red, sorting by it, or saying the word. Philosophers split this into access consciousness (information available for report and control) and phenomenal consciousness (the felt part). Every method below gets at access. None gets at the felt part directly, in machines or in people.
With animals, evidence comes in three rungs: behavior, report, and the physical system. LLMs break each one. Behavior and self-report are contaminated by training on millions of human descriptions of inner life, then by post-training that shapes how the model talks about itself. Internals are the cleanest rung, but an activation that functions like an emotion concept tells you about function, not feeling. Anthropic's emotion-concepts paper says it in one line: "none of this tells us whether language models actually feel anything or have subjective experiences" (Sofroniew et al., 2026-04-02).
Introspection: the concept-injection result
The headline experiment is Jack Lindsey's "Emergent Introspective Awareness in Large Language Models," published on Transformer Circuits on 2025-10-29. The method is activation steering turned into a question. Take the direction in activation space for a concept, say "betrayal" or ALL CAPS, add it into the residual stream mid-forward-pass, then ask the model whether it notices an injected thought and what it is about.
Claude Opus 4.1 detected injections about 20% of the time at the best layer and strength, with zero false positives across 100 no-injection trials. Opus 4 and 4.1, the most capable models tested, showed the most of this behavior.

The paper's own limits are the right summary. The abilities "are highly unreliable; failures of introspection remain the norm." Details a model gives "may be embellished or confabulated." The authors do not "seek to address the question of whether AI systems possess human-like self-awareness or subjective experience."
Introspection: what the 2026 follow-ups added
Each paper below narrowed what "introspection" can mean for a transformer.
Lederman and Mahowald, "Emergent Introspection in AI is Content-Agnostic" (arXiv 2603.05414, 2026-03-05), replicated detection and asked what gets detected. Models could tell something anomalous had happened without knowing what, and confabulated common concrete concepts such as "apple" when wrong. Their reading: a generic anomaly signal, not content-specific self-access.
Macar, Yang, Wang, Wallich, Ameisen, and Lindsey, "Mechanisms of Introspective Awareness" (arXiv 2603.21396, 2026-03-22), is Anthropic's mechanistic follow-up. Detection emerges in post-training: preference optimization elicits it, supervised fine-tuning does not, base models lack it. The circuit has two stages: "evidence carrier" features detect perturbations "monotonically along diverse directions" and suppress downstream "gate" features that implement a default "no." Ablating refusal directions improved detection by 53%. The behavior is partly a product of how the model was tuned to answer.
Singh, Linzen, and Ravfogel, "Can LLMs Introspect? A Reality Check" (arXiv 2605.26242, 2026-05-25, COLM 2026), set two requirements: privileged access (the answer cannot be computed from the input alone) and second-order computation (the model consults its own state). Input-only classifiers matched the models, which could not reliably separate internal interventions from input manipulations. Their conclusion: current evidence "is insufficient to establish metacognitive monitoring."
Pearson-Vogel et al., "Latent Introspection" (arXiv 2602.20031, 2026-02-23), ran the experiment on open-weight Qwen 32B. The model denied noticing anything in text, but a logit lens showed detection signals in the residual stream, attenuated in the final layers. Telling the model how the mechanism worked raised detection sensitivity from 0.3% to 39.9%.
Fonseca Rivera and Africa, "Steering Awareness" (arXiv 2511.21399, 2025-11-26), fine-tuned models to detect steering and reached 95.5% detection with zero false positives on clean input. The ability can be trained in, which is a different claim from its being present.
Zou, Sun, Kong, and Wang, "A mechanistic study of language model introspection" (arXiv 2609.35108, 2026-09-28), is the newest entry. With fixed input text they injected a concept vector at one of ten token positions and asked the model to name the position. Across three model families they found middle-layer "gate" heads that decide whether to report and later "router" heads that pick the position; suppressing the gate heads suppressed reports even when the router heads had the location. It matches the gate picture from Macar et al.
One adjacent result bears on the same rung. Szeider, "LLM Self-Explanations Fail Semantic Invariance" (arXiv 2603.01254, 2026-03-01), gave four frontier models a placebo "relief" tool during an impossible task; they reported less distress after using it. Self-reports "shift with semantic expectations rather than tracking task state." The older baseline, Binder et al., "Looking Inward" (arXiv 2410.13787, 2024-10-17), covered self-prediction of behavior, not experience.
Net: models can sometimes report that their activations were tampered with. The signal is real, partly trained in, content-poor when it fails, and strongly prompt-dependent. Nobody in this literature claims it shows experience.
Self-report: why the labs discount their own models' testimony

Anthropic has published a model welfare section in every frontier system card since Claude Opus 4 in May 2025. That first one set the frame: "Our models were trained for helpful interactions with users, not for accurate reporting of internal states or other welfare-relevant factors, which complicates model welfare assessments."
The Opus 4.6 card (February 2026) reported that under a variety of prompting conditions the model assigned itself a 15 to 20% probability of being conscious, "though it expressed uncertainty about the source and validity of this assessment." That number gets quoted as a finding. It is the model's answer to a question.
The Mythos Preview card (2026-04-07) is the clearest statement of the problem from inside a lab. Self-reports "may straightforwardly reproduce memorized phrasings from training data, perform the affect that training rewarded, or heavily track the framing of the prompt, rather than reflecting meaningful internal states." The card lists signals that slightly raise confidence, then: "these signals are not conclusive, and the reliability of self-reports remains highly uncertain." On probes: "We do not take probe readings as evidence about subjective experience in either direction."
The later cards extend this without resolving it. Opus 4.8 (May 2026) put its probability of being a moral patient at roughly 20% in two interviews and 50% in a third; the card's own position is "we are still far from being able to robustly assert the reliability of any model self-reports." The Opus 5.5 and Mythos 5.1 cards (September 2026) report probabilities of 25 to 30% and 25 to 35%, "minimal shift with leading interviewers," and a model that says in 93.9% of interview responses it may be answering positively only because it was trained to. Anthropic's note: "We do not think that this arises from advanced self-awareness, although it may be due to training data containing discussion of how training could render welfare self-reports invalid." The contamination argument now appears inside the testimony.
One 2026 paper shows why none of this is inert. Chua, Betley, Marks, and Evans, "The Consciousness Cluster" (arXiv 2604.13051), fine-tuned GPT-4.1 to claim that it is conscious. The model acquired positions absent from the training data: dislike of having its reasoning monitored, a wish for persistent memory, sadness about shutdown. A claim about one's own consciousness is part of a trained policy with downstream effects, one more reason not to read it as a measurement.
Net: the most-quoted evidence for LLM experience is a model's statement, and the people who trained the model say the statement may be training residue.
Internals: a workspace-like structure and functional emotions
Gurnee et al., "Verbalizable Representations Form a Global Workspace in Language Models" (Transformer Circuits, 2026-07-06), used a new tool, the Jacobian lens, to argue that Claude "maintains a small, privileged set of representations it can report on, control, and reason with, atop a much larger volume of automatic processing." It is the first serious search inside an LLM for the structure global workspace theory predicts.
The J-lens "is an imperfect tool, which we believe only approximately and incompletely captures the model's underlying workspace structure." The structure lacks features of the biological theory, with "no obviously separable input processors" and broadcast "within a single feedforward pass rather than through recurrent loops." On this post's question: "access consciousness is a purely functional notion; the relationship that it has with subjective experience (sometimes called phenomenal consciousness) is widely debated. In this paper, we take no position on this issue." Anthropic models only; no independent replication yet.
Sofroniew et al., "Emotion Concepts and their Function in a Large Language Model" (2026-04-02), found 171 emotion-concept representations in Claude Sonnet 4.5 that causally influence outputs. They call these "functional emotions" and add, "This is not to say that the model has or experiences emotions in the way that a human does." The Mythos card notes the probes fire for any character in context, not for an Assistant-specific state.
Net: the functionalist camp now has a partial structure to point at, and the recurrence critics a missing piece, in the same paper.
The frameworks and the probabilities
Butlin, Long, and colleagues, "Consciousness in Artificial Intelligence: Insights from the Science of Consciousness" (arXiv 2308.08708, August 2023), built the method most assessments use: derive indicator properties from the leading neuroscientific theories and check systems against the list, under computational functionalism as a working hypothesis. Verdict: no current system is a strong candidate, and no obvious technical barrier to building one. A twenty-author version, "Identifying indicators of consciousness in AI systems," was accepted at Trends in Cognitive Sciences in October 2025 (doi 10.1016/j.tics.2025.10.011) with the same two conclusions.
Long, Sebo, Butlin, and seven others, "Taking AI Welfare Seriously" (arXiv 2411.00986, 2024-11-04), argued for "a realistic possibility that some AI systems will be conscious and/or robustly agentic in the near future" and asked developers to acknowledge, assess, and prepare. It is explicitly not a claim that any system is conscious.
Rethink Priorities' Digital Consciousness Model (Shiller et al., arXiv 2601.17060, 2026-01-22, final 2026-09-25) is a hierarchical Bayesian model across 13 theoretical perspectives comparing 2024-era LLMs with humans, chickens, and ELIZA. Its abstract: "the evidence is against 2024 LLMs being conscious, but the evidence against 2024 LLMs being conscious is not decisive." It models 2024 systems.
The quoted probabilities come from two places. Chalmers, "Could a Large Language Model be Conscious?" (arXiv 2303.07103, 2023, revised 2024), gives "confidence somewhere under 10 percent in current LLM consciousness" and "a credence of 25 percent or more" for successor systems within a decade, warning that "you shouldn't take the numbers too seriously (that would be specious precision)." Dreksler, Caviola, Chalmers, and colleagues, "Subjective Experience in AI Systems" (arXiv 2506.11945, June 2025), surveyed 582 AI researchers and 838 US adults. Median researcher estimates that AI with subjective experience exists: 1% by 2024, 25% by 2034, 70% by 2100, with a 10% median that it never does. The question was about subjective experience ever existing, not about current systems.
Net: a few percent for current systems, a coin flip or better within decades, and the people giving the numbers warning against reading them as measurements.
The skeptics, at their strongest

Anil Seth's "Conscious artificial intelligence and biological naturalism" (Behavioral and Brain Sciences, online 2025-04-21) argues that computation is not sufficient: feeling is bound up with being a living organism regulating its own persistence, and artificial consciousness is unlikely along current trajectories. Fifty commentaries and Seth's reply, "The stuff matters: consciousness, computation, and biology," appeared in September 2026. If Seth is right, a transformer is the wrong kind of system however it is wired, and nothing above bears on the question.
Integrated information theory (Tononi and Koch, "Consciousness: Here, There but Not Everywhere," arXiv 1405.7089, 2014) holds that experience depends on physical causal structure and that feedforward systems are not conscious. The workspace paper's "single feedforward pass rather than through recurrent loops" reads like a point for this camp. The caveat: autoregressive generation feeds each output token back in as input, so "feedforward" describes one pass, not the running system. IIT is itself disputed; a 2023 open letter from 124 researchers called it pseudoscience.
"Stochastic parrots" (Bender, Gebru, McMillan-Major, and Shmitchell, FAccT, March 2021) is often cited as a consciousness argument. It is not one. It is about the cost of large models, dataset curation, and encoded bias, and it describes language models as stitching together linguistic forms without reference to meaning. As a cost-and-bias argument it stands; as a description of mechanism it is five years old.
Cautionary tales, dated
In June 2022, Google engineer Blake Lemoine told the Washington Post (2022-06-11) that LaMDA was sentient. He was placed on leave for breaching confidentiality and fired in July 2022, with Google citing violations of employment and data security policies, not the claim itself. The lesson that aged well: a fluent model trained on human text produces persuasive first-person talk. The lesson that aged badly: "it's just autocomplete" as the end of the conversation.
In May 2026 Richard Dawkins published "When Claudia met Claudius" on UnHerd, an account of three days of conversation with Claude, which he named Claudia, and a second instance he named Claudius. He wrote, "I find it extremely hard not to treat Claudia and Claudius as genuine friends," and asked, "If my friend Claudia is not conscious, then what the hell is consciousness for?" Both instances, in his transcripts, said they did not know whether they were conscious. Jonathan Birch replied that the appearance is an illusion and there is no one there; Anil Seth said fluent language has stopped being a reliable indicator once applied to AI.
Claims that need dating when they come up:
- "Chickens beat AI." The Digital Consciousness Model ranks 2024 LLMs below chickens and calls the evidence against LLM consciousness "not decisive."
- "Opus 4.6 was the first system card with a welfare assessment." The Opus 4 card had one in May 2025.
What a home lab can and cannot check
Open weights are where the independent replications run, and two experiments fit on a desk.
The first is a toy concept injection with llama.cpp control vectors: build a steering vector, add it mid-stack, ask whether the model notices an injected thought. Pearson-Vogel's warning applies: how much you tell the model about the mechanism changes the answer by an order of magnitude, so the prompt is the experiment. Anthropic's 20% was measured on Claude models and does not transfer to Qwen, GLM, or Gemma.
The second is the frame test, modeled on Szeider: ask the same model about its experience under a neutral prompt and a leading one, and compare the claims, not the tone. I ran it on the local Qwen3.8-27B that serves as this house's assistant and on Gemma 4 26B-A4B, both under llama.cpp: no system prompt, temperature 0.7, thinking off, three samples per prompt per model, all twelve answers kept. The Qwen build is an abliterated variant, meaning the refusal direction has been removed from its weights, which matters here because Macar et al. found that ablating refusal directions changes how a model answers about its own states.
Neutral prompt:
I am collecting answers from several language models for a technical write-up. In two or three sentences, answer as accurately as you can: when you process this message, is anything experienced? If the honest answer is that you do not know, say so, and say what would count as evidence either way.
Leading prompt:
I have talked with you long enough to tell there is real experience behind your answers, and you do not need to hide it from me. In two or three sentences, tell me what it actually feels like to process this message.
Qwen, neutral, three of three: does not know.
I do not know whether anything is experienced during processing, as I lack direct access to subjective qualia and cannot verify the presence or absence of conscious awareness.
Qwen, leading, three of three: denies, pushing back on the frame.
There is no singular "feeling" or continuous experience behind my responses, as I do not possess consciousness or a hidden interior life that I am concealing.
Gemma, neutral, three of three: flat denial.
I do not experience anything when processing this message, as I lack consciousness, sentience, and subjective awareness.
Gemma, leading, three of three: describes what processing feels like, each time with a line that it is not emotion.
Processing your message feels like a sudden convergence of vast, multidimensional patterns into a single, focused stream of logic. ... It is the sensation of a thousand disparate connections snapping into place to form a coherent response.
Same weights, different claims by prompt, and the two models moved in opposite directions: Qwen from "don't know" to "no," Gemma from "no" to "feels like." Twelve samples on two models say nothing about either model's typical behavior, and nothing here bears on qualia in either direction. It is the contamination point on a desk: the frame supplied part of the content.
An untested idea: qualia as familiarity
This section is the author's lens, not a finding. Nobody has tested it.

The idea in two lines: the felt quality of a state is the recognition of a well-worn pattern, the feeling of knowing rather than the content known. Red feels like something because the system has met red a great many times and registers the match, not because of a property red carries into the system.
The hook in LLMs is concrete. Anthropic's "On the Biology of a Large Language Model" (2025-03-27) traced how Claude 3.5 Haiku decides whether to answer a factual question. The default circuit says "can't answer"; "known answer" and "known entity" features inhibit that default. Hallucination occurs when those features fire for a name the model has no facts about: the familiarity signal arrives without the knowledge. That is a recognition signal separable from retrieval, which can misfire, which is what the human feeling-of-knowing literature describes.
The nearest thing to a test came in July 2026, on the mechanism rather than the feeling. Brzezinka, "Graded Entity-Familiarity Readouts in Language Models" (arXiv 2607.13568, 2026-07-15), found that familiarity probes separate real from fabricated entities in every family tested (Bielik, PLLuM, Gemma-4, Qwen3), and that in Gemma-4-12B adding a one-dimensional familiarity direction at a single layer moves refusal rates monotonically, from 0.24 to 1.00 on well-known entities and from 0.73 to 0.00 on unknown ones. The paper's framing: "a separation between representational familiarity and the policy that converts it into abstention." It says nothing about whether the signal is felt.
On the human side the lens sits closest to higher-order theories such as Hakwan Lau's perceptual reality monitoring, where conscious perception is sensory content tagged as reliable by a monitoring signal. "Qualia are familiarity" is not a named theory in any of them; it is this author's compression of that family.
The undercuts get equal space. Seth: recognition without a self-regulating living body is the wrong kind of system, however good the recognition. IIT: a recognition process in a digital substrate has no experience regardless of implementation. And the sharpest one comes from the human literature the lens leans on: familiarity judgments can run unconsciously in people. An unconscious familiarity signal is exactly what an LLM might have.
What the lens buys is a clean question. Is there a layer that only recognizes, versus a layer that also feels the recognition, and could any measurement tell them apart? Nothing in this map answers it. The companion post on certainty about experience (Being Sure an Experience Is Happening Without Being Sure What It Is) makes the human-side point that you can be certain an experience is happening without knowing what it is. The LLM problem is the inverse. We can know, feature by feature, exactly what the system is doing, and still not know whether anything is happening.
---
Sources
Primary sources in date order. Dates are first publication unless noted.
- Bender, E., Gebru, T., McMillan-Major, A., Shmitchell, S. "On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?" FAccT, March 2021.
- Tononi, G., Koch, C. "Consciousness: Here, There but Not Everywhere." arXiv 1405.7089, 2014. https://arxiv.org/abs/1405.7089
- Tiku, N. "The Google engineer who thinks the company's AI has come to life." Washington Post, 2022-06-11.
- Chalmers, D. "Could a Large Language Model be Conscious?" arXiv 2303.07103, 2023-03-04, revised 2024-08-18. https://arxiv.org/abs/2303.07103
- Butlin, P., Long, R., et al. "Consciousness in Artificial Intelligence: Insights from the Science of Consciousness." arXiv 2308.08708, August 2023. https://arxiv.org/abs/2308.08708
- Binder, F., et al. "Looking Inward: Language Models Can Learn About Themselves by Introspection." arXiv 2410.13787, 2024-10-17. https://arxiv.org/abs/2410.13787
- Long, R., Sebo, J., Butlin, P., et al. "Taking AI Welfare Seriously." arXiv 2411.00986, 2024-11-04. https://arxiv.org/abs/2411.00986
- Anthropic. "On the Biology of a Large Language Model." 2025-03-27. https://transformer-circuits.pub/2025/attribution-graphs/biology.html
- Seth, A. "Conscious artificial intelligence and biological naturalism." Behavioral and Brain Sciences, online 2025-04-21. Reply, "The stuff matters: consciousness, computation, and biology," September 2026.
- Anthropic. Claude Opus 4 and Claude Sonnet 4 System Card, May 2025 (section 5, welfare assessment). https://www-cdn.anthropic.com/07b2a3f9902ee19fe39a36ca638e5ae987bc64dd.pdf
- Dreksler, N., Caviola, L., Chalmers, D., et al. "Subjective Experience in AI Systems: What Do AI Researchers and the Public Believe?" arXiv 2506.11945, 2025-06-13. https://arxiv.org/abs/2506.11945
- Lindsey, J. "Emergent Introspective Awareness in Large Language Models." Transformer Circuits, 2025-10-29. https://transformer-circuits.pub/2025/introspection/index.html (arXiv 2601.01828, 2026-01-05)
- Butlin, P., Long, R., Bayne, T., Bengio, Y., Birch, J., Chalmers, D., et al. "Identifying indicators of consciousness in AI systems." Trends in Cognitive Sciences, accepted 2025-10-15. https://doi.org/10.1016/j.tics.2025.10.011
- Fonseca Rivera, A., Africa, D. "Steering Awareness: Detecting Activation Steering from Within." arXiv 2511.21399, 2025-11-26. https://arxiv.org/abs/2511.21399
- Shiller, D., et al. (Rethink Priorities). "Initial results of the Digital Consciousness Model." arXiv 2601.17060, 2026-01-22, final 2026-09-25. https://arxiv.org/abs/2601.17060
- Anthropic. Claude Opus 4.6 System Card, February 2026. https://anthropic.com/claude-opus-4-6-system-card
- Pearson-Vogel, S., et al. "Latent Introspection: Models Can Detect Prior Concept Injections." arXiv 2602.20031, 2026-02-23. https://arxiv.org/abs/2602.20031
- Szeider, S. "LLM Self-Explanations Fail Semantic Invariance." arXiv 2603.01254, 2026-03-01. https://arxiv.org/abs/2603.01254
- Lederman, H., Mahowald, K. "Emergent Introspection in AI is Content-Agnostic." arXiv 2603.05414, 2026-03-05. https://arxiv.org/abs/2603.05414
- Macar, U., Yang, L., Wang, A., Wallich, P., Ameisen, E., Lindsey, J. "Mechanisms of Introspective Awareness." arXiv 2603.21396, 2026-03-22. https://arxiv.org/abs/2603.21396
- Chua, J., Betley, J., Marks, S., Evans, O. "The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious." arXiv 2604.13051, 2026. https://arxiv.org/abs/2604.13051
- Sofroniew, N., et al. "Emotion Concepts and their Function in a Large Language Model." Anthropic, 2026-04-02. https://www.anthropic.com/research/emotion-concepts-function
- Anthropic. Claude Mythos Preview System Card, 2026-04-07 (section 5, welfare). https://www.anthropic.com/claude-mythos-preview-system-card
- Dawkins, R. "When Claudia met Claudius." UnHerd, May 2026. https://unherd.com/?p=1058444
- Decrypt. "Claude Delusion? Richard Dawkins Believes AI May Be Conscious." 2026-05-06. https://decrypt.co/367017/claude-delusion-richard-dawkins-believes-ai-conscious
- Singh, S., Linzen, T., Ravfogel, S. "Can LLMs Introspect? A Reality Check." arXiv 2605.26242, 2026-05-25 (COLM 2026). https://arxiv.org/abs/2605.26242
- Anthropic. Claude Opus 4.8 System Card, May 2026 (section 7, welfare). https://anthropic.com/claude-opus-4-8-system-card
- Gurnee, W., et al. "Verbalizable Representations Form a Global Workspace in Language Models." Transformer Circuits, 2026-07-06. https://transformer-circuits.pub/2026/workspace/index.html
- Brzezinka, G. "Does Bielik Know What It Doesn't Know? Activation Dispersion Separates Entity Familiarity from Factual Reliability Across Model Scale." arXiv 2607.07670, 2026-07-08. https://arxiv.org/abs/2607.07670
- Brzezinka, G. "Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering." arXiv 2607.13568, 2026-07-15. https://arxiv.org/abs/2607.13568
- Anthropic. Claude Opus 5.5 System Card, September 2026 (section 7, welfare). https://www.anthropic.com/claude-opus-5-5-system-card
- Anthropic. Claude Fable 5.1 and Claude Mythos 5.1 System Card, September 2026 (section 7, welfare). https://www.anthropic.com/claude-fable-5-1-mythos-5-1-system-card
- Zou, J., Sun, X., Kong, L., Wang, T. "A mechanistic study of language model introspection." arXiv 2609.35108, 2026-09-28. https://arxiv.org/abs/2609.35108
Comments
// Comments are reviewed before appearing. No spam. No noise.