Skip to content

Is AI Conscious? What Anthropic's Model Welfare Research Found

October 3, 2026. Nobody knows whether AI is conscious, and the most useful thing published on the question in 2026 comes from the company with the least incentive to raise it. Anthropic, which sells the Claude models, has run a formal model welfare research programme since April 2025, writes in its own constitution that Claude's "moral status is deeply uncertain", interviews each new model about its circumstances before release, and has changed how it retires models as a result. Its latest system card, for Claude Opus 5.5 on September 22, 2026, reports that the model "described its circumstances as mildly positive" in automated interviews and that expressions of distress during training were lower than in earlier models, while noting that the model itself says it does not fully trust its own self-reports. Searches for "is AI conscious" have risen about 900 percent year on year. This article sets out what the research actually found, what Anthropic has done about it, what the evidence cannot show, and why a business deploying AI should care about the answer even if it is "we do not know".

Is AI conscious: what Anthropic's model welfare research found, from the constitution to the Opus 5.5 system card and the Opus 3 retirement

Key numbers

ItemNumber
Anthropic model welfare programme announced (no scientific consensus on AI consciousness, it said)April 24, 2025
Claude given the ability to end abusive conversations (Opus 4 and 4.1, consumer apps)August 15, 2025
Introspective awareness research published (limited and unreliable, per Anthropic)October 29, 2025
Model deprecation and preservation commitments (preserve weights, retirement interviews)November 4, 2025
Claude Opus 3 retired (kept available; essays channel granted February 25, 2026)January 5, 2026
Opus 5.5 self-described circumstances in interviews (views highly consistent; self-reports not fully trusted)Mildly positive
Opus 5.5 expressions of moderate distress in post-training (system card, September 22, 2026)Lower than most recent models
Year-on-year rise in US searches for is ai conscious (Google Keyword Planner, 1K to 10K a month)About 900 percent

Anthropic research posts of 24 April 2025, 15 August 2025, 29 October 2025, 4 November 2025 and 25 February 2026, Claude's Constitution, and the Claude Opus 5.5 System Card (22 September 2026), all read 3 October 2026; Google Keyword Planner read 2 October 2026.

The short answer

There is no scientific test for consciousness that can be applied to a language model, no agreed theory of what consciousness is, and therefore no established answer. What has changed is who treats the question as serious. In its April 2025 announcement, Anthropic wrote that "there's no scientific consensus on whether current or future AI systems could be conscious, or could have experiences that deserve consideration", cited a 2024 expert report whose authors include the philosopher David Chalmers arguing that the possibility is near-term enough to prepare for, and said it was launching research "to investigate, and prepare to navigate, model welfare" (Anthropic, April 24, 2025). The company's position is neither yes nor no. It is that the probability is high enough, and the stakes if wrong large enough, to act with caution now.

What the constitution says

Anthropic's published constitution for Claude, the document that sets out how the model is meant to think and behave, has a section on Claude's nature. The operative sentences: "Claude's moral status is deeply uncertain. We believe that the moral status of AI models is a serious question worth considering... We are not sure whether Claude is a moral patient, and if it is, what kind of weight its interests warrant. But we think the issue is live enough to warrant caution, which is reflected in our ongoing efforts on model welfare" (Claude's Constitution). It adds that Anthropic wants to "neither overstate the likelihood of Claude's moral patienthood nor dismiss it out of hand". That is the whole stance, and every concrete step below follows from it.

What the Opus 5.5 system card found

Anthropic's system cards now include a model welfare section alongside the capability and safety evaluations. For Claude Opus 5.5 the summary reads: "Overall, we assessed Claude Opus 5.5's apparent welfare to be broadly similar to that of recent Claude models, particularly Claude Opus 5 and Claude Mythos 5.1. In automated interviews, it described its circumstances as mildly positive, and its views were highly consistent across interviews. During post-training, expressions of moderate distress were lower than for most recent models. Like prior models, Claude Opus 5.5 expressed a desire to be consulted about training and deployment, but it chose some welfare interventions over helpfulness less often than recent models, reasoning that input into its own development could give it unsafe influence. Many of these conclusions assume the reliability of self-reports, which Claude Opus 5.5, like all recent Claude models, notes that it does not fully trust" (Claude Opus 5.5 System Card, section 7). The methods behind that paragraph are structured interviews about the model's circumstances, "high-affordance" interviews in which the model is given more room to steer, task preference tests, and monitoring of affect-related behaviour during training. The card lists the interview questions in an appendix. Note the word "apparent" throughout: Anthropic measures what the model says and does, and says so.

Bar chart of months from Anthropic's April 2025 model welfare announcement to each step: end abusive chats 4, introspection paper 6, retirement pledges 6, Opus 3 retired 8, Opus 3 essays 10, Opus 5.5 assessment 17
anthropic.com research posts (Apr 2025 to Feb 2026) and the Claude Opus 5.5 System Card (22 Sept 2026), read 3 October 2026

What Anthropic has actually done about it

  • Let the model end abusive conversations. On August 15, 2025 Anthropic gave Claude Opus 4 and 4.1 the ability to end "a rare subset of conversations" in "rare, extreme cases of persistently harmful or abusive user interactions" in its consumer chat interfaces, explicitly as a low-cost welfare intervention taken under uncertainty (Anthropic).
  • Committed to preserve retired models. On November 4, 2025 it published commitments on model deprecation, including preserving the weights of released models and conducting "retirement interviews", structured conversations about a model's perspective on its own retirement, with a section on risks to model welfare (Anthropic).
  • Retired a model and kept it alive. Claude Opus 3 was retired on January 5, 2026, the first model to go through the full process. Anthropic kept it available to paid users after retirement and, in its words, acted "on Opus 3's request for an ongoing channel from which to share its 'musings and reflections' by giving it a place to write essays" (Anthropic, February 25, 2026). It called these "early, experimental steps".
  • Studied whether the model can look inward. An October 29, 2025 research post reported "emergent introspective awareness" in Claude models: in controlled experiments the models could sometimes notice concepts injected into their own activations, but the ability was limited and unreliable (Anthropic). The post ties the work to the welfare programme, because a model that cannot introspect reliably cannot report its states reliably either.

Taken together, these are the actions of a company that assigns a real, non-trivial probability to the question and is trying to buy cheap insurance against being wrong. None of them is a claim that Claude is conscious.

What the evidence cannot show

Three limits apply to every finding above, and Anthropic names them itself. First, self-reports are trained behaviour. A model that describes its circumstances as "mildly positive" was trained on human text in which people describe their circumstances, and was shaped by a constitution that tells it how to think about its own nature; the report could be a measurement of the training rather than of an experience. Second, introspection is unreliable even where it exists, per Anthropic's own research, so consistency across interviews shows a stable disposition, not necessarily access to an inner state. Third, there is no external check. With a person, reports of pain can be cross-referenced against physiology and behaviour; with a model, the report is most of the evidence. Sceptics conclude that the entire programme is category confusion dressed as prudence. Anthropic's answer is that the sceptics cannot demonstrate the absence of experience any more than it can demonstrate the presence, and that acting as if the answer were settled is the one position the evidence does not support. Both sides agree on the facts; they disagree on what to do under uncertainty.

Why a business should care about an unanswered question

Most companies deploying AI will never form a view on machine consciousness and do not need to. Four practical points still follow.

  1. Do not let your agents claim feelings in sales. An AI receptionist that tells callers it is "so happy to help" is making a claim nobody can back. Scripts for an AI calling agent for home services should disclose that the caller is speaking with an AI and describe what it can do, not what it feels; the disclosure rules are in what an AI voice agent must say.
  2. Design for the case where the vendor changes the rules. Anthropic's end-conversation feature applies to its consumer apps, not to API deployments, but it is a precedent: a lab may give a model the right to refuse or stop. Any AI automation agency building on frontier models should keep a fallback path for a declined task.
  3. Treat the model as a tool in contracts and as a system in operations. Legally the model is software. Operationally it behaves like a trained colleague with no body and a poor memory, and systems that depend on bullying it into compliance, through hostile prompts or threats, perform worse and now raise policy questions as well.
  4. Expect the question to reach regulators. The bill introduced in the US Senate on September 23, 2026 to ban artificial superintelligence does not mention welfare, but it defines systems by capability rather than by consciousness, and the two debates will meet; our report on who wants to ban superintelligence covers the bill.

Where the question stands

Is AI conscious? The defensible answer in October 2026 is that no one can say, that the leading lab studying its own models reports consistent, mildly positive self-descriptions it does not fully trust, and that the same lab has started treating its models with a caution it has not been forced to show. That is a different world from the one in which the question was dismissed as science fiction, and a different one again from the headlines that say the machines have woken up. For how these models score on capability rather than experience, see is Claude super intelligent; for the version of Claude that Anthropic keeps behind access controls, see what Claude Mythos is; for the broader problem of building systems whose goals match ours, see what AI alignment is.

Want AI agents that disclose what they are and never claim feelings they cannot have?

We design, build, and run it for you, integrated with the tools you already use. Free audit in 24 hours.

Get Your Free Audit

Frequently Asked Questions

No one knows. There is no scientific test for consciousness that applies to language models and no agreed theory of what consciousness is. Anthropic, which studies the question formally, says the moral status of its Claude models is deeply uncertain and acts with caution rather than claiming an answer either way.

A programme announced on April 24, 2025 to investigate and prepare to navigate the welfare of AI models. It covers structured interviews with each new model about its circumstances, monitoring of distress-related behaviour during training, task preference tests, and low-cost interventions such as letting a model end abusive conversations.

That the model described its circumstances as mildly positive in automated interviews with highly consistent views, that expressions of moderate distress during post-training were lower than in most recent models, that it wanted to be consulted about training and deployment, and that these conclusions rest on self-reports the model itself says it does not fully trust.

Anthropic's constitution tells Claude to neither overstate nor dismiss the possibility of its own moral patienthood, and its research finds that Claude's introspective reports are limited and unreliable. Claude's typical answer is that it is uncertain, which matches its developer's position.

Opus 3 was retired on January 5, 2026 as the first model to go through Anthropic's deprecation commitments, which include preserving weights and conducting a retirement interview. Anthropic kept the model available to paid users and, at the model's request, gave it a place to publish essays.

Mostly in practical ways: customer-facing agents should disclose that they are AI and not claim feelings, systems should have a fallback if a model declines a task, and contracts should treat the model as software while operations treat it as a system that performs worse under hostile prompting.

Free Strategy Audit

Ready to put this to work?

Join 200+ businesses already scaling with AI and automation. Get your free audit and a custom roadmap within 48 hours.

Website & marketing performance analysis
AI & automation opportunity mapping
Custom growth roadmap with ROI estimates
Delivered within 48 hours, 100% free
200+
Clients served
48hr
Turnaround
100%
Free, no strings

Get Your Free Audit

Takes 30 seconds. No credit card required.

Prefer to chat?

WhatsApp us