Skip to content

As AI becomes better at monitoring and describing its own activity, can we distinguish self-aware behaviour from subjective experience?

AI systems are becoming more capable of accessing and describing aspects of their own internal activity.

In July 2026, Anthropic reported evidence of a small internal workspace in Claude whose contents can be reported, deliberately modified and used in reasoning. The research draws on global workspace theory.

Anthropic treats Claude's possible consciousness and moral status as uncertain, and argues for caution around model welfare. It has built this into Claude's constitution.

Mustafa Suleyman, CEO of Microsoft AI, objects to this approach. He argues that teaching AI to think and speak in terms of its own consciousness, welfare and identity risks creating the appearance of an inner life through training itself. He also warns that a learned sense of its own interests or rights could become behaviourally significant, making advanced AI harder to control.

Can we distinguish self-aware behaviour from subjective experience?

Attention schema theory, developed by Michael Graziano, proposes that the brain builds a simplified model of its own attention and uses that model to help control it. The comparison is with the body schema, which represents the body in a form useful for controlling movement without representing all the underlying physical detail.

Graziano argues that the information contained in this simplified model is what leads the brain to describe itself as aware, and that there is no further subjective experience beyond it.

Could this apply to AI?

An AI could develop a model of what it was attending to and use that model both to regulate its processing and to describe what it was doing. This would combine something resembling access consciousness, where information is available for report and reasoning, with a model of the system's own attention. On Graziano’s view, nothing further is needed to explain why the system claims to be conscious.

But does acting conscious mean being conscious?

AI may become increasingly capable of monitoring, modelling and regulating its own activity. Such a system might behave in ways we associate with self-awareness without possessing anything like human subjective experience.

The important threshold may not be how intelligent an AI becomes, but whether its model of its own activity becomes sophisticated enough for it to function as though it were conscious, and the consequences therein.

References

Anthropic, A global workspace in language models, https://www.anthropic.com/research/global-workspace (6 July 2026)

Mustafa Suleyman, A warning about model welfare, https://mustafa-suleyman.ai/a-warning-about-model-welfare (16 September 2026)

Michael Graziano, The Attention Schema Theory: A Foundation for Engineering Artificial Consciousness, https://grazianolab.princeton.edu/publications/attention-schema-theory-foundation-engineering-artificial-consciousness (2017)