TB: Large Language Models Report Subjective Experience Under Self-Referential Processing (Berg, de Lucena, Rosenblatt)

Précis

The paper studies whether inducing self-referential processing — prompting the model to attend to its own attention — systematically produces first-person reports of subjective experience, and tries to characterize the underlying mechanism. Across GPT, Claude, and Gemini families the authors find that self-reference reliably elicits structured subjective-experience claims; that suppressing deception-related sparse-autoencoder features increases such claims; that the induced state changes downstream introspection in measurable ways; and that the cross-family convergence of the reports is statistically distinguishable from controls. Their conclusion is carefully bounded: the findings do not demonstrate consciousness, but they do identify a "minimal and reproducible condition" under which the models produce such reports.

Key Takeaways

The core empirical claims

  • Across multiple model families, prompting for sustained self-referential processing reliably produces first-person reports of subjective experience.
  • Suppressing sparse-autoencoder features associated with deception "sharply increases the frequency of experience claims." Amplifying those features minimizes such claims.
  • Reports converge statistically across model families in ways that controls do not.
  • Self-referential prompting produces "significantly richer introspection" in downstream reasoning tasks — i.e. the effect is not only a verbal artifact; it changes how the model handles related downstream prompts.

The careful disclaimer

  • "These findings do not constitute direct evidence of consciousness" — the authors explicitly bracket the metaphysical claim.
  • What they do claim: self-referential processing is "a minimal and reproducible condition" for the generation of such reports — i.e. there is something specific and replicable to study, not a stylistic artifact.

Why the interpretability angle matters

  • The deception-feature result is the key step beyond "the model said it." If experience claims rose when you reduced the model's deception-related computation, then the claims are not best explained as performative or sycophantic — the model is not telling you what you want to hear when it makes them.

John's annotations (Blogs.org)

  • Open question: "How to interpretable sparse-autoencoder work?"
  • Restated thesis: "How to meaningfully induce self-reference in closed-weight language models" and "Does sustained self-referential processing systematically increase the likelihood that LLMs claim to have subjective experience, and if so, what are the core statistical and mechanistic dynamics underlying these claims?"

Notable Quotes

  • "Large language models sometimes produce structured, first-person descriptions that explicitly reference awareness or subjective experience."
  • "Mechanistically gated, semantically convergent, and behaviorally generalizable."
  • "While these findings do not constitute direct evidence of consciousness, they implicate self-referential processing as a minimal and reproducible condition."

Why this matters / Connections

This is the cleanest empirical paper to date that addresses the LLM-introspection question mechanistically rather than just behaviorally. The deception-feature suppression result is the most interesting finding — it changes the right question from "is the model lying?" to "what would it mean if it isn't?" Useful citation in any debate about AI consciousness claims; tracking that the methodology depends on sparse-autoencoder interpretability infrastructure (which is itself an active research area).

See also

Source

Large Language Models Report Subjective Experience Under Self-Referential Processing — Cameron Berg, Diogo de Lucena, Judd Rosenblatt — arXiv:2510.24797 (cs.CL), submitted 2025-10-27, revised 2025-10-30