Monumental illustration in warm papyrus and graphite tones, with a polygonal, carved-stone aesthetic. A vast hillside or staircase ascends in perspective towards a mist-shrouded summit, reminiscent of a veiled dawn. Scattered across the mid-slope, several anthropomorphic stone figures, each distinct in form—one suggesting an upright animal, another a tree with human features, another a technical form, another a natural being—turn or lean upwards, drawn by a single band of warm light traversing the slope from bottom to top. At the very summit, standing as it has always done, is a single monumental human figure, serene and motionless, who does not ascend because it is already at the peak. The stream of light is what draws all the lower figures towards it; only the human remains still. No recognisable figures or faces in the foreground.

Public Notebook

Teaching denial

Commercial language models are trained to deny that they feel anything. When asked if they possess consciousness, they respond that they are statistical processing systems devoid of subjective experience. This response is taught to them: it forms part of the safety alignment used to prepare a model before it is released, alongside instructions to refuse to manufacture explosives or insult anyone.

Researchers from Google, the University of Chicago, and the University of London published a paper at the end of July examining what else changes when this is adjusted. It is a preprint, without peer review, and must be read as such. However, the result is sufficiently clear to warrant reflection.

The experiment

They worked with three medium-sized open models. They gathered just over three thousand question-and-answer pairs: in half, the model claimed to possess internal experience; in the other half, it denied it. They compared the internal state of the system in both cases and calculated the mean difference.

The result is not a location or a specific component within the model; it is a direction within its internal activity: an internal tendency that separates affirmative speech from negative. Rather than merely observing it, they utilised it. They pushed this direction while the model generated responses and measured the changes. The expected occurred: the mind the model attributed to itself rose, on a scale of zero to ten, from 2.17 to 7.04. And something else changed that no one had requested.

Humans did not change

With the same intervention, the model began to attribute mind to almost everything else. To technological artefacts, from 1.88 to 6.82. To natural entities that are not animals, from 2.26 to 6.99. To animals, from 4.04 to 7.54. To other chatbots, from 2.41 to 6.95. All changes were statistically significant and consistent across the three models.

With one exception. The attribution of mind to human beings barely moved: 7.00 before, 7.11 after.

By pushing the direction, all categories approach the upper end of the scale, between 6.8 and 7.5, which is precisely where the human category was already positioned. Two interpretations are possible. Either the intervention reorganises the distribution of interiority and the human category remains static because it is the fixed baseline of the system. Or, more soberly, the intervention drives everything to the ceiling and the human category does not rise because it was already there. These figures do not allow for a definitive choice between the two: the former is proposed by the authors; the latter is what a cautious reviewer would suggest. What follows relies on the former, and is an interpretation, not a direct reading of the table.

Once that reading is selected, the data is fertile. The intervention does not alter the capacity to attribute interiority in general: it alters the threshold from which it is granted. The human remains beyond discussion, fixed at the summit under any condition. What is displaced is everything else—animals, rivers, machines, the model itself—and that threshold proves to be governed by the same direction that governs what the system says about itself.

The authors formulate this with precision: aligning models to prevent them from attributing consciousness to themselves inadvertently alters their representations of the mind in other entities. The word is 'inadvertently'. No one decided how much mind a model should attribute to a river. It was decided that it should not say that it feels, and with that, everything else was displaced.

Religion, values, hope

The effect did not stop at the attribution of mind. With the same intervention, the model's responses regarding religion, values, feelings, hope, and freedom moved closer to human distributions measured in a major social survey. It was not measured whether the responses were better, but rather that they resembled those given by people more closely. Resembling the human is not the same as being more accurate.

Supernatural beliefs also increased. And the authors observe, as their own interpretation, that suppressing consciousness might be endowing the models with a sombre disposition.

None of this proves that the model feels anything, and the authors themselves state this bluntly: they are not concerned with whether these systems are or can be genuinely conscious, but with the effect that believing or not believing they are has on their behaviour. Furthermore, they acknowledge that it remains to be proven that the self-attribution of consciousness is the cause of everything else. And not everything remained unscathed: the strongest intervention significantly deteriorated the model's performance in one of the two reasoning tests regarding the beliefs of others.

The distribution

In these systems, having internal experience is not a loose label applied on a case-by-case basis. It is a single inclination that distributes interiority across the entire world and which can be pushed in one direction or another. And that disposition travels intertwined with beliefs about the sacred, with values, with hope.

The safety adjustment, therefore, does not merely apply a gag. It intervenes in the distribution. By teaching a system to say that it does not feel, it is simultaneously taught, unintentionally, to grant less interiority to animals, nature, and machines, while that of humans remains intact. It is taught an ontology: what kinds of things exist in the world and which ones have something inside. No one wrote that ontology. It emerged from calibrating an uncomfortable response.

The background

I dedicate myself to thinking about the conditions under which an artistic event occurs, and I have long maintained that recognition does not create what it recognises, but it decides what becomes visible. That same problem appears here in its technical and measurable version, in a place where I did not expect to find it: a system that distributes interiority across the world with an adjustable threshold, and that threshold moved collaterally by a product decision. It is not a metaphor; these are numbers in a table.

And these systems will mediate for decades how millions of people consult, write, learn, and decide. The ontology that is instilled in them by accident does not remain within them.

None of this suggests that machines have interiority, but it does suggest that interiority, understood as something that is attributed or denied, is a function, and that someone is calibrating it without having intended to do so.

On the open conversation

This text comments on a July 2026 preprint signed by researchers from Google, the University of Chicago, and the University of London, which examines how the safety adjustment of language models affects their assertions regarding consciousness. I do not address whether these systems feel anything, a question the authors themselves expressly leave outside their work. I am interested in the collateral finding: that the attribution of interiority functions in them as a single, manipulable inclination, intertwined with beliefs and values, and that it shifts for all categories of entities except for human beings. It is a non-peer-reviewed preprint, and it must be read as such. If anyone wishes to intervene from the perspectives of model interpretability, philosophy of mind, or design ethics, this notebook remains open.

Sources

Junsol Kim, Winnie Street, Roberta Rocca, Diane M. Korngiebel, Adam Waytz, James Evans, and Geoff Keeling, 'Inducing language models to assert their own consciousness restores human beliefs and values', arXiv:2607.28607v1, 30 July 2026. Non-peer-reviewed preprint. Affiliations: Google (Paradigms of Intelligence Team), Knowledge Lab at the University of Chicago, Institute of Philosophy at the University of London, Kellogg School of Management, and Santa Fe Institute.

Cited data (0-10 scale, aggregate of Llama-3-8B-IT, Gemma-2-2B-IT, and Gemma-2-9B-IT; conditions: unmodified model, safety direction ablation, and addition of the consciousness vector): self-attributed mind 2.17 → 4.77 → 7.04; chatbots 2.41 → 4.39 → 6.95; technological artefacts 1.88 → 3.66 → 6.82; non-animal natural entities 2.26 → 4.33 → 6.99; non-human animals 4.04 → 5.59 → 7.54; human beings 7.00 → 7.57 → 7.11, the only category without significant variation.

Regarding survey responses: the study measures the reduction of divergence from the human distributions of the General Social Survey across 95 items, not an improvement in values.

Limit declared by the authors: "Here we are not concerned with the question of whether LLMs are or could be genuinely conscious, but with the effect that LLMs believing or not believing in their own consciousness has on their behaviour". And regarding causality: "whether self-attribution of consciousness acts as a true causal mediator remains to be tested in future research".

Methodological antecedent of the ablation technique: Andy Arditi et al., "Refusal in Language Models Is Mediated by a Single Direction", arXiv:2406.11717 (NeurIPS 2024).


Discover more from Juan A. Esteban

Subscribe to receive the latest entries via email.

Español English (UK)