Illustration 'The box that does not read itself': a person seen from behind observes a large abstract painting composed of overlapping triangles and circles with a luminous void in the centre.

Public notebook

The box that does not read itself

On the limits of interpretability and the opacity of the machine relevance regime

This line of enquiry has maintained, across four texts, a single thesis: aesthetic comprehension rests upon regimes of relevance, and the relevance regime of a machine differs from that of a human. I demonstrated that this divergence is measurable (in attention), that the core of human aesthetic judgement—self-relevance—requires a self that the machine does not possess, and that what human and machine can indeed share is the process, not the meaning.

One final point remains, and it is the most uncomfortable. Throughout the series, I assumed that we could at least describe the machine's relevance regime from the outside: to state where it looks, what it prioritises, what it weights. This text examines to what extent that is true. And the answer is harsher than it appears: the relevance regime of a system not only differs from our own, but is opaque even to those who possess the entire model—its weights, its activations, its internal mechanisms. The machine is not a black box only to us. It is, in a precise sense that warrants development, a box that does not read itself.

This is not a mystical assertion regarding the essential inscrutability of artificial intelligence. It is a technical conclusion, supported by recent research into interpretability, and it should be argued with care so as not to descend into the grandiloquence that the subject invites.

1. The Myth of Total Access

There is a widespread intuition, reasonable at first glance: if we have complete access to a model—every weight, every connection, every activation—then, in principle, we could know why it does what it does. The opacity of current systems would be a transient problem, resolvable once interpretability tools mature sufficiently.

This intuition is, in large part, false, and research from recent years has demonstrated this on two distinct fronts. The first, older, examines whether a model's attention mechanisms—the parts we would most naturally interpret as 'where it looks'—serve as a reliable explanation of its decisions. The second, more recent and profound, examines how concepts are represented internally, and finds a structural reason why reading those representations is intrinsically difficult.

It is worth exploring both, as together they sustain the thesis without the need for exaggeration.

2. Attention is not Explanation

The first front was born from a concrete question in language processing: do a model's attention weights tell us which parts of the input were important for its decision?

Jain and Wallace (2019), in a work with a deliberately provocative title—Attention is not Explanation—demonstrated two uncomfortable facts. First, that attention weights often correlate weakly with other measures of feature importance: what the attention mechanism highlights does not coincide with what, by other methods, proves decisive for the output. Second, and more significantly, that it is possible to construct entirely different attention distributions that produce the exact same prediction. If two opposing attention maps yield the same result, then the map cannot be the explanation for the result. It is, at best, one of many configurations compatible with it.

Serrano and Smith (2019) reached a convergent conclusion from another angle: attention magnitudes are, at best, noisy indicators of the importance of elements for prediction. They are not devoid of relation to what matters, but they are far from being a faithful reading.

The debate did not conclude in total negation, and it is important to state this to avoid caricature. Wiegreffe and Pinter (2019) responded with a paper titled, also deliberately, Attention is not not Explanation: they argued that the explanatory value of attention depends on how one defines 'explanation' and how the experiment is designed, and that dismissing it entirely was excessive. The balanced conclusion of the debate is not 'attention says nothing', but something more nuanced: the visible internal mechanisms of a model do not automatically equate to faithful explanations of their causal relevance. They may provide guidance; they cannot provide certification.

A note of caution regarding scope is advisable. These three works originate from the 2019 debate on language models and should not be mechanically extrapolated to the entire universe of contemporary systems. However, their underlying lesson holds: observing where an attention mechanism concentrates is not the same as knowing what weighed causally in a decision. And if this holds for attention—the most legible mechanism of all—it holds with greater force for the rest.

3. The Structural Reason: Superposition

The second front is more profound, as it does not examine a specific mechanism, but rather how concepts are represented within the model. And it offers a structural reason why internal reading is intrinsically difficult, not merely technically inconvenient.

The work of Elhage and colleagues (2022) on superposition formulates this with precision. Deep models must represent far more distinct concepts than there are physical dimensions available in their layers. To achieve this, they compress information using superimposed representations: each internal dimension does not encode a unique concept, but rather a combination of many unrelated concepts. The consequence is polysemanticity: a single internal unit—an artificial 'neuron'—activates in response to things as disparate as a visual texture, a fragment of code, and a grammatical feature, without that coincidence signifying anything in understandable terms.

This changes the nature of the problem. If every internal unit blends multiple compressed concepts, then reading a single neuron, a single weight, or a single attention head as if it were the transparent representation of a concept is epistemologically risky. It is not that tools are lacking: it is that the information is not stored in a form that permits direct reading. The model's relevance regime is distributed, compressed, and intertwined across millions of parameters, in a form that lacks a readable centre.

Research into interpretability has advanced precisely by attempting to undo that compression—projecting activations into spaces where concepts are better separated. This is serious and promising work. But even in its best version, what it extracts is a fragmented list of isolated features, not a unified synthesis of 'what mattered to the model and why'. The reason is that there is no mechanism within the system that performs such a synthesis. The model processes; it does not observe itself processing.

4. The Box That Does Not Read Itself

Herein lies the point that gives this text its title, and it should be formulated without sliding into metaphysics.

A current system contains no component that reads its own relevance regime. It computes an output from an input, internally weighting a vast quantity of distributed factors. But there is, nowhere in its architecture, a process that takes that distributed weighting and unifies it into something akin to 'this is what I have considered important, and for these reasons'. The model does not read itself because it has no vantage point from which to do so: it lacks an internal observer, a point of synthesis, an instance that integrates its own states into a representation of what it is doing.

This carries a dual consequence for the thesis of the entire series. The first we were already aware of: the machine’s regime of relevance differs from that of the human. The second is new and more radical: that regime is opaque even when one has total access to the system. It not only fails to share our priorities; it does not present its own in any legible form, neither to us nor to itself. It diverges and, furthermore, it conceals itself—not through a desire to conceal, for it possesses none, but due to the manner in which it is constructed.

It is prudent to mark the limit of the argument so as not to exceed it. I do not assert that it is impossible, in principle, to reconstruct aspects of a model’s operation from the outside: mechanistic interpretability does exactly that, with increasing success in bounded problems. I assert something more precise: that there exists, today, no reliable and direct access to “what mattered” in a specific decision, and that there are structural reasons—not merely instrumental ones—why that reading is difficult. The box may be partially reconstructed from the outside with significant effort; what it does not do is read itself and deliver the result to us.

5. The Contrast with the Human Regime

The contrast with the human case illuminates what is at stake, and it is prudent to frame it with care to avoid falling into facile analogies.

When a human judges a work as important, something occurs that the previous text in this line documented: the work enters into a relationship with their memory, their identity, their history. But there is an additional trait, relevant here: the human, to some extent, can account for that relationship. Not completely—a good portion of aesthetic judgment is opaque to us as well, and introspection is notoriously unreliable. But there is a degree of reflective access: we can state, albeit imperfectly and sometimes mistakenly, why something moved us, what memory it activated, what mattered to us about it.

I do not wish to exaggerate this contrast, because human introspection is limited and frequently confabulates reasons. But the difference in degree is real and relevant. The human possesses a point of synthesis—call it a self, call it reflective consciousness—from which their own experience is presented to them, if only partially. The model possesses this not at all. It is not that it has worse introspection than our own: it is that it has none, because it lacks the instance that would make it possible.

This closes the arc of the series. The human regime of relevance is partially accessible to itself; the machine’s is not at all. And that asymmetry is not a technical detail: it is what separates a system that selects the world from some location—a self, however imperfect it may be—from a system that weighs factors without there being anyone at home for whom that weighting signifies anything.

6. What this does not authorise us to conclude

Three closures, for the sake of rigour.

This does not authorise the conclusion that models are useless or that interpretability is a vain endeavour. On the contrary: understanding the limits of internal legibility is a condition for using these systems responsibly, above all when decisions that affect people are delegated to them. Knowing that an attention map is not a faithful explanation is valuable operational knowledge, not a lament.

This does not authorise the conclusion that opacity is eternal or metaphysical. I assert that there are structural reasons why internal reading is difficult today, and that current systems lack self-observation. I do not assert that no future system can possess forms of internal synthesis or reflective access. I describe the current state and its reasons; I do not prophesy the future.

This does not authorise the conclusion that the human is transparent to themselves. Human introspection is limited, biased, and often confabulatory. The difference with the machine is one of degree and structure—the human possesses an imperfect point of synthesis; the machine possesses none—not the difference between total transparency and total opacity. Maintaining that precision is part of the rigour that this line demands.

7. Conclusions

Three conclusions close this text and the series.

The first, technical. The visible internal mechanisms of a model do not equate to faithful explanations of its decisions. Attention is not explanation; internal representations are superimposed and polysemous; there is no direct and reliable access to “what mattered” in a specific decision. Opacity has structural reasons, not merely instrumental ones.

The second, regarding the regime of relevance. That of the machine not only differs from the human, but is opaque even with total access to the system. It diverges and conceals itself, not by will, but by construction. The box does not read itself because it lacks a point of synthesis from which to read itself.

The third, regarding the final asymmetry. The human regime of relevance is partially accessible to itself, however imperfect introspection may be; the machine’s is not at all. That asymmetry is what separates a system that selects the world from a self from a system that weighs factors without anyone for whom that weighting signifies anything.

The box that does not read itself concludes the trajectory of this line, returning the question to its origin. Throughout five texts, I have maintained that aesthetic comprehension is a method of selecting the world, that such selection requires a self, and that the machine operates in its absence. This final text adds the missing piece: the machine not only lacks a self for which anything might matter; it also lacks a self capable of reading what it weighs within itself. Consequently, when we seek a gaze, a taste, or a judgement within the machine, we encounter only our own reflection projected onto a process that returns nothing to us. Aesthetic judgement remains, as far as the evidence allows us to assert, a matter for those who possess a world to select and who, at least in part, can account for their selection. The machine performs the former in its own blind fashion. The latter, as yet, it does not perform at all.

On open conversation

This text concludes the series on regimes of relevance that I initiated with 'What Matters and What Is Relevant' and developed in 'Where the Machine Looks', 'The Self Effect', and 'Processes Without Meaning', within NeuroArt: Cognitive Surplus. The five texts share a single thesis: aesthetic comprehension is a method of selecting the world, it rests upon regimes of relevance, and the distance between the human and the machine regime is not one of degree but of nature.

Should anyone wish to intervene from the perspectives of model interpretability, the philosophy of mind, neuroaesthetics, or artistic practice, this notebook remains open. The line is, for the moment, closed in its initial trajectory; its natural continuation is research, and eventually, an academic work that formalises it.

Sources

Jain, S., & Wallace, B. C. (2019). Attention is not Explanation. Proceedings of NAACL-HLT 2019, 3543–3556. https://doi.org/10.18653/v1/N19-1357

Serrano, S., & Smith, N. A. (2019). Is Attention Interpretable? Proceedings of ACL 2019, 2931–2951. https://doi.org/10.18653/v1/P19-1282

Wiegreffe, S., & Pinter, Y. (2019). Attention is not not Explanation. Proceedings of EMNLP-IJCNLP 2019, 11–20. https://doi.org/10.18653/v1/D19-1002

Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., Grosse, R., McCandlish, S., Kaplan, J., Amodei, D., Wattenberg, M., & Olah, C. (2022). Toy Models of Superposition. arXiv:2209.10652. https://doi.org/10.48550/arXiv.2209.10652

Esteban Ruiz, J. A. What matters and what is relevant. Public notebook at juanesteban.art, 2026.

Esteban Ruiz, J. A. Where the machine looks. Public notebook at juanesteban.art, 2026.

Esteban Ruiz, J. A. The effect of itself. Public notebook at juanesteban.art, 2026.

Esteban Ruiz, J. A. Processes Without Meaning. Public notebook at juanesteban.art, 2026.

Esteban Ruiz, J. A. Art as Structural Surplus: Toward a Relational Ontology Beyond Human Authorship (V2.3). PhilArchive and Zenodo, 2026.


Discover more from Juan A. Esteban

Subscribe to receive the latest entries via email.

Español English (UK)