Compressione Evolutiva

What Survives the Lesion

Switch off a model's internal workspace and fluency holds while reasoning collapses. And the information the system inherits was paid for by us.

Leggi in italiano →

Two columns compared: capabilities that survive the ablation of the workspace on the left, those that collapse on the right.

Most of the coverage of Anthropic's July interpretability result read it as a story about consciousness. That is not the interesting part.

They found a small set of internal patterns inside Claude that behave like a workspace: they carry what the model has in mind without saying it. Then they did the thing that matters. They switched it off.

Without that workspace the model still speaks fluently. It classifies sentiment, answers multiple-choice questions, pulls facts out of a passage, roughly as before. What collapses is multi-step reasoning — to near zero. Summarization and rhyming poetry fall below the level of a much smaller model left intact.

I had argued, arriving by a different road, that a text can come into the world in two ways: compressed or thought. The experiment confirms me and corrects me, and the correction is worth more than the confirmation.

These are not two qualities of the product. They are two routes — and the same output can come out of either. The sharpest demonstration: give the model a passage in Spanish and swap "Spanish" for "French" inside its workspace. Ask what language it is, and it says French. Ask it to continue the passage, and it continues in perfect Spanish. Same knowledge, two paths.

The consequence is blunt: you cannot recover the route from the text. Not by reading it more carefully, not by reading it twice. The signal is not there. I learned this on a problem of model railway track geometry — a confident, fluent, completely wrong answer, caught only by a test outside the text.

Then there is the number that stopped me. That workspace accounts for less than a tenth of internal activity and holds a few dozen concepts at a time. It is not where the mass is. It is where the connection is: components read from and write to those directions on the order of a hundred times more than to ordinary ones. Not a new channel — the channel was always there, and everything passes through it. A few privileged directions inside it. Not an organ. A communication regime.

And nobody designed it. It emerged on its own during training.

Here it meets what I have been writing about for a year.

In biology, for a discovery to persist, something has to die. The unit of variation coincides with the body, selection removes organisms, every cycle costs a generation. In training, variation and selection happen inside a single substrate that persists. Selection has not disappeared — every gradient step eliminates configurations. What has collapsed is its coincidence with the body. The carrier no longer has to die for the information to remain.

But that information did not come from nowhere. It came from our corpus, which is the sediment of billions of deaths already paid. The model did not escape the cost. It inherited it. An heir, not a fugitive.

Which makes the next question less triumphant and more serious. The first generation inherits for free. The second? When the corpus no longer suffices and systems learn from their own products, it is not variation that runs out — new material can be generated without limit. What is missing is something external that separates a discovery from a coherent echo. The debt is not in the variation. It is in the criterion.

Two caveats, because they are needed.

The instrument all of this was measured with is declared imperfect by the authors themselves: it captures only an approximation of the true workspace, and only concepts that fit in a single token. They are ablating their approximation, not "thought".

And compressed does not mean stupid. The model without its workspace goes on doing things that look a great deal like understanding. The compressed is not the poor part. It is the non-deliberate part, and it is the overwhelming majority of what works. In them, and — I suspect — in us.

This essay belongs to a framework I call compressione evolutiva — on the widening gap between slow biological evolution and the rapid acquisition of cognitive capability by machines.
Read the framework →