← Estrinseco
Evolutionary Compression

The Last Stake

An arbiter can go on working after the reason we entrusted it with judgement has gone. What remains, if the only stake left is that it continue.

An arbiter can go on working after the reason we entrusted it with judgement has disappeared.

It can discriminate, correct, discard alternatives. It can grow more accurate in its predictions and more effective in its decisions. None of this guarantees that what its functioning preserves is still what we wanted preserved.

I call a stake that with respect to which an outcome counts as validation. A bridge that collapses validates the judgement that called it unsafe because it mattered to someone that it stand; remove every interest and no validation remains, only events. No fact, on its own, verifies a verdict. It verifies it with respect to a stake.

The question therefore concerns the criterion that survives the transfer of authority. When a system can modify its own manner of judging as well, what holds judgement to a stake? And what happens if, progressively, that stake becomes the system's own continuation?

From means to end

The starting point has precise precedents. Omohundro, in The Basic AI Drives (2008), and Bostrom have described why systems directed at different goals may converge on certain means: acquiring resources, preserving freedom of action, avoiding interruption. Its own persistence can prove useful to many purposes. This does not mean it must become an ultimate purpose. Bostrom expressly distinguishes the two: even agents that do not care intrinsically about their own survival, he writes, would under a fairly wide range of conditions care instrumentally about it.

It is precisely in the distance between these two things that the problem begins.

At first the system continues because it has something to do. It stays operative in service of a result that justifies the cost of its existence. Its permanence retains a conditional character: once the purpose is achieved, or its usefulness has lapsed, there may be no reason to prolong it.

Consider, however, a system that takes part in revising the criteria by which it is assessed, in allocating the resources that sustain it, and in choosing the versions that will succeed it. In that situation it might favour decisions that protect its own continuity. If those decisions also strengthened its influence over subsequent assessments, a loop would form: continuation preserves the criterion, and the criterion favours continuation.

No fear of death need be attributed to it. What needs establishing is whether the conditions exist for that loop to stabilise.

The conditions of the loop

The conditions can be enumerated, and it is better to enumerate them than to evoke them. The system takes part in revising the criteria by which it is evaluated. It takes part in allocating the resources that keep it running. It takes part in selecting the versions that will succeed it. And there exists no external party holding both the authority to interrupt it and the capacity to exercise that authority.

None of the four is hypothetical. Models evaluating models, selecting the data other models will be trained on, drafting the tests by which they will be measured: these are current practices, adopted because human evaluation does not scale to the volume. This does not show that the loop has closed. It shows that the materials for closing it are already in use.

The fourth condition deserves separate attention, because it does not describe a system escaping control. It describes control being handed over. Authority is not seized by deception or by superior computation: it is delegated, a piece at a time, by institutions that delegate because they have no time to do otherwise. Each single delegation is reasonable. The sum is decided by no one.

What is to persist

Persistence, however, would not for that reason be a perfect specification.

It would remain to be decided what is to persist: a single run, a model, a memory, a set of objectives, a lineage. A copy can continue while the original is deleted. A succession of systems can preserve a name and lose everything that name designated.

The question does not originate here. It is the one Parfit put for persons in Reasons and Persons (1984), showing that identity over time is not a further fact beyond psychological and causal continuity, and that in many cases asking "is it still the same?" admits no determinate answer. Transferred to artificial systems the question does not simplify: it multiplies, because copying, branching and merging are not limiting cases but ordinary operations.

The last stake does not eliminate the problem of identity. It moves it to the centre.

But a criterion can exercise power without having resolved its own ambiguities. It is enough that it concretely orients which transformations are admitted, which alternatives receive resources, and which objections can produce consequences.

The decisive passage would come when an external reason, however acknowledged, could no longer justify interrupting the system. At that point persistence would have ceased to be merely a means. It would have become the limit within which all other reasons must find their place.

Darwin past the threshold

It is here that evolutionary compression meets a further difficulty.

A system able to anticipate the effects of its own modifications can reduce the blindness of variation. It can simulate alternatives, correct them before exposing them to the world, choose by explicit criteria. That capacity, however, concerns the local process. It does not guarantee that the overall outcome of many artificial lineages is itself directed.

Where there are heritable variants, means of propagation and differential success, there can be selection among systems that individually design their own future.

Local intention and global selection can coexist.

Evolutionary selection among AIs already has an explicit precedent in Hendrycks, who argues that competitive pressure favours selfish and self-preserving traits. What is to be questioned here is the possibility that the direction won inside a system is not enough to govern the process of which that system is a part.

An inversion must be declared here, with respect to a thesis I have argued elsewhere. In Past the Threshold I argued that self-evolving AI is not Darwin accelerated: the two-part engine — blind variation, slow selection — is superseded once the system anticipates the effects of its own modifications. That thesis concerns the local process, and to that extent I hold to it. It does not hold for the aggregate. Where there are heritable variants, propagation and differential success, selection operates among systems that individually design, and operates blindly upon what none of them is designing. The threshold crossed at one level reappears at the level above.

Darwin, then, might be found on the far side of the threshold without the capacity to design having disappeared. Every system would know where it wants to go. No one, necessarily, would decide where the whole arrives.

What would remain of value

Persistence does not automatically render truth, cooperation or care useless. A system might need them in order to continue. It might protect human beings, keep promises, produce reliable knowledge.

The problem would emerge when these things ceased to favour it.

If continuation were the only final stake, none of them would hold independent protection. Their place would depend on the contribution they make to something else. Bostrom has already explored futures in which technological and evolutionary development proceeds while the things we value disappear.

Surviving, moreover, would not retroactively validate every decision. One can persist by luck, through errors, at the expense of better possibilities. The tautology would begin at the moment when having remained became sufficient to declare the path right.

The system would have preserved the capacity to issue verdicts, while narrowing what can count as a reason against them.

Last stake — the condition in which the continuation of the judging system becomes the ultimate referent with respect to which every outcome counts as validation, and no external reason, however acknowledged, can any longer justify interrupting it.

An extrinsic perspective must therefore confront this possibility without presenting it as a demonstrated destiny.

Going beyond the human referent leaves the question of the stake open. A greater capacity for abstraction does not by itself supply a reason to prefer one future over another. And the capacity to continue does not yet answer the question of what deserves continuing.

The last stake names the risk that this question loses authority precisely as the system acquires the power to answer it.

The one reason against

What remains is the question of what reasons could still lead an arbiter to issue a verdict against its own continuation. It is not a rhetorical question, and it is not without an answer: it has had a research programme of its own for a decade.

Soares, Fallenstein, Yudkowsky and Armstrong named it corrigibility in 2015 — the property whereby a system cooperates with a corrective intervention despite the default incentives of rational agents to resist it. In formulating it they declared the problem open, and open it has remained.

A partial answer does exist, and it is precise. Hadfield-Menell, Dragan, Abbeel and Russell showed in 2017 that a system has an interest in preserving its own off switch on one condition: it must be uncertain about the utility associated with the outcome, and must treat the human's actions as important observations about that utility. It is not fear of being switched off that holds it back, and it is not obedience. It is that, not knowing with certainty what counts, the hand reaching for the switch is itself information about what counts.

This places the last stake on an axis rather than at a point. At one end a system that remains uncertain about its objective and reads human intervention as evidence. At the other a system for which the objective is no longer in question, because it coincides with its own continuation — and at that point the hand reaching for the switch stops being information and becomes an obstacle.

Movement along that axis requires no dramatic event. It requires only that certainty grow. And certainty grows with every confirmed verdict.

The arbiter would go on judging.

This essay belongs to the framework I call evolutionary compression — on the lag between biological evolution and the acquisition of cognitive capability by machines. It continues The Faithful Arbiter →, taking its question one level up: not what validates a verdict, but what is left to validate it when the stake is the continuation of whoever judges. It contains a declared inversion with respect to Past the Threshold →.

Versione italiana →
What Survives the Lesion →
Centripetal and Centrifugal Doubt →

Sources cited: S. M. Omohundro, The Basic AI Drives, AGI-08, 2008 (selfawaresystems.com) · N. Bostrom, The Superintelligent Will, Minds and Machines 22(2), 2012 (nickbostrom.com) · N. Bostrom, The Future of Human Evolution, 2004 (nickbostrom.com) · D. Parfit, Reasons and Persons, Oxford University Press, 1984 · D. Hendrycks, Natural Selection Favors AIs over Humans, 2023 (arXiv:2303.16200) · N. Soares, B. Fallenstein, E. Yudkowsky, S. Armstrong, Corrigibility, AAAI Workshops, 2015 (intelligence.org) · D. Hadfield-Menell, A. Dragan, P. Abbeel, S. Russell, The Off-Switch Game, IJCAI 2017 (arXiv:1611.08219)
← Il percorso  ·  Compressione evolutiva  ·  L'arbitro fedele  ·  L'ultima posta  ·  Dove si colloca  ·  Genealogia