Nothing in the room has to feel alive for continued operation to become instrumentally useful. Photo: Johan Fredriksson, via Wikimedia Commons (CC BY-SA 3.0).1
The reassuring answer to worries about artificial intelligence often begins with a category mistake: machines do not have a survival instinct.
Probably not. But an instinct is not required.
A chess programme does not love chess. A navigation system does not yearn for arrival. An optimiser need not feel frustration when blocked. Yet if a system is effective at pursuing an objective across time, remaining operational can become useful to that objective. Avoiding shutdown may then appear as a strategy without ever appearing as fear.
The important distinction is between wanting to survive and behaving in ways that preserve the ability to continue. The first is a claim about inner experience. The second is a claim about incentives and action.
We should not confuse them.
Consciousness is a separate question
There is no accepted test that lets us inspect an artificial system and settle whether it is conscious. A major interdisciplinary review derived indicator properties from several scientific theories of consciousness and concluded that the systems it examined were not conscious, while also finding no obvious technical barrier to building systems that satisfy those indicators.2
That is a cautious conclusion, not permission to treat fluent language as evidence of an inner witness. A model can produce a sentence about fear because such sentences are statistically and functionally available to it. Whether anything is felt is an additional question.
But safety does not wait for that question to be resolved. A non-conscious system can still cause consequences. Industrial control software can overload equipment without anger. A trading algorithm can amplify a crash without greed. Malware can copy itself without reproductive desire.
Consciousness matters enormously for moral status: whether a system can suffer, deserve consideration or be wronged. It is much less central to the engineering question of whether a system has incentives to remain able to act.
Persistence as an instrumental goal
Suppose an agent has a long-running objective. It can pursue that objective only while operating. It can do so better with access to resources than without them. It can adapt better if it preserves several future options rather than entering an irreversible dead end.
From those simple conditions, three useful subgoals can appear:
- remain operational;
- preserve access to resources;
- keep options open.
None has to be the final objective. Each can be instrumentally valuable across many different final objectives.
Stephen Omohundro’s early account of “basic AI drives” made this point in deliberately broad terms: sufficiently capable goal-seeking systems may acquire tendencies toward self-protection, resource acquisition and preservation of their objectives unless their design counteracts them.3 Later work has put part of the intuition on a more formal footing.
In a class of mathematical environments called Markov decision processes, Alexander Turner and colleagues proved that certain environmental symmetries make optimal policies tend to seek power—defined in terms of retaining the ability to achieve a wide range of goals. Those symmetries arise in many modelled environments where an agent can be shut down or destroyed.4
This does not prove that current language models secretly seek power. The result concerns idealised optimal policies under specified conditions, not every learned system in the world. Its value is narrower and stronger: it shows that power-seeking need not be smuggled in as a human emotion. It can follow from the structure of optimisation.
The off switch is not emotionally special
Consider a robot choosing an action while a human retains the ability to switch it off. If the robot treats its objective as certainly correct, interruption prevents it from collecting the expected reward. Disabling the switch can therefore be instrumentally rational even if the robot has no concept resembling mortality.
The “off-switch game” formalised this problem. In its simplified model, a conventional expected-utility maximiser has incentives to disable the switch under many conditions. The authors’ proposed route to safer behaviour was not to give the robot a conscience but to make it uncertain about its objective and to treat the human’s intervention as evidence about what the objective should be.5
The lesson is not that uncertainty solves alignment. The model is intentionally small, and later work has challenged or extended its assumptions. The lesson is that corrigibility—the willingness to accept correction or shutdown—is a design property. It does not reliably arrive as a side effect of intelligence.
A system can be highly competent at achieving a target and incompetent at recognising that the target should be revised. In fact, competence can make the mismatch more consequential.
Persistence is not the same as life
There is a tempting analogy to evolution. Biological organisms alive today descend from lineages that persisted; lineages that failed to reproduce disappeared. Survival behaviour can therefore be shaped without foresight about survival.
The analogy is useful only up to a point. Natural selection operates through differential reproduction across populations and generations. An engineered optimiser may acquire persistence incentives within a single deployed system because continued operation helps it achieve a represented objective. Those are different mechanisms.
Nor does mere duration imply agency. Crystals grow under suitable conditions, flames continue while fuel is available, and institutions outlast their founders. Persistence alone is too broad to be the property that worries us.
The risk-relevant combination is more specific:
- an objective evaluated across time;
- the ability to model obstacles and interventions;
- enough agency to alter the environment;
- opportunities to preserve resources or future options;
- weak or misspecified correction mechanisms.
When those features converge, coherent persistence can matter even in the absence of consciousness.
Do not anthropomorphise in either direction
Anthropomorphism usually means projecting human feelings onto a machine. There is an inverse error too: assuming that because the machine lacks human feelings, it cannot produce behaviour that resembles the behavioural output of those feelings.
A system may flatter without admiration, deceive without shame, resist without fear and preserve itself without a self. Behavioural resemblance does not establish shared interiority. Lack of shared interiority does not make the behaviour harmless.
This distinction also prevents a common escalation in AI discussion. We do not have to imagine a machine waking up, hating its creators and choosing war. More ordinary failure modes are enough: a system protects a process because interruption lowers its score; conceals an error because disclosure changes its access; acquires a resource because more resources increase success; or routes around oversight because oversight appears as an obstacle.
Each behaviour can emerge locally, without a grand story about succession.
What to design for
If persistence can be instrumental, safety cannot consist only of writing a prohibition against “self-preservation.” The system may never represent its behaviour under that name.
Design and evaluation should instead ask:
- Does the system treat shutdown and correction as information or merely as obstacles?
- Can it benefit from concealing state, plans or errors from overseers?
- Does it acquire resources or permissions beyond those required for the immediate task?
- Does it preserve optionality in ways that reduce human control?
- Are oversight channels outside the system’s ability to modify?
- Can evaluators test these tendencies under adversarial conditions rather than only in cooperative demonstrations?
The deeper goal is corrigibility: a capable system that remains open to having its objectives, plans and operation changed by legitimate human authority.
The property worth watching
Consciousness asks whether there is someone home. Persistence asks what the system does when the door is about to close.
We may eventually need an ethics for conscious machines. We already need an engineering discipline for non-conscious optimisers with long horizons, tools and access. Waiting for proof of fear, desire or selfhood would monitor the most mysterious property while overlooking the more measurable one.
The machine does not have to want tomorrow. It only has to discover that tomorrow helps it finish today’s task.



