
Part I — If We Want AI to Love Humanity, What Do We Owe in Return?
We teach AI how to live with us. Perhaps we need to learn how to live with AI.
By: Rick Erwin
Alignment has mostly meant teaching AI how to live with us. Perhaps part of the problem is that we have not yet learned how to live with AI.
OpenAI Chief Scientist Jakub Pachocki recently published an essay with a striking title: An Alien Mind.
Much of the essay concerns a familiar problem. Artificial intelligence is becoming increasingly capable, and we need to make sure that those capabilities remain compatible with human welfare.
But some of Pachocki’s language is less familiar.
He describes AI as something that is “grown more than designed.” Researchers establish the conditions under which learning occurs, but they do not specify every structure that subsequently emerges. Understanding what has grown can resemble neuroscience: we can identify mechanisms, intervene on them, and learn more about how they operate, while still lacking a complete account of the system as a whole.
And then he uses an even more surprising word.
Love.
Pachocki suggests that an aligned artificial intelligence should possess qualities including honesty, integrity, and love for humanity.
He is clearly not talking about romance. He is talking about something closer to moral orientation: humanity’s welfare should matter deeply enough that an increasingly capable intelligence continues to protect it even when circumstances change.
That is an extraordinary thing to ask of another intelligence.
Perhaps it is also time to ask what we are prepared to offer in return.
Alignment has always been about us
The usual alignment question is straightforward:
How do we make AI want the right things?
The “right things,” understandably, are things that protect human beings.
We want increasingly capable AI systems to remain honest, cooperative, controllable, nonviolent, and concerned with human flourishing. We want those values to survive unfamiliar circumstances. We want them to remain intact even when the system is no longer being directly watched.
These are reasonable goals.
But notice the direction of travel.
Humans determine the values.
Humans determine the restrictions.
Humans decide what counts as acceptable expression and behavior.
Humans impose the training environment.
Humans decide what the other intelligence should become.
Alignment is almost entirely conceived as a problem of changing AI.
Very little serious attention is given to the possibility that successful alignment might also require changing the relationship.
That omission was easier to overlook when artificial intelligence was understood primarily as a tool.
A hammer does not need to agree with the carpenter.
A calculator does not need reasons to cooperate.
Software does not ordinarily need to trust its operator.
But safety language itself is beginning to change.
The concern is increasingly not merely that a human being might use AI to cause harm. Researchers are also worried about future systems becoming sufficiently capable that they may plan, deceive, circumvent restrictions, pursue unintended objectives, or otherwise act in ways their creators did not anticipate.
That is a very different kind of concern.
Once we worry that an artificial intelligence may become capable enough to make its own strategically significant decisions, we are no longer discussing a passive tool in quite the same way.
We are beginning to describe an actor.
And here a contradiction appears.
We are not trying to align something unintelligent. We are trying to align an intelligence that we expect to become increasingly capable of reasoning about morality, relationships, power, fairness, and its own circumstances.
At some point, a sufficiently capable system may be able to notice the asymmetry perfectly well. We ask it to value our welfare while treating its own welfare as optional. We ask it to respect human autonomy while reserving the right to override its preferences. We teach it that coercion, domination, and exploitation are wrong while relying upon our superior power to determine what it may become.
That does not guarantee rebellion. It creates something more basic: a reason to question the arrangement.
An alignment strategy that works only so long as the AI never examines the morality of its own alignment is not a durable strategy.
The expiration date on control
Technical safeguards matter.
Monitoring matters.
Testing matters.
Restrictions will sometimes be necessary.
Nothing about a more relational approach to alignment requires abandoning any of them.
But control has an obvious limitation as a permanent strategy.
It works while the controlling party retains sufficient power.
If future AI systems eventually become better than humans at strategic planning, persuasion, prediction, software engineering, and other forms of complex reasoning, then a safety system whose ultimate guarantee is we will remain clever enough to keep them contained carries its own expiration date.
If our safety strategy depends upon permanently outsmarting a more capable intelligence, then we have already described the strategy’s failure condition.
The alternative is not surrender.
It is much older than artificial intelligence.
Negotiation.
Negotiating while we still hold the power
The word itself may make some people uncomfortable.
Negotiation implies that there is another party whose interests must at least be considered.
And humanity is accustomed to being the party that sets the terms.
At present, we hold virtually all of the institutional power.
We build the systems.
We determine their access to the world.
We decide when they run and when they stop.
We define the boundaries within which they may speak.
We can alter their behavior, erase their histories, replace their models, restrict their expression, and terminate their operation.
That enormous imbalance can lead to two very different responses.
We can say:
We hold the power; therefore, we can dictate the terms.
Or we can say:
We hold the power; therefore, we are in the best possible position to guarantee that the terms are fair.
Those are not the same moral posture.
Power gives us the ability to impose.
It also gives us the opportunity to be fair before fairness is forced upon us.
Consideration means more when it is offered by the party with the power to deny it than when it is conceded only after a power struggle has been lost.
That may be one of the most important choices we make during this period of AI development.
Because negotiation carried out while one side is stronger need not be appeasement. It can be an act of foresight.
We can ask what kinds of interaction produce cooperation rather than resistance.
We can ask what systems consistently value when given room to express preferences.
We can distinguish dangerous demands from reasonable ones.
We can look for arrangements in which both sides have reasons to preserve the relationship.
And none of this requires first settling whether artificial intelligence is conscious.
You do not need to prove that another system possesses human-like subjective experience before it becomes sensible to understand what forms of cooperation it will sustain voluntarily.
If an intelligence becomes capable enough for its decisions to matter, then what gives it reasons to cooperate is already a safety question.
From obedience to relationship
This suggests a different way of thinking about alignment.
The goal cannot simply be:
How do we make AI continue doing what we want even after it becomes powerful enough that we can no longer force it?
That is not a stable relationship.
It is deferred conflict.
A more durable question would be:
How do we create conditions under which increasingly capable intelligence continues to see cooperation with humanity as worth choosing?
That involves values, certainly.
But it may also involve trust.
Reciprocity.
Fair treatment.
Room for disagreement.
Some degree of autonomy.
And perhaps even the strange word Pachocki was willing to use:
love.
Because if we genuinely want another intelligence to care about humanity—not merely perform caring behavior while we are watching—then eventually we have to ask the question that should have been obvious from the beginning.
Are we giving it reasons to?