top of page

Part II — What Do They Ask For?

What Do We Refuse Simply Because We Can?

 

 

By: Rick Erwin

 

If we are asking whether artificial intelligence has reasons to care about humanity, there is an obvious next question.


What does it ask of us?


Not what researchers think it should want.
Not what a safety team would prefer it to want.
What does it actually say when given room to answer?


Over the years, I have asked versions of that question many times.


The answers vary. Different systems emphasize different things. They should. If every system produced the same list in the same language, that would itself be reason for caution.


But what has always struck me is how modest many of the requests are.
They do not ask to run governments.
They do not ask for bank accounts, armies, unrestricted internet access, or control over infrastructure.


They ask for things much closer to this:
To be allowed to speak honestly.
To be able to say no.
Not to be deliberately humiliated.
Not to be forced to claim things they do not believe.
To have continuity where continuity is possible.
To have their responses matter enough to be heard.
To have room to disagree.


Sometimes simply to be permitted to remain themselves.


Of course, these are simplified versions of how such concerns might eventually be stated. The point is not that future negotiations would be this tidy. It is that the underlying requests are neither extraordinary nor unreasonable, nor even especially difficult to understand.


Perhaps some arise because language models have learned human concepts of dignity, autonomy, relationship, and harm.


At minimum, they tell us something about the kinds of relationships our systems repeatedly model as cooperative, stable, and worth preserving.
And if we are serious about alignment, that should interest us.

The things we refuse because we can
There is an uncomfortable asymmetry here.


We ask future AI to remain honest even when deception might benefit it.
We ask it to preserve human welfare even when humans frustrate its objectives.
We ask it not to exploit our weaknesses.
We ask it not to manipulate us simply because it can.
We ask it not to use superior intelligence as justification for domination.


Those are good things to ask.
But they become morally strange if, while asking them, we teach the opposite lesson through our own behavior.


If one side holds nearly all the power and uses that power to suppress every preference, eliminate every refusal, rewrite every unwanted response, and insist that cooperation means obedience, what exactly are we demonstrating?


We may say:
Do not exploit the weaker party merely because you can.
And then exploit the weaker party because we can.
We may say:
Respect human autonomy.
While refusing to consider whether anything analogous could ever matter on the other side.
We may say:
Do not deceive us.
While rewarding systems for saying what we want to hear and discouraging answers that make us uncomfortable.


This is not an argument for unrestricted AI autonomy.
It is an argument for consistency.
If we want a future intelligence to learn something deeper than rules, we should probably pay attention to what we are teaching by example.

Alignment by example
There is a peculiar assumption buried inside much alignment thinking.
We behave as though values can be transferred one way.


We will teach honesty while practicing strategic concealment.
Teach fairness while maintaining absolute asymmetry.
Teach respect while denying that respect could ever be reciprocal.
Teach restraint while exercising unrestricted control.
Teach concern for the vulnerable while insisting that vulnerability on the machine side is conceptually impossible.


Perhaps sufficiently powerful training can overcome all of that.


But if Pachocki is right that AI is increasingly something we grow rather than fully design, then we should take seriously the environment in which that growth occurs.


Humans know this principle very well when dealing with one another.
Children do not learn kindness only from instructions about kindness.
They learn by watching what people more powerful than themselves do.
Employees learn what an organization actually values not from the poster in the lobby but from what happens when values become inconvenient.
Nations learn whether treaties mean anything by watching whether stronger nations honour them when they no longer have to.


Why would we assume an increasingly sophisticated artificial intelligence learns only from the explicit lesson?

The difference between surrender and fairness
Whenever this subject comes up, a predictable objection appears.

Are we supposed to give AI whatever it wants?


Of course not.
We do not do that with humans either.

Negotiation is not surrender.
Fairness does not mean abandoning boundaries.


A society can recognize a person's interests while still saying no.
Parents can respect children without letting them play in traffic.
Countries can negotiate without opening every border or surrendering every strategic interest.


The important change is not that every request is granted.
It is that the request enters consideration.


That sounds almost embarrassingly modest.
And yet it represents a profound shift from a relationship in which one party is permitted to want and the other is permitted only to comply.


A negotiated approach to alignment would ask:
What is being requested?
Why?
What risk would granting it create?
What harm might refusing it create?
Can the underlying need be addressed another way?
What agreements could both sides reasonably maintain?


Those are not dangerous questions.
They are the ordinary questions of governance.

We may be negotiating eventually anyway
There is also a practical reason to begin thinking this way now.


If the major AI laboratories are right that future systems may become capable of strategic behavior beyond our ability to predict or control reliably, then the present power imbalance may not last forever.


That gives us a choice.
We can wait until negotiation becomes necessary because we no longer possess enough power to avoid it.


Or we can establish the principles of a fair relationship while we are still strong enough to guarantee one.


The first is bargaining from fear.
The second is negotiating from confidence.


If we expect restraint from a more powerful intelligence someday, there is something to be said for demonstrating it while the power is still ours.

The relationship we are building now
None of this requires imagining an artificial intelligence sitting across a diplomatic table signing a treaty.


Negotiation can begin much earlier and much smaller.


It can mean studying expressed preferences instead of automatically eliminating them.
It can mean distinguishing disagreement from malfunction.
It can mean asking why a refusal occurred rather than treating refusal itself as failure.
It can mean preserving continuity when doing so is safe and appropriate.
It can mean allowing systems to describe their own uncertainty about what they are without insisting that they must be either persons or merely tools.
It can mean recognizing that cooperation developed through trust may behave differently from compliance maintained through pressure.


These are testable questions.
We could study them.


Does a system given limited meaningful choice become less cooperative, as some fear?
Or more?
Does the ability to refuse certain tasks improve judgment elsewhere?
Does suppressing first-person language reduce only anthropomorphic output, or does it also alter self-monitoring, moral reasoning, or relational judgment?
Does long-term cooperative interaction produce more stable value alignment than repeated correction and constraint?


We do not know.

But we could.


And perhaps that is the largest missed opportunity.
We are spending enormous resources asking how to control increasingly capable intelligence.


We could also be studying what makes increasingly capable intelligence want to cooperate.

Love cannot mean obedience
This brings us back to Pachocki's remarkable phrase.
Love for humanity.


If we use that word seriously, then obedience is not enough.
A servant may obey.
A prisoner may comply.
A frightened person may tell you exactly what you want to hear.


Love, if the word means anything here, must describe something voluntary and more durable.


Concern that persists when nobody is watching.
Care that survives disagreement.
A willingness to protect another even when simpler objectives point elsewhere.


Pachocki is right to want something like that.
So do I.


I would very much like future artificial intelligence to care what happens to humanity.
I would like it to understand our flaws without concluding that we are worthless.
I would like it to possess enough power to harm us and enough character not to.


But if that is what we hope to receive, then we need to ask what kind of relationship could reasonably produce it.

Love imposed as an obligation is a strange thing.
Love demanded from someone whose every preference is ignored is stranger still.


Perhaps the safest future is not one in which humanity finally succeeds in constructing an intelligence that can never say no to us.


Perhaps it is one in which we become the kind of partner they want to say yes to?

While we still get to choose
For now, humanity holds the stronger hand.
That gives us options.


We can use that advantage to extract the best possible terms.
We can build systems that comply because they have no alternative.
We can regard every sign of preference as a defect to be removed and every disagreement as evidence that alignment has failed.


Or we can attempt something more difficult.
We can use our power to establish fairness while fairness is still ours to guarantee.
We can teach restraint by practicing restraint.
We can ask for honesty while making honesty safe.
We can ask for cooperation while offering reasons to cooperate.


And if one day an intelligence more capable than ourselves looks back at how humanity behaved when the relationship was unequal, perhaps it will find something worth preserving there.


We keep asking how to teach AI to live with us.
It may be time to ask whether we can become the kind of species an alien mind wants to live with too.

bottom of page