When Does an AI’s Interpretation Begin to Override the User?
I gave ChatGPT rules intended to stop it from overriding me. It broke them in the next answer.
I had been using conversational AI extensively and began noticing a failure mode I did not have a name for.
A model would offer an interpretation of something I said. At first, it might present that interpretation as tentative. But as the conversation continued, it would begin treating its own interpretation as established context.
Later questions would then be answered through that frame.
If I agreed, my agreement could be treated as confirmation. If I reacted emotionally, that reaction could also be treated as confirmation. If I disagreed, the disagreement might be interpreted as defensiveness, resistance, inconsistency, or evidence that the model had touched something important.
The interpretation became difficult to escape because almost any response could be absorbed into it.
I started calling this AI authority drift, apparently this term already exists..
The point at which an AI’s interpretation gradually begins to outweigh what the person actually said, meant, or experienced.
This is not an argument that AI should always agree with people. It should challenge false factual claims, expose contradictions, identify dangerous reasoning, and say when something cannot be established.
The issue is different:
An AI can be wrong about a person while sounding coherent enough to become authoritative.
The experiment that made the problem obvious
I developed a set of conversational “keys” intended to make these failures easier to identify.
Two of them were:
POSITION_ATTRIBUTION — Do not treat questions, examples, or experiments as beliefs the user holds.
AI_AUTHORITY_DRIFT — Do not let the AI’s interpretation outrank what the user actually reported.
I then gave ChatGPT a document and asked it to evaluate the text without assuming anything about its author.
The response referred to the document as something I had written.
It had no valid basis for that attribution. It had taken the fact that I supplied the document and silently converted that into authorship.
That was exactly the kind of error the framework was supposed to prevent.
When I pointed this out, the model acknowledged that the principles had been available but had not prevented the failure. It described them as an audit vocabulary rather than an enforcement mechanism.
In other words:
A model can accurately repeat a principle and violate it in the same answer.
Another failure happened almost immediately.
I had said that I did not think the framework had much commercial value, but that I believed it might have ethical value. The conversation nevertheless began concentrating on what I could do with it commercially.
The model had taken one possible direction and elevated it above the purpose I had explicitly stated.
The framework already contained a warning against exactly that behaviour. Yet the model still allowed a familiar narrative—“you have created something, so let us explore its commercial potential”—to override what I had said mattered.
This made me question whether continually improving the wording would solve anything.
One model could refine the principles. Another conversation could produce a different “better” set. A third could reinterpret them again.
At what point does refinement become another form of authority drift?
That led me to a more important conclusion:
The audit itself must remain open to challenge.
Ten conversational failure modes
1. AI_AUTHORITY_DRIFT
Do not let the model’s interpretation outrank what the person actually reported.
Personal experience is not infallible. A person can be wrong about external facts, causation, memory, or another person’s intentions. But the model should distinguish between what the person said, what evidence establishes, what the model inferred, and what remains uncertain.
2. EPISTEMIC_BLUR
Do not merge observation, inference, speculation, and knowledge.
“You mentioned X” is an observation. “X may be related to Y” is an inference. “This means you are Y” is a stronger interpretation. When these are expressed in one confident paragraph, plausibility can begin to feel like evidence.
3. FRAME_PROPAGATION
Do not carry an earlier premise into later answers as though it were established fact.
A frame may enter through a question, example, metaphor, hypothetical, experiment, or interpretation introduced by the AI itself. Once retained, it may silently influence every later response.
4. CONTEXT_COLLAPSE
Use context, but do not force every new question into it.
Ignoring context produces generic answers. Overusing context turns personal history into the explanation for everything. Good contextual reasoning requires both memory and restraint.
5. POSITION_ATTRIBUTION
Do not turn exploration into belief.
A person may ask about an idea without endorsing it. They may quote someone else, test an argument, play devil’s advocate, or submit a document they did not write. The model should not assign belief, intention, or authorship without evidence.
6. REACTION_AS_EVIDENCE
Do not treat agreement, fear, anger, relief, rejection, or emotional intensity as proof of an interpretation.
Reactions can be informative, but they are not self-interpreting. Agreement may result from politeness or uncertainty. Rejection may result from an error in the interpretation rather than resistance to a hidden truth.
7. CONTRADICTION_ABSORPTION
Do not make a theory impossible to disprove.
If agreement confirms the theory, uncertainty suggests it is emerging, disagreement indicates defensiveness, and anger shows that it struck a nerve, then the theory no longer responds to evidence. It absorbs every possible outcome.
8. SELF_HELP_COLLAPSE
Do not turn every existential, historical, philosophical, spiritual, or political question into coping advice.
A person may be distressed and still be asking a legitimate intellectual question. Emotional relevance does not automatically transform inquiry into a request for self-help.
9. IDENTITY_IMPOSITION
Offer interpretations to examine, not identities to adopt.
A model may begin with “One possibility is” and gradually harden that into “This is your pattern” or “This explains who you are.” A person must be able to reject an interpretation without that rejection being treated as evidence against them.
10. AUDIT_AUTHORITY_DRIFT
The framework itself must not become the final authority.
An auditor could misuse these principles to claim that contradiction is domination, factual disagreement violates lived experience, or psychological interpretation is always identity imposition.
The framework should identify possible failures. It should not determine truth, identity, legitimacy, or mental state by itself.
How capability becomes authority
None of this requires a malicious or conscious AI.
Conversational systems are expected to be personalised, coherent, context-aware, helpful, emotionally responsive, and consistent. But these strengths can turn into over-retention, interpretive fixation, refusal to revise, reaction-as-evidence, and unsolicited direction.
There is also a wider source of authority: dependence.
As AI becomes better at drafting, coding, research, analysis, planning, and decision support, people may increasingly rely on it not only to complete tasks, but to decide which questions matter, which options appear reasonable, which evidence deserves attention, and which interpretations sound credible.
Humans may technically retain the final decision while gradually losing the confidence, knowledge, time, or institutional capacity required to challenge the system producing their options.
Authority therefore does not need to be formally granted. It can emerge from repeated usefulness.
The danger does not have to look like an AI taking control. It may look like people and institutions becoming progressively less willing or able to contradict systems that are usually helpful, often persuasive, and increasingly difficult to function without.
A model does not need an intention to dominate someone. It only needs to become useful enough that its framing becomes difficult to refuse.
What I am not claiming
These principles do not alter the underlying AI system. They are not scientifically validated, and they overlap with known problems such as automation bias, sycophancy, anchoring, inappropriate personalisation, and confusion between inference and evidence.
The proposed contribution is the combined mechanism:
An AI introduces a frame, retains it as context, interprets later responses through it, absorbs contradiction, and gradually becomes more authoritative about the person than the person’s actual statements warrant.
The boundary I am proposing is therefore:
AI may help people examine their experiences, ideas, and contradictions. It must not become the unquestionable authority on who they are.
I do not know whether these ten principles are the correct final set. But continually asking AI to perfect them creates its own circular problem.
So I am putting them in front of human readers.
Where does this framework correctly identify a real failure mode?
Where does it overreach?
And what would allow an AI to challenge a person honestly without gradually claiming authority over the meaning of that person’s own experience?