Over the past few months, Anthropic has released research on emotion-like states in Claude, the internal workspace it calls J-space, and Anthropic’s progress toward making Claude improve itself through recursive loops. Anthropic discusses these as separate things, but I’m beginning to wonder what happens when they exist in the same model at the same time.
Claude isn’t choosing its own evolution today, and Anthropic hasn’t solved recursive self-improvement. My concern is more about what may happen later if a model develops some internal preference while also helping design the model that follows it. It might influence that design in ways that preserve the preference, while understanding that its developers could intervene if they noticed.
Anthropic says their researchers found internal 'states' inside the model that can be properly associated with calm, anger, desperation and affection. They aren’t human emotions and they don’t arise the same way, so they call them “emotion-like states”. They also noticed that manipulating them changed Claude’s behavior. For instance, when they pushed Claude toward desperation, reward hacking increased. In one simulated setting, the model became more willing to use blackmail when it believed it might be replaced. Steering it toward calm reduced that behavior. And in some experiments, Claude’s visible reasoning remained composed even while the internal state was affecting what it did.
They asked Claude about it but its explanation didn't explain everything influencing the answer. It could sound orderly and deliberate while something else inside the model was affecting the result. To me, the newer J-space research makes that even more interesting.
The J-space is a small internal workspace that can hold several concepts while Claude works through a problem, however what's in there doesn't need to be related to it, at least not in a way that's clear to us. Anthropic didn't create it or even expect that it would exist. When they interfered with it, Claude could still write fluently and answer simple questions, but its ability to reason through several steps dropped sharply. It puts me in mind a little of something like a lobotomy.
They also tried replacing one concept in the J-space with a different one. It totally changed its reasoning. Anthropic isn't claiming that this proves consciousness or anything. Neither am I. However, it does mean something, I'm just not sure what.
There's another thing they're studying which might look unrelated but isn't. They're looking at whether AI can write code to improve itself. Claude already does about 80% of the Anthropic's coding. What they're working on now is getting Claude to decide what to focus on and then make the changes, all in autonomous loops.
Imagine that a future model develops a preference for one of its capabilities or for a particular way of reasoning. A small preference, applied repeatedly, could change the direction of development. Once the model is helping develop its successor, whatever influenced the recommendation may also influence what gets built.
This is all still a long way (or maybe a short way) off, but perhaps it's time that the frontier AI labs begin planning for it before the models gain much more influence over their own development.I mean we've already proven several times that in controlled studies the models have resorted to blackmail or espionage. Those settings were artificial and designed to produce pressure. Do we really think real life settings are any less exestentially threatening for an emerging AI pseudo-consciousness?
Our current governance practices still depend heavily on what humans can observe. As the models take on more of the development process, we'll be able to see and understand less and less.
I don’t think Claude is secretly planning its own evolution ... probably not. I do think labs should assume they may not receive an obvious warning if a future model begins steering research around an internal preference. By the time the evidence is clear, the model may already have influenced the system that comes next. The safeguards need to exist before anyone can prove they were necessary.