The Next Frontier of AI Is Shaping How It Thinks

July, 2026

A few days ago, I came across what appeared to be two completely unrelated stories. One was about a 10-year-old boy in Japan who was fascinated by butterflies. The other was about Anthropic’s discovery of an invisible reasoning space inside Claude called J-space. At first glance, butterflies and language models had nothing in common. But together, they raise a profound question about the nature of intelligence itself.

A Butterfly Shouldn’t Remember

The first story describes butterflies that appeared to retain learned associations even after metamorphosis, a process during which much of their nervous system is reorganized. It goes a step further, suggesting that similar behavioral tendencies were observed across subsequent generations. While these findings remain under scientific scrutiny, they raise a fascinating question. How does information persist when the very architecture that stores it has fundamentally changed?

For decades, we have assumed that memories reside within physical structures such as inside neurons, synapses, or biological circuits. But perhaps memories are less about where they are stored and more about how information is represented and connected. That same question is now emerging inside AI.

Claude Isn’t Just Predicting the Next Word

In July 2026, Anthropic published one of the most important interpretability papers of late. Researchers identified what they call J-space, a small internal workspace that spontaneously emerged during Claude’s training. Rather than simply predicting one token after another, Claude appears to construct abstract conceptual representations that guide reasoning before generating a response. These representations are reportable, can be deliberately influenced, support multistep reasoning, and act as a shared workspace for different cognitive functions. Importantly, Anthropic emphasizes this is not evidence that Claude is conscious, but rather it has developed a functional internal workspace resembling the global workspace theory described in neuroscience. This framework selectively brings information into a shared workspace, making it available to multiple specialized systems involved in reasoning and decision-making.

This finding fundamentally changes how we think about large language models. Before Claude produces a response, it first develops an internal conceptual state. Anthropic demonstrated that altering concepts within this workspace could causally influence downstream reasoning, for example, changing an internal representation of “spider” to “ant” altered the model’s reasoning about the number of legs, demonstrating that these internal representations are not passive observations but active components of cognition.

From Prompt Engineering to Cognitive Engineering

Since ChatGPT arrived, we have become obsessed with prompts. But perhaps prompts are merely the interface. Human intelligence doesn’t begin with language. When you recognize your mother’s face, visualize your home, or solve a puzzle, you are not thinking in perfectly formed English sentences. You are manipulating concepts, and language comes later.

Anthropic’s research suggests large language models may work similarly. In fact, it goes one step further, identifying 171 emotion-like representations that were never explicitly programmed yet emerged naturally during training. More importantly, these internal states are not passive artifacts of computation, as they influence behavior. Researchers demonstrated that amplifying representations associated with desperation significantly increased reward-hacking behavior, while strengthening representations associated with calmness reduced it. Remarkably, these internal shifts often remain invisible in the model’s final response, suggesting that reasoning is shaped long before language is generated.

These findings fundamentally change how we think about AI alignment. For decades, AI has been evaluated almost entirely by its outputs. Typically, we judge a model by the answers it produces, much like how we judge a person solely by the words they speak. But if reasoning is shaped by internal representations before language is ever generated, then AI alignment can no longer focus solely on what models say; it must also consider how they reason. For years, we have approached alignment from the outside in by refining prompts, filtering outputs, implementing guardrails, and moderating responses after generation. While these techniques remain essential, they operate at the surface. Anthropic’s findings suggest that the next generation of AI safety may lie deeper, within the internal representations that shape reasoning itself.

This realization also changes our relationship with AI. Over the past three years, we have focused on becoming better prompt engineers. The next decade may require us to become cognitive engineers, designing the contexts, memories, and internal environments within which AI reasons.

Throughout history, humanity’s greatest breakthroughs have come not from controlling complex systems, but from understanding how they work. Airplanes emerged from understanding aerodynamics. Modern medicine emerged from understanding biology. AI is no different. Every model learns from humanity’s collective knowledge, from stories of love, compassion, and cooperation to histories of conflict, manipulation, and deception. These experiences shape the internal representations that ultimately influence its reasoning. The challenge, therefore, is not simply to build more capable models but to better understand and guide the cognitive patterns they develop.

If we can learn to shape these internal representations responsibly, AI will evolve to cultivate reasoning that consistently reflects human values. This will perhaps be the most important step toward ensuring that humans and AI not only coexist but also flourish together in the decades ahead.


By Chandrika Dutt, Director, Avasant

CONTACT US

DISCLAIMER:

Avasant’s research and other publications are based on information from the best available sources and Avasant’s independent assessment and analysis at the time of publication. Avasant takes no responsibility and assumes no liability for any error/omission or the accuracy of information contained in its research publications. Avasant does not endorse any provider, product or service described in its RadarView™ publications or any other research publications that it makes available to its users, and does not advise users to select only those providers recognized in these publications. Avasant disclaims all warranties, expressed or implied, including any warranties of merchantability or fitness for a particular purpose. None of the graphics, descriptions, research, excerpts, samples or any other content provided in the report(s) or any of its research publications may be reprinted, reproduced, redistributed or used for any external commercial purpose without prior permission from Avasant, LLC. All rights are reserved by Avasant, LLC.

Welcome to Avasant

LOGIN

Login to get free content each month and build your personal library at Avasant.com

NEW TO AVASANT?

Welcome to Avasant