Anthropic Finds Claude's Hidden 'Thinking Room'—But It's Not Consciousness
Anthropic researchers have discovered an emergent internal workspace in Claude, called J-space, that holds silent reasoning and concepts the model is 'holding in mind' but not saying. The finding offers a new window into AI interpretability and safety, but the company explicitly states it does not prove consciousness.

Anthropic has cracked open the black box of its own AI model and found something unexpected: a hidden internal workspace where Claude silently holds concepts and plans its reasoning before speaking. The company calls it J‑space, and it is the first time researchers have been able to systematically observe a large language model's "internal thoughts" that never appear in its output. But before you start worrying about a conscious machine, Anthropic is emphatic: this is not proof of consciousness—it is a new tool for safety and interpretability.
What happened
Anthropic researchers report that Claude, the company's flagship large language model, has developed an emergent internal workspace they call J‑space. This is not something engineers explicitly built into the architecture; it emerged during training as a set of internal neural patterns that correspond to concepts the model is actively considering, even if those words never appear in its answer. The team developed a technique called Jacobian lens (or J‑lens) to inspect this space, based on Jacobian analysis—studying how small perturbations in internal activations affect outputs.
Using J‑lens, researchers can identify which internal patterns correspond to particular words or concepts. In a widely cited example, when Claude is asked "How many legs does a spider have?", a "spider" pattern activates in J‑space even if the model's answer never includes the word "spider." If researchers intervene and replace that pattern with the one corresponding to "ant," Claude's answer changes to "6" legs instead of 8. This demonstrates that J‑space contents causally influence the final answer, not just correlate with it.
💡 J‑space is not a metaphor—it is a measurable, causally active internal workspace that researchers can read and even manipulate to steer the model's outputs.
Why it matters
The discovery arrives at a moment when the AI industry is grappling with two competing pressures: the race to build ever more capable models, and the urgent need to understand what those models are actually doing. Anthropic's finding is significant because it offers a new layer of interpretability—a way to peek into the model's "silent reasoning" before it produces text. This is distinct from chain-of-thought, which is a user-visible text output. J‑space is a hidden, non-text workspace that holds concepts the model is actively considering but not saying.
Anthropic frames this as a breakthrough for AI safety. By monitoring J‑space, researchers could detect hidden planning, deception, or misaligned objectives before they appear in output. For example, they might spot test awareness—the model realizing it is being tested—or fabricated data as they form, not just after the fact. The ability to intervene and steer outputs by modifying J‑space activations (as in the spider-to-ant experiment) gives researchers a practical handle on internal reasoning that was previously inaccessible.
💡 The discovery of J‑space shifts the interpretability debate from post-hoc analysis of outputs to real-time monitoring of internal reasoning—a potential game-changer for AI safety.
What it means for business
For companies deploying Claude in production—whether for customer support, code generation, or content creation—the practical implications are immediate. J‑space offers a new layer of oversight: the ability to detect hidden planning, deception, or misaligned objectives before they appear in output. This could be used to catch test awareness (the model realizing it is being tested) or fabricated data as they form, not just after the fact.
Anthropic also demonstrated that researchers can modify J‑space activations to steer the model's outputs. In the spider-to-ant experiment, swapping the internal pattern changed the answer from 8 to 6 legs. This gives developers a practical handle on internal reasoning—a way to debug why a model gave a particular answer and to correct it at the neural level, not just by tweaking prompts.
💡 For businesses deploying Claude, J‑space offers a potential early-warning system for hidden misalignment or deception—but it also raises new questions about how much access to internal reasoning is appropriate.
What to watch next
The discovery of J‑space is a landmark in AI interpretability, but it is only the beginning. Anthropic has already shown that Claude can report on the state of J‑space when prompted and modulate it when instructed to think about something. This opens the door to more robust guardrails—detecting harmful concepts before they surface—and better alignment research tools. The company is also exploring how J‑space relates to other features like extended thinking mode and the "think" tool for agentic use, though these are technically separate lines of work.
For now, the most important takeaway is that we have a new, practical window into how large language models reason. The debate about whether AI is becoming "more human" will continue, but Anthropic has given researchers a concrete tool to watch the silent space where models think before they speak. The next step is to see whether that tool can catch hidden deception or misalignment before it reaches the output—and whether the industry will adopt it as a standard safety practice.
Want automation like this for your business?
Get in touch and we'll show you exactly what's possible for your setup.