Garp Independent AI & technology journalism
Saturday, August 8, 2026 Sign In · Join Subscribe
Latest Naïve raises $28.5M to automate the grunt work of setting up and running a company

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  AI News  /  Claude’s hidden inner monologue is now readable thanks to Anthropic’s new Jacobian Lens

AI News

Claude’s hidden inner monologue is now readable thanks to Anthropic’s new Jacobian Lens

Claude’s hidden inner monologue is now readable thanks to Anthropic’s new Jacobian…

Anthropic has found that Claude developed an internal working memory on its own during training. The company calls it “J-Space” and can now read it using a new analysis tool called J-Lens.

It shows that Claude has developed a small set of internal neural patterns that play a distinct role compared to the rest of its processing. The researchers call it “J-Space” and classify it under Global Workspace Theory from consciousness research. That theory holds that conscious thought relies on a kind of central working memory. The work builds on the company’s earlier interpretability research. Using an “AI microscope,” Anthropic had already shown that Claude activates language-independent concepts and works through multi-step questions in individual reasoning steps. Every pattern in J-Space is linked to a word or concept without the model having to output it, similar to internal thinking in words. According to Anthropic, Claude can report on the stored content, modify it on request, and use it for multi-step inferences. The company had already explored reading out and steering internal states in a previous study on self-awareness in language models. When the concept “spider” is stored in J-Space, Claude derives the number of legs from it. Swap that representation for “ant,” and the model answers “6” instead of “8.” The same holds for country names. If “France” is active, Claude can flexibly derive the capital, language, continent, or currency. Replace “France” with “China,” and the answers shift to Beijing, Chinese, Asia, and yuan. Anthropic had already shown that individual concept representations can be isolated and swapped this way with its “Persona Vectors.” When J-Space is suppressed, Claude still speaks fluently, classifies sentences, and answers simple factual questions. But it loses multi-step inferences, summaries, and the ability to compose rhymes. In one test with a Spanish text passage, the model kept writing fluent Spanish after the manipulation but incorrectly called the language French and attributed it to Victor Hugo instead of Garcia Marquez. In a blackmail scenario from earlier studies on agentic misalignment, J-Lens shows that Claude Sonnet 4.5 recognizes the setup as fabricated before producing any output.