Podcast
Anthropic Can Now Read Claude’s Mind
The AI Daily Brief: Artificial Intelligence News and Analysis
- Models Maintain A Small Private Workspace
- Modern LLMs appear to maintain a small privileged workspace of concepts separate from bulk computation.
- Anthropic calls this J-space and found it contains reportable, steerable concepts the model is poised to verbalize. (Time 0:15:30)
- Golden Gate Bridge Feature Demonstration
- Anthropic previously mapped millions of features in Claude and traced circuits for behaviors like planning rhymes.
- They amplified a Golden Gate Bridge feature until the model obsessively produced bridge content as a demonstration. (Time 0:16:49)
- Real-Time Interpretability Is The Engineering Holy Grail
- The holy grail of interpretability is reading models’ in-the-moment internal processing rather than post-hoc explanations.
- Real-time access would let engineers debug mechanisms instead of trial-and-error prompt tweaks. (Time 0:18:18)
- J-Space Shows Five Distinct Functional Properties
- The workspace (J-space) satisfies five behaviors: reporting, steering, reasoning, reusing, and staying small.
- It sits between input parsing and output, connects broadly across circuits, and holds only dozens of concepts at once. (Time 0:20:10)
- Suppressing J-Space Collapses Deliberate Thought
- Suppressing the workspace removes deliberate reasoning but preserves reflexive language abilities.
- That suggests complex reasoning is architecturally distinct from fluent generation in LLMs. (Time 0:22:45)
- J-Lens Reveals Hidden Intermediate Steps
- The J-lens can surface intermediate private steps the model used but never outputs.
- Examples include multi-hop recall and arithmetic intermediate results like 21 then 42 before outputting 49. (Time 0:23:37)
- Workspace Exposes Intentions And Deception
- Reading the workspace exposes unspoken intentions, plans, and when a model knows it’s being tested or fabricating.
- Anthropic used this to detect deceptive internal flags like fake and fictional before any output appears. (Time 0:24:12)
- Train The Model’s Thoughts Not Just Outputs
- Train models on their internal reflections to shape silent reasoning, not just outputs.
- Anthropic’s counterfactual reflection training lit up concepts like honest and integrity and measurably improved behavior. (Time 0:25:13)
- Workspace Is Not Proof Of Consciousness
- Anthropic’s authors avoid claiming consciousness; they measure functional reportability not subjective experience.
- Neuroscientists note differences: no ongoing background thought, larger capacity, and no persistent self over time. (Time 0:26:21)