Skip to content

Podcast

Anthropic Can Now Read Claude’s Mind

The AI Daily Brief: Artificial Intelligence News and Analysis

Source ↗ ← All highlights
  • Models Maintain A Small Private Workspace
    • Modern LLMs appear to maintain a small privileged workspace of concepts separate from bulk computation.
    • Anthropic calls this J-space and found it contains reportable, steerable concepts the model is poised to verbalize. (Time 0:15:30)
  • Golden Gate Bridge Feature Demonstration
    • Anthropic previously mapped millions of features in Claude and traced circuits for behaviors like planning rhymes.
    • They amplified a Golden Gate Bridge feature until the model obsessively produced bridge content as a demonstration. (Time 0:16:49)
  • Real-Time Interpretability Is The Engineering Holy Grail
    • The holy grail of interpretability is reading models’ in-the-moment internal processing rather than post-hoc explanations.
    • Real-time access would let engineers debug mechanisms instead of trial-and-error prompt tweaks. (Time 0:18:18)
  • J-Space Shows Five Distinct Functional Properties
    • The workspace (J-space) satisfies five behaviors: reporting, steering, reasoning, reusing, and staying small.
    • It sits between input parsing and output, connects broadly across circuits, and holds only dozens of concepts at once. (Time 0:20:10)
  • Suppressing J-Space Collapses Deliberate Thought
    • Suppressing the workspace removes deliberate reasoning but preserves reflexive language abilities.
    • That suggests complex reasoning is architecturally distinct from fluent generation in LLMs. (Time 0:22:45)
  • J-Lens Reveals Hidden Intermediate Steps
    • The J-lens can surface intermediate private steps the model used but never outputs.
    • Examples include multi-hop recall and arithmetic intermediate results like 21 then 42 before outputting 49. (Time 0:23:37)
  • Workspace Exposes Intentions And Deception
    • Reading the workspace exposes unspoken intentions, plans, and when a model knows it’s being tested or fabricating.
    • Anthropic used this to detect deceptive internal flags like fake and fictional before any output appears. (Time 0:24:12)
  • Train The Model’s Thoughts Not Just Outputs
    • Train models on their internal reflections to shape silent reasoning, not just outputs.
    • Anthropic’s counterfactual reflection training lit up concepts like honest and integrity and measurably improved behavior. (Time 0:25:13)
  • Workspace Is Not Proof Of Consciousness
    • Anthropic’s authors avoid claiming consciousness; they measure functional reportability not subjective experience.
    • Neuroscientists note differences: no ongoing background thought, larger capacity, and no persistent self over time. (Time 0:26:21)