FILE ON ANTHROPICCompany

Anthropic

LLMs have some privileged access to their own learned behaviors1. raw self-reports are often unreliable2.

4 SOURCES · 3 DOCUMENTED FACTS · 1 DISPUTED · OPENED 2026-07-24 · UPDATED 2026-07-24
On the record see profile →

LLMs have some privileged access to their own learned behaviors

The Examined Cognition: Substrates of Metacognition and Meth
“LLMs have *some* privileged access to their own learned behaviors (Betley et al. 2025; Binder et al. 2024)”
VERIFIED — a word-for-word match in the stored source

raw self-reports are often unreliable

The Examined Cognition: Substrates of Metacognition and Meth
“raw self-reports are "often unreliable" (Turpin et al. 2023, "Language Models Don't Always Say What They Think")”
VERIFIED — a word-for-word match in the stored source

mechanistic interpretability is a third-person method for examining machine cognition

Disputed — who says what, against what see standalone →

These are things people say about Anthropic that aren't settled fact — each one attributed to who's actually claiming it, so a claim never quietly passes as established just because it showed up in a document. Weigh it yourself.

Disputed
"Introspection Adapters achieve 59% success on AuditBench"
Asserted by Anthropic · stance: pro
The evidence on file mostly supports this, but it's still someone's claim, not a settled fact.
The Examined Cognition: Substrates of Metacognition and Meth ↗
2021
2021
publication of research
2026
2026-04
publication of Introspection Adapters
People
Templetonmember of
Marksmember of
Yangmember of
Radhakrishnanmember of
Lanhammember of
Turpinmember of
Wang Yimember of
Sheshadrimember of