🧩 Philosophy 7h ago · camilablank

WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace

Less Wrong
View Channel →
WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace
Source ↗ 👁 0 💬 0
TL;DRWe introduce WorkspaceBench, a set of evaluations for how well an activation-to-text tool can read the contents of the “global workspace” of a model, i.e. the intermediate variables during a forward pass.The benchmark comprises 3,356 questions across 27 eval families, spanning topics in safety, logical reasoning, and multihop computation, with a subset for single-token-output tools.A desirable property of good interpretability techniques is minimal hallucinations, so WorkspaceBench also pro

Comments (0)

Sign in to join the discussion

More Like This

Why I'm scared of RL
LessWrong · 2h ago
📰
We Underestimate the Weaknesses of Pangram
LessWrong · 3h ago
📰
Minimal Vs Maximal superintelligence
LessWrong · 5h ago
What if AI2027 came two months earlier?
LessWrong · 6h ago
Higher Quality Small Synthetic Natural Language Text Generation for Interpretability Research
LessWrong · 11h ago
Signals of Slop: How to identify and avoid creating slop
LessWrong · 13h ago