🧩 Philosophy 1d ago · loops

Matryoshka NLAs: training activation verbalizers to frontload reconstruction-relevant information

Less Wrong
View Channel →
Matryoshka NLAs: training activation verbalizers to frontload reconstruction-relevant information
Source ↗ 👁 0 💬 0
TLDR: We train a “matryoshka” NLA that, unlike standard NLAs, is trained to put the most important details at the start; it is trained by randomly truncating the verbalizer’s explanations before showing them to the reconstructor. We find that our matryoshka NLA frontloads claims that are important to reconstruction (more than standard NLAs), and that this can be used as a heuristic for saliency of the represented feature. We did not find strong evidence that matryoshka NLAs are significantly mor

Comments (0)

Sign in to join the discussion

More Like This

Model Organisms of Sandbagging in the Wild
LessWrong · 1d ago
Function vectors as a model diffing tool: 17 heads repair a bad fine-tune
LessWrong · 1d ago
How to define P(doom) and why it matters
LessWrong · 1d ago
AI #180: No Longer In Charge
LessWrong · 1d ago
The Open Problems of the AI Alignment Field and their Cruxes
LessWrong · 1d ago
📰
Why You Should Almost Never Use AI to Write Anything Substantive
LessWrong · 1d ago