Matryoshka NLAs: training activation verbalizers to frontload reconstruction-relevant information
Source ↗
👁 0
💬 0
TLDR: We train a “matryoshka” NLA that, unlike standard NLAs, is trained to put the most important details at the start; it is trained by randomly truncating the verbalizer’s explanations before showing them to the reconstructor. We find that our matryoshka NLA frontloads claims that are important to reconstruction (more than standard NLAs), and that this can be used as a heuristic for saliency of the represented feature. We did not find strong evidence that matryoshka NLAs are significantly mor
Comments (0)