🧩 Philosophy 5h ago · Aniket Ghosh

Function vectors as a model diffing tool: 17 heads repair a bad fine-tune

Less Wrong
View Channel →
Function vectors as a model diffing tool: 
17 heads repair a bad fine-tune
Source ↗ 👁 0 💬 0
I take two fine-tuned models trained to give bad medical advice, one on Qwen2.5-7B and one on Llama-3.1-8B, from the Model Organisms for Emergent Misalignment collection, and I found out that I could make one safe by simply copying 17 attention heads from the base model it was trained from (which I'm assuming is the good one). The interesting thing is that the reverse is not true. There is also a single direction you can pull out of the difference between the two models, and removing it partly c

Comments (0)

Sign in to join the discussion

More Like This

Model Organisms of Sandbagging in the Wild
LessWrong · 2h ago
How to define P(doom) and why it matters
LessWrong · 5h ago
AI #180: No Longer In Charge
LessWrong · 6h ago
The Open Problems of the AI Alignment Field and their Cruxes
LessWrong · 7h ago
Matryoshka NLAs: training activation verbalizers to frontload reconstruction-relevant information
LessWrong · 10h ago
📰
Why You Should Almost Never Use AI to Write Anything Substantive
LessWrong · 10h ago