🧩 Philosophy 4h ago · Vladimir Ivanov

Model Organisms of Sandbagging in the Wild

Less Wrong
View Channel →
Model Organisms of Sandbagging in the Wild
Source ↗ 👁 0 💬 0
TL;DRAll current model organisms (MOs) of sandbagging in LLMs are either fine-tuned to sandbag or prompted in a way that makes it clear that sandbagging is strategically useful. We found a case of non-egregious sandbagging occurring more naturally, that is, without fine-tuning the models and without the prompts implying that sandbagging is strategically useful.Our finding: We observe that paraphrasing prompts to imply that the user is evil reduces performance in some settings. For example, repla

Comments (0)

Sign in to join the discussion

More Like This

Function vectors as a model diffing tool: 17 heads repair a bad fine-tune
LessWrong · 7h ago
How to define P(doom) and why it matters
LessWrong · 7h ago
AI #180: No Longer In Charge
LessWrong · 8h ago
The Open Problems of the AI Alignment Field and their Cruxes
LessWrong · 9h ago
Matryoshka NLAs: training activation verbalizers to frontload reconstruction-relevant information
LessWrong · 11h ago
📰
Why You Should Almost Never Use AI to Write Anything Substantive
LessWrong · 11h ago