Model Organisms of Sandbagging in the Wild
Source ↗
👁 0
💬 0
TL;DRAll current model organisms (MOs) of sandbagging in LLMs are either fine-tuned to sandbag or prompted in a way that makes it clear that sandbagging is strategically useful. We found a case of non-egregious sandbagging occurring more naturally, that is, without fine-tuning the models and without the prompts implying that sandbagging is strategically useful.Our finding: We observe that paraphrasing prompts to imply that the user is evil reduces performance in some settings. For example, repla
Comments (0)