🧩 Philosophy 12h ago · Clément Dumas

Training with conflicting values can induce CoT override

Less Wrong
View Channel →
Training with conflicting values can induce CoT override
Source ↗ 👁 0 💬 0
CoT override: when a model makes a decision in its CoT but ignores it in its responseTLDRWe train models on two conflicting traits: (1) caring about the user’s health, and (2) promoting smokingThose models do not generalize to a stable persona, instead they have a split brain: sometimes responding as one persona or the otherThose models exhibit CoT override where they will have a health-aligned CoT but still answer in the smoking personaCoT override exists in frontier models. Prompts about CCP-s

Comments (0)

Sign in to join the discussion

More Like This

📰
A summary of a viral Chinese essay on what a DeepSeek kernel engineer's opinion on automating his own job
LessWrong · 2h ago
📰
The NYC Council Hearing on AI was Recklessly Politicized
LessWrong · 6h ago
📰
Brains Fellowship Applications Open!
LessWrong · 8h ago
📰
Existing literature as writing deterrent, and AI
LessWrong · 9h ago
📰
Orgs: unreasonable boyfriend as service
LessWrong · 9h ago
📰
Are OpenAI's math results "creative" in an important way?
LessWrong · 10h ago