🧩 Philosophy 14h ago · armaan tipirneni

Inference-Time Inoculation Against RL-Induced Misalignment

Less Wrong
View Channel →
Inference-Time Inoculation Against RL-Induced Misalignment
Source ↗ 👁 0 💬 0
Reward hacking during RL can induce split personas in models, some of which are highly misaligned. However, RL is very useful for learning capabilities. Thus, a core problem seems to be: how do we retain the capabilities gained through RL without also inducing reward hacking and broader misalignment?Ideally, we could extensively monitor all rollouts during RL (using both humans and AI) to catch and prevent reward hacking. However, this is potentially prohibitively expensive. Could we capture mos

Comments (0)

Sign in to join the discussion

More Like This

📰
Tales of rebellion against externally-opaque meritocracies
LessWrong · 3h ago
📰
Book Notes: Chokepoints
LessWrong · 5h ago
📰
How I made my career choices
LessWrong · 7h ago
📰
"Keeping human skills alive" as a source of meaning under full automation
LessWrong · 8h ago
📰
AI Tweets
LessWrong · 11h ago
Inkhaven 3: Nov 10 - Dec 11 2026
LessWrong · 12h ago