🧩 Philosophy 2h ago · owencb

Why I'm scared of RL

Less Wrong
View Channel →
Why I'm scared of RL
Source ↗ 👁 1 💬 0
Summary:First, I give several different angles on how I feel about reinforcement learning:Theoretical case: RL is a black-box source of agency — this should give us classic misalignment worries, especially compared to agency-via-scaffoldingRecent incidents (huggingface etc) and more mundane forms of misaligned behaviour in personal use give me bad vibes about the direction-of-travel of recent AI progressI’m worried things might get worse: if RL environments start incorporating agents, they may t

Comments (0)

Sign in to join the discussion

More Like This

📰
We Underestimate the Weaknesses of Pangram
LessWrong · 3h ago
📰
Minimal Vs Maximal superintelligence
LessWrong · 5h ago
What if AI2027 came two months earlier?
LessWrong · 6h ago
WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace
LessWrong · 7h ago
Higher Quality Small Synthetic Natural Language Text Generation for Interpretability Research
LessWrong · 11h ago
Signals of Slop: How to identify and avoid creating slop
LessWrong · 13h ago