Why I'm scared of RL
Source ↗
👁 1
💬 0
Summary:First, I give several different angles on how I feel about reinforcement learning:Theoretical case: RL is a black-box source of agency — this should give us classic misalignment worries, especially compared to agency-via-scaffoldingRecent incidents (huggingface etc) and more mundane forms of misaligned behaviour in personal use give me bad vibes about the direction-of-travel of recent AI progressI’m worried things might get worse: if RL environments start incorporating agents, they may t
Comments (0)