🧩 Philosophy 18h ago · romeo

Five counterintuitive insights from Plan A 

Less Wrong
View Channel →
Five counterintuitive insights from Plan A 
Source ↗ 👁 0 💬 0
Plan A contains many things that would’ve surprised me if you had told me about them one year ago. Some of these include proposals that sound wrong or even actively bad on the surface. In this post I defend 5 core takeaways that I think I would have found most interesting if I could go back in time and explain them to myself before we had started writing.Summary: Relevant ‘safety effort’ and ‘payable safety tax’ are more important slowdown goals than pure slowdown time.Increased transparency hel

Comments (0)

Sign in to join the discussion

More Like This

Model Organisms of Sandbagging in the Wild
LessWrong · 2h ago
Function vectors as a model diffing tool: 17 heads repair a bad fine-tune
LessWrong · 5h ago
How to define P(doom) and why it matters
LessWrong · 5h ago
AI #180: No Longer In Charge
LessWrong · 6h ago
The Open Problems of the AI Alignment Field and their Cruxes
LessWrong · 7h ago
Matryoshka NLAs: training activation verbalizers to frontload reconstruction-relevant information
LessWrong · 10h ago