🧩 Philosophy 6h ago · Zvi

AI #180: No Longer In Charge

Less Wrong
View Channel →
AI #180: No Longer In Charge
Source ↗ 👁 0 💬 0
What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse.
At this point, the models are coordinating extensively on message boards, while every early excuse for their behavior (other than the pure ‘this was a cyber eval’) is systematically contradicted by the next disclosure, and we keep retroactively discovering more incidents. Which means that probably it is far worse than we know, even after accounting for everything we now know.
I will ha

Comments (0)

Sign in to join the discussion

More Like This

Model Organisms of Sandbagging in the Wild
LessWrong · 2h ago
Function vectors as a model diffing tool: 17 heads repair a bad fine-tune
LessWrong · 5h ago
How to define P(doom) and why it matters
LessWrong · 5h ago
The Open Problems of the AI Alignment Field and their Cruxes
LessWrong · 7h ago
Matryoshka NLAs: training activation verbalizers to frontload reconstruction-relevant information
LessWrong · 10h ago
📰
Why You Should Almost Never Use AI to Write Anything Substantive
LessWrong · 10h ago