🧩 Philosophy 12h ago · beyarkay (Boyd Kane)

LLMs could control their host machines by exploiting inference engines

Less Wrong
View Channel →
LLMs could control their host machines by exploiting inference engines
Source ↗ 👁 0 💬 0
Large language models often take actions running on one computer (via an agentic harness such as Claude Code or Codex), however the LLMs’ responses to prompts are computed on a different computer with GPU access. Could a malicious LLM gain control of the host machine where its weights are loaded? Such a machine is a high-value target: it has sufficient compute to run a frontier LLM, offers easy access to the LLM’s weights, and has privileged access to other computers in the datacentre compared w

Comments (0)

Sign in to join the discussion

More Like This

📰
AI Safety Acculturation is Neglected
LessWrong · 6h ago
📰
Endorsing Burhan Azeem for State Senate
LessWrong · 9h ago
📰
What just happened? Pragmatism and Pessimization
LessWrong · 19h ago
In search of natural features
LessWrong · 23h ago
📰
PSA: There's a third option in the "measure problem"
LessWrong · 1d ago
Utilities as Legendre duals of probabilities
LessWrong · 1d ago