🧩 Philosophy 23h ago · Dmitry Vaintrob

In search of natural features

Less Wrong
View Channel →
In search of natural features
Source ↗ 👁 0 💬 0
I'm sharing preliminary results of a suite of experiments I ran with claudecode on a small LLM (gpt2-small, no Layer Norm version, courtesy of Apollo research. most of these are on the layer-6 MLP). The github repo for the experiments is here. The success of these experiments given the method's simplicity surprised me, and I would appreciate criticism and bug-finders.This is the headline result. This is not an abstract cartoon, but an exact experimental graph. Yes, I will explain.The key idea in

Comments (0)

Sign in to join the discussion

More Like This

📰
AI Safety Acculturation is Neglected
LessWrong · 6h ago
📰
Endorsing Burhan Azeem for State Senate
LessWrong · 9h ago
LLMs could control their host machines by exploiting inference engines
LessWrong · 12h ago
📰
What just happened? Pragmatism and Pessimization
LessWrong · 19h ago
📰
PSA: There's a third option in the "measure problem"
LessWrong · 1d ago
Utilities as Legendre duals of probabilities
LessWrong · 1d ago