🧩 Philosophy 5h ago · Richard Juggins

Learning new facts can change LLM behaviour

Less Wrong
View Channel →
Learning new facts can change LLM behaviour
Source ↗ 👁 0 💬 0
TL:DR: I use synthetic document fine-tuning to train an LLM to believe that in 2027 ‘long-horizon’ frontier LLMs count as moral persons. I find the model scores highly on measures of belief depth, and that prompting alone is also effective. Furthermore, I find this new belief can have substantial consequences on downstream behaviour, although this is highly context-dependent. When audited in a scenario specifically about model welfare, the fine-tuned model argued with the auditor about its belie

Comments (0)

Sign in to join the discussion

More Like This

📰
Mom's Advice For Hosting A Class Reunion
LessWrong · 3h ago
📰
I'm starting a interview series of people working in Lean / formal methods / math formalization
LessWrong · 5h ago
📰
Nuclear physics of Alex Zhao's comment for "Pacing the Frontier"
LessWrong · 5h ago
📰
All Utilitarians Should Be Classical Utilitarians
LessWrong · 6h ago
On Dwarkesh Patel’s Podcast With Ryan Greenblatt
LessWrong · 6h ago
📰
Rerunning AI safety papers on every frontier release would be pretty easy and valuable
LessWrong · 14h ago