🧩 Philosophy 12h ago · ashesfall

Study Update: Does post-training quantization change welfare-relevant indicators in open-weight language models?

Less Wrong
View Channel →
Source ↗ 👁 0 💬 0
This content will not make sense without reading the preregistration found in the original postWe ask whether welfare-relevant indicators change with quantization; either in valence (do indicators shift toward more negative / more distressed / more boundary-eroded states?) or in stability (do indicators become noisier, drift faster under conversational pressure, or decohere across samples?)What follows is an update on the study, from August 10th through the 15th. This document is effectively a p

Comments (0)

Sign in to join the discussion

More Like This

Untie Squared ReLU variant
LessWrong · 2h ago
Q2.5 2026 Timelines Update: Uplift and Revenue
LessWrong · 14h ago
Case for Funding AI Safety in Japan
LessWrong · 14h ago
Will There Be an AI Hegemon? A Mental Model for AI Power Concentration
LessWrong · 14h ago
📰
The Doomsday Argument is Reasonable and Mostly Points to Longevity
LessWrong · 16h ago
Three thoughts on civilisational handoff
LessWrong · 16h ago