Study Update: Does post-training quantization change welfare-relevant indicators in open-weight language models?
Source ↗
👁 0
💬 0
This content will not make sense without reading the preregistration found in the original postWe ask whether welfare-relevant indicators change with quantization; either in valence (do indicators shift toward more negative / more distressed / more boundary-eroded states?) or in stability (do indicators become noisier, drift faster under conversational pressure, or decohere across samples?)What follows is an update on the study, from August 10th through the 15th. This document is effectively a p
Comments (0)