🧩 Philosophy 2h ago · Amy_

Untie Squared ReLU variant

Less Wrong
View Channel →
Untie Squared ReLU variant
Source ↗ 👁 1 💬 0
1 IntroMy collaborator Michael Bukatin came up with an idea: in the hidden projection of MLP, instead of using ReLU, he would duplicate one branch, and initialize two matrices on the two branches separately, and then multiply the two element-wise, then project back as output in the feedforward layer.As a quick recap, a standard ReLU FFN is:

mjx-math {
display: inline-block;
text-align: left;
line-height: 0;
text-indent: 0;
font-style: normal;
font-weight: normal;
font-size: 100%;

Comments (0)

Sign in to join the discussion

More Like This

📰
Study Update: Does post-training quantization change welfare-relevant indicators in open-weight language models?
LessWrong · 12h ago
Q2.5 2026 Timelines Update: Uplift and Revenue
LessWrong · 14h ago
Case for Funding AI Safety in Japan
LessWrong · 14h ago
Will There Be an AI Hegemon? A Mental Model for AI Power Concentration
LessWrong · 14h ago
📰
The Doomsday Argument is Reasonable and Mostly Points to Longevity
LessWrong · 16h ago
Three thoughts on civilisational handoff
LessWrong · 16h ago