💻 Technology Aug 13, 2026 · TechNode Feed

WeChat AI Team Details WeLM Models Scaling to 617B Parameters

TechNode
TechNode China tech
View Channel →
WeChat AI Team Details WeLM Models Scaling to 617B Parameters
Source ↗ 👁 16 💬 0
Tencent’s WeChat AI team has detailed a new scaling approach for its WeLM model family. The team trained WeLM-HD4-80B and WeLM-HD4-617B models using a method called Hidden Decoding, which expands each token into multiple internal computation streams without increasing the main Transformer backbone.
The 80B model activates 3 billion parameters, while the 617B model activates 23 billion. Both models outperformed their matched autoregressive baselines across nine shared benchmarks in the team’s tes

Comments (0)

Sign in to join the discussion

More Like This

📰
Placeholder domain used in dev docs now serves ClickFix attacks
BleepingComputer · Sep 23, 2026
Oracle shifts agent controls toward data-layer security
SiliconANGLE · Sep 23, 2026
📰
New RemControl Android banking malware targets users in Europe and Canada
BleepingComputer · Sep 23, 2026
Workforce economics emerges as AI reshapes the C-suite
SiliconANGLE · Sep 23, 2026
📰
Debian Inference Portal Launches To Provide Free AI/LLM Inferencing To Debian Developers
Phoronix · Sep 23, 2026
📰
Qualcomm Talks Up Linux On Snapdragon X2 Laptops
Phoronix · Sep 23, 2026