WeChat AI Team Details WeLM Models Scaling to 617B Parameters
Source ↗
👁 0
💬 0
Tencent’s WeChat AI team has detailed a new scaling approach for its WeLM model family. The team trained WeLM-HD4-80B and WeLM-HD4-617B models using a method called Hidden Decoding, which expands each token into multiple internal computation streams without increasing the main Transformer backbone.
The 80B model activates 3 billion parameters, while the 617B model activates 23 billion. Both models outperformed their matched autoregressive baselines across nine shared benchmarks in the team’s tes
The 80B model activates 3 billion parameters, while the 617B model activates 23 billion. Both models outperformed their matched autoregressive baselines across nine shared benchmarks in the team’s tes
Comments (0)