Home: Motoring > Alibaba's Tongyi Qianwen Team Open-Sources Qwen3.8-Flash Multimodal AI Model

Alibaba's Tongyi Qianwen Team Open-Sources Qwen3.8-Flash Multimodal AI Model

From:Internet Info Agency 2026-08-26 21:01:00

On August 26, Alibaba's Qwen team announced the release of Qwen3.8-Flash, open-sourcing its model weights on Hugging Face and ModelScope platforms at 23:00 that day, along with an FP8 quantized version. The model features a multimodal Mixture-of-Experts (MoE) architecture with 125 billion main model parameters, complemented by a 51-billion-parameter N-gram Embedding module. It activates 6 billion parameters per token and natively supports a context length of 262,144 tokens, extendable to 1 million tokens via YaRN. Qwen3.8-Flash introduces systematic upgrades across four key areas—Attention, Residual connections, Embedding, and Optimization: it adopts a hybrid GDN+QSA attention architecture; integrates a GatedResidual mechanism supporting FP8 storage; employs a 51-billion-parameter N-gram Embedding; and applies the MuonOptimizer alongside multiple refined training strategies. Compared to Qwen3.7-Plus, it significantly reduces training costs and demonstrates stronger performance on coding and office-related tasks. The model will soon be available via API on the Qwen AI platform, priced at RMB 1 per million input tokens and RMB 3 per million output tokens. Also released on the same day, Qwen3.8-Flash-Next defaults to 1M-token context support and includes official tools built-in.

Editor:NewsAssistant