September 26, 2026

Alibaba’s Qwen3.8-Flash-Next Previews Qwen4 Architecture

Alibaba's Qwen team unveiled Qwen3.8-Flash-Next, a mixture-of-experts model that previews Qwen4's architecture while undercutting rivals on training cost and price.
Alibaba’s Qwen3.8-Flash-Next Previews Qwen4 Architecture

Alibaba’s Qwen team has released Qwen3.8-Flash-Next, a multimodal mixture-of-experts model described as an architecture preview of the upcoming Qwen4, according to the-decoder.com. The model has 125 billion total parameters but activates only 6 billion per token, and includes a new 51 billion parameter N-gram embedding layer that stores common word groups in a kind of phrase dictionary. This layer can run on regular system RAM instead of GPU memory at relatively low added cost, one of the architectural innovations planned for Qwen4.

The model natively supports a 262,144 token context window, expandable to one million tokens using YaRN. Its technical report is posted on GitHub, with weights available on Hugging Face and ModelScope. The production version, Qwen3.8-Flash, is offered through QwenCloud at $0.16 per million input tokens and $0.47 per million output tokens, with API access expected soon.

According to Qwen’s published benchmarks, Flash-Next outperforms the larger Qwen3.7-Plus model at roughly one-ninth the training cost, with the biggest gains in coding and office productivity tasks. It also outperforms DeepSeek-V4-Flash and Anthropic’s Claude Opus 4.6 on most tested benchmarks, including agentic coding tasks like DeepSWE and SWE-bench Pro, and office tasks like CoWorkBench and JobBench. Claude Opus 4.6 leads only on Humanity’s Last Exam, a benchmark of difficult multidisciplinary problems, though the-decoder.com notes it is an older Anthropic model.

Flash-Next performs just below Alibaba’s flagship Qwen3.8-Max, introduced in early August, while costing about one-twelfth as much. The-decoder.com frames the pricing as continued pressure on OpenAI and Anthropic, noting OpenAI recently cut prices on its GPT-5.6 line in response to competition from Chinese providers.

Based on reporting by the-decoder.com.

Leave a Reply

Your email address will not be published. Required fields are marked *