SyncAI.news, a Varaisys broadcasting
Qwen 3.8 Flash-Next is Cheap, But There Are Complicating Factors
ES

Esther Shittu

· 1 min read

BusinessAI Business

Qwen 3.8 Flash-Next is Cheap, But There Are Complicating Factors

Chinese AI tech giant Alibaba introduced a new variant of its flagship model on Wednesday, aiming to compete on price with U.S. vendors and to provide enterprises with lower-cost AI models. However, the latest Alibaba release serves as a reminder to enterprises that price is not always the main driver of model choice.

Qwen 3.8-Flash-Next is a multimodal model that provides an early preview of the architecture used in the upcoming Qwen 4. It is a 125B-parameter model, with an active 6B parameters per token and a mixture-of-experts (MoE) architecture. In comparison, Qwen 3.8 Max has 2.4 trillion parameters. Qwen 3.8-27B has 27B parameters. Rival Chinese AI vendor Moonshot’s Kimi K3 MoE model has 2.8 trillion parameters.

Alibaba highlighted that compared to Qwen 3.7-Plus, the Flash-Next version has lower training and inference costs. The model excels at computer use and can interact with complex APIs, calculators and custom external databases using visual and text prompts, according to the vendor

Qwen 3.8-Flash-Next is yet another example of model providers appealing to enterprises’ need for cheaper and more cost-efficient models, amid a price war among top AI vendors and escalating tension between open source and proprietary model providers. The vendor priced Flash-Next at $0.16 per million input tokens and $0.47 per million output tokens. It is open weight, unlike frontier models from OpenAI and Anthropic.

“They’ve tried to keep the model very competitive from a pricing perspective,” said Arun Chandrasekaran, an analyst at Gartner. He said that Alibaba can do that because of the model’s inference efficiency: although it contains many parameters, only a few are active, keeping inference costs low.

“They’re positioning this as a workhouse for a very broad category of enterprise workloads, where this model is super competitive,” Chandrasekaran said, noting that the model is optimized for agentic workloads such as tool calling and coding applications.

Original source

This story was published by AI Business and written by Esther Shittu. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on aibusiness.com

Similar News