Alibaba and Z.ai opened lower-cost models
Alibaba opened Qwen3.8-Flash-Next, a multimodal mixture-of-experts model and early preview of its Qwen4 design. The company says six billion parameters activate per token, and training used about one-ninth the compute of Qwen3.7-Plus. Z.ai released GLM-5.3-Flash with open weights, 18 billion active parameters, and a one-million-token context window.
Verified 2:36 AM PDT · 4 original sources
The evidence
What the reporting establishes
What happened
Alibaba opened Qwen3.8-Flash-Next, a multimodal mixture-of-experts model and early preview of its Qwen4 design. The company says six billion parameters activate per token, and training used about one-ninth the compute of Qwen3.7-Plus. Z.ai released GLM-5.3-Flash with open weights, 18 billion active parameters, and a one-million-token context window.
Pressure point
The model makers supplied the efficiency and benchmark numbers. Open weights let outsiders test the models, but they do not remove the hardware needed to run them. Open weights also do not prove the claimed cost advantage in production.
What to watch
Watch independent cost and quality tests on the same hardware. Watch license terms, serving support, and evidence that Chinese chips can host GLM-5.3-Flash at the claimed scale.
Audit the story
Original sources
Company claims remain company claims. Follow the reporting and judge the evidence directly.
Continue the morning