Thursday, August 27, 2026HotTea verified storyVerified 2:36 AM PDT
← Back to the Thursday, August 27, 2026 edition

Alibaba and Z.ai opened lower-cost models

Alibaba opened Qwen3.8-Flash-Next, a multimodal mixture-of-experts model and early preview of its Qwen4 design. The company says six billion parameters activate per token, and training used about one-ninth the compute of Qwen3.7-Plus. Z.ai released GLM-5.3-Flash with open weights, 18 billion active parameters, and a one-million-token context window.

Verified 2:36 AM PDT · 4 original sources

The evidence

What the reporting establishes

What happened

Alibaba opened Qwen3.8-Flash-Next, a multimodal mixture-of-experts model and early preview of its Qwen4 design. The company says six billion parameters activate per token, and training used about one-ninth the compute of Qwen3.7-Plus. Z.ai released GLM-5.3-Flash with open weights, 18 billion active parameters, and a one-million-token context window.

Pressure point

The model makers supplied the efficiency and benchmark numbers. Open weights let outsiders test the models, but they do not remove the hardware needed to run them. Open weights also do not prove the claimed cost advantage in production.

What to watch

Watch independent cost and quality tests on the same hardware. Watch license terms, serving support, and evidence that Chinese chips can host GLM-5.3-Flash at the claimed scale.

Audit the story

Original sources

Company claims remain company claims. Follow the reporting and judge the evidence directly.

  1. Alibaba QwenQwen3.8-Flash-Next
  2. The DecoderAlibaba releases Qwen3.8-Flash-Next, targeting ultimate cost efficiency
  3. Z.aiGLM-5.3-Flash
  4. SiliconANGLEZ.ai open-sources Ox Alpha model as GLM-5.3-Flash

Continue the morning

Five stories. One sourced briefing.

Read the full editionListen to the daily audio →