Qwen3.8-Flash-Next Previews Qwen4 Architecture With 6B Active Parameters

Covered by 5 sources
Read full post
Alibaba's Qwen team unveiled Qwen3.8-Flash-Next, an experimental 125B-parameter model activating only 6B per token, previewing the Qwen4 architecture focused on cost-efficient inference with hybrid attention and long context support.

Covered by 5 sources

More on this story


More in LLM & Text Generation

Peter Thiel-Backed AI Startup Cognition Raises Funds at $48 Billion Valuation

Covered by 2 sources

Build more natural voice experiences with GPT‑Live‑1 in the API

Covered by 2 sources