Q

Qwen3.8-Flash-Next

Multimodalby Qwen·Model page

Qwen's 180B MoE vision-language model with a 1M-token context, built for fast multimodal chat.

Rankings
#71on OpenRouter
124.8B tokens
View rankings
Max output

Most tokens the model can return in a single response.

131Ktokens
Share:

Pricing

per 1M tokens
Input$0.15 /1M
Output$0.47 /1M
Cache read$0.016 /1M
Cache write$0.2 /1M

Model Description

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format.

These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc.

[!Tip] For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Qwen Cloud.

In particular, Qwen3.8-Flash is the official version based on Qwen3.8-Flash-Next with more production features, e.g., 1M context length by default, official built-in tools. For more information, please refer to the Qwen3.8-Flash Overview.

As the frontier of foundation models pushes toward ever-larger parameter counts and ever-longer context windows, the question is no longer just how much we can scale, but how efficiently we can do so. Sustainable progress toward artificial general intelligence (AGI) that benefits everyone demands architectural innovation. Today, we are sharing a concrete step in that direction: Qwen3.8-Flash-Next.

Qwen3.8-Flash-Next Architecture

This experimental preview of the architecture that will underpin Qwen4 is built around a fundamental rethinking of how the core components of modern large language models (LLMs) interact at scale.

Highlights

The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces:

  • Hybrid Attention with QSA: The Gated DeltaNet and Gated Attention pairing has been reworked into Gated DeltaNet and Qwen Sparse Attention (QSA). Rather than selecting individual tokens for processing, QSA operates at the micro-block level. This cuts long-context latency significantly, a critical gain as agentic workloads increasingly dominate real-world usage.
  • Gated Residual: Residual streams with normalization are what make deep LLM training manageable. Gated Residual modulates information flowing through widened residual streams via an element-wise, data-dependent read gate and a per-branch scalar write gate. This brings finer-grained expressiveness across layers while preserving training stability and keeping inference overhead low.
  • N-gram Embedding: Embeddings provide a unique axis for parameter scaling that requires less computation and is more amenable to offloading than Mixture-of-Experts (MoE). By indexing with short n-grams, this approach makes parameter scaling highly efficient for memory-constrained accelerators without sacrificing quality.
  • Tailored Training Recipe: The Muon and AdamW optimizers are applied to specific weight categories to maximize efficiency. Guided by refitted scaling laws, we eliminate traditional batch-size warmups and start directly at the target batch size, substantially reducing total optimizer steps while safely supporting larger learning rates for robust convergence.

For more details, please refer to our blog post Qwen3.8-Flash-Next and the technical report.

We are excited to embark on this next chapter with you and welcome your feedback as we build what comes next.

Model Overview

  • Type: Causal Language Model with Vision Encoder
  • Training Stage: Pre-training & Post-training
  • Language Model
    • Number of Parameters: 125B with 6B activated, plus 51B n-gram embedding and 4B MTP
    • Hidden Dimension: 2560
    • Token Embedding: 248320 (Padded)
    • N-gram Embedding: 20,000,000 (bigrams/trigrams at layer 2)
    • Number of Layers: 48
    • Hidden Layout: 12 × (3 × (Gated DeltaNet → MoE) → 1 × (Qwen Sparse Attention → MoE))
    • Gated DeltaNet:
      • Number of Linear Attention Heads: 48 for V and 16 for QK
      • Head Dimension: 128
    • Qwen Sparse Attention:
      • Number of Attention Heads: 24 for Q and 2 for KV
      • Head Dimension: 256
      • Rotary Position Embedding Dimension: 64
      • Indexer Structure: MQA with 4 Query Heads and 1 Shared Key Head
      • Indexer Head Dimension: 128
      • Budget: 512 blocks or 2048 tokens
    • Mixture Of Experts
      • Number of Experts: 512
      • Number of Activated Experts: 10 Routed + 1 Shared
      • Expert Intermediate Dimension: 640
    • Gated Residual:
      • Number of Branches: 4
      • Bottleneck Rank: 320
    • LM Output: 248320 (Padded)
    • MTP: 1 layer, trained with multi-steps
  • Context Length: 262,144 natively and extensible up to 1,000,000 tokens.

Benchmark Results

Language

125B 27B 397B 284B -- 6B 27B 17B 13B -- 51B -- -- -- -- Coding 58.7 42.2 16.5 54.4 -- 62.5 61.7 55.8 56.0 53.4 81.0 73.8 75.8 -- 77.5 48.1 42.3 41.1 54.2 47.6 Agent 73.9 70.7 65.1 45.1 68.2 55.7 33.4 27.6 41.3 36.6 -- 73.5 67.1 50.6 70.3 -- General 81.3 79.5 79.1 79.2 62.5 91.7 89.2 90.3 90.8 91.3 35.9 30.8 34.7 33.8 40.0 91.9 90.3 89.6 90.6 88.8

Vision Language

49.9 47.1 30.2 -- 84.5 81.9 81.0 62.0 -- 64.0 62.9 42.1 -- General Multimodal Intelligence 72.3 65.5 69.8 40.8 76.6 72.4 76.2 63.0 88.5 85.9 86.9 73.9

Quickstart

For streamlined integration, we recommend using Qwen3.8-Flash-Next via APIs.

Serving Qwen3.8-Flash-Next

[!Important] Inference efficiency and throughput vary significantly across frameworks. We recommend using the latest framework versions to ensure optimal performance and compatibility. For production workloads or high-throughput scenarios, dedicated serving engines such as SGLang, KTransformers or vLLM are strongly recom

Author
Q
Qwen
Organization
Qwen
Details
Downloads474.7K
Likes5K
AccessOpen Source
Context1M tokens
Input price$0.15 /1M
Output price$0.47 /1M
Taskimage-text-to-text
Parameters180B
Trending354
Licenseother
Librarytransformers
CreatedAug 24, 2026
UpdatedAug 27, 2026
View on Hugging Face
Get the full context.

Sign up to read complete case studies, access detailed metrics, and unlock all use cases.

Qwen3.8-Flash-Next — AI Model Details | Applied