Q

Qwen3.8-27B

Multimodalpor Qwen·Página del modelo

Modelo de visión y lenguaje de 27B de Qwen que procesa texto, imagen y video con un contexto de 1M de tokens.

Rankings
#56en OpenRouter86.7%
377B tokens
Ver rankings
Salida máxima

Máximo de tokens que el modelo puede devolver en una sola respuesta.

131Ktokens
Share:

Precios

por 1M tokens
Entrada$0.42 /1M
Salida$3 /1M
Lectura de caché$0.085 /1M

Descripción del Modelo


library_name: transformers license: apache-2.0 pipeline_tag: image-text-to-text

Qwen3.8-27B

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format.

These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc.

[!Tip] For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Qwen Cloud. In particular, Qwen3.8-27B will be available as a hosted version with more production features, e.g., 1M context length by default, official built-in tools. For more information, please refer to the Qwen3.8-27B Overview. The service is coming soon. Stay tuned for updates.

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date.

Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability.

Qwen3.8 Highlights

Qwen3.8-27B features the following enhancements:

  • Core Capabilities: Comprehensive improvements across coding, professional work, research, and long-horizon agentic tasks.
  • Agent Execution: Stronger autonomous planning and better handling of environment feedback, leading to more reliable end-to-end task completion.
  • Downstream Compatibility: Broader support for popular harnesses and development tools, making it easier to integrate into your existing stack.
  • Flexible Thinking Control: Thinking mode is on by default and can be disabled per request; reasoning depth can be tuned with reasoning_effort, and reasoning context from historical messages is retained via preserve_thinking.
  • Vision-Language Understanding: Native support for image and video understanding, from STEM diagrams and documents to hour-scale videos.

Model Overview

  • Type: Causal Language Model with Vision Encoder
  • Training Stage: Pre-training & Post-training
  • Language Model
    • Number of Parameters: 27B
    • Hidden Dimension: 5120
    • Token Embedding: 248,320 (Padded)
    • Number of Layers: 64
    • Hidden Layout: 16 × (3 × (Gated DeltaNet → FFN) → 1 × (Gated Attention → FFN))
    • Gated DeltaNet:
      • Number of Linear Attention Heads: 48 for V and 16 for QK
      • Head Dimension: 128
    • Gated Attention:
      • Number of Attention Heads: 24 for Q and 4 for KV
      • Head Dimension: 256
      • Rotary Position Embedding Dimension: 64
    • Feed Forward Network:
      • Intermediate Dimension: 17,408
    • LM Output: 248,320 (Padded)
    • MTP (Multi-Token Prediction): trained with multiple steps
  • Context Length: 262,144 natively and extensible up to 1,000,000 tokens.

Benchmark Results

Text Performance

73.0 63.4 64.0 51.7 78.2 61.7 53.5 57.6 51.2 53.4 42.3 36.2 41.1 -- 47.6 42.2 13.3 14.2 -- -- 79.0 49.3 59.2 -- 63.8 Agent 70.7 61.0 65.1 -- 68.2 33.4 21.8 27.6 -- -- -- -- General 79.5 69.1 79.1 77.0 62.5 89.2 87.8 90.3 83.5 91.3 30.8 24.0 34.7 22.0 40.0 90.3 83.9 89.6 -- 88.8

VL Performance

84.363.973.365.972.7 64.848.855.3---- 81.970.381.0--62.0 47.129.830.2---- -- 38.625.730.0--27.1 62.945.042.1---- General Multimodal Intelligence -- -- 78.8 91.189.491.475.886.6 85.984.186.9--73.9 65.562.569.8--40.8

Quickstart

For streamlined integration, we recommend using Qwen3.8 via APIs.

Serving Qwen3.8

[!Important] Inference efficiency and throughput vary significantly across frameworks. We recommend using the latest framework versions to ensure optimal performance and compatibility. For production workloads or high-throughput scenarios, dedicated serving engines such as SGLang, vLLM, or TokenSpeed are recommended.

Qwen3.8 can be deployed with popular inference frameworks, e.g.:

Autor
Q
Qwen
Organización
Qwen
Detalles
Descargas6.4M
Me gusta14.2K
AccesoCódigo Abierto
Contexto1M tokens
Precio entrada$0.42 /1M
Precio salida$3 /1M
Tareaimage-text-to-text
Parámetros27.8B
Tendencia511
Licenciaapache-2.0
Libreríatransformers
Creado5 ago 2026
Actualizado14 ago 2026
Ver en Hugging Face
Benchmarks
Inteligencia41.4
Código68.1
Agéntico46.8
Entiende todo el contexto.

Regístrate para leer casos de estudio completos, acceder a métricas detalladas y recibir todos los reportes.