¿Quién creó DeepSeek-V4-Pro?

DeepSeek-V4-Pro fue publicado por DeepSeek en Hugging Face.

DeepSeek-V4-Pro

Name: DeepSeek-V4-Pro
Author: DeepSeek

DeepSeek-V4-Pro es un modelo de lenguaje de 861.000 millones de parámetros de DeepSeek para generación de texto conversacional con precisión FP8.

Rankings

#6en OpenRouter▼11.6%

10474.4B tokens

Ver rankings →

Salida máxima

Máximo de tokens que el modelo puede devolver en una sola respuesta.

384Ktokens

Design Arena

Categoría	ELO	Tasa de victoria	Puesto
3D	1319	59.3%	#10
Game dev	1288	55.8%	#21
Code	1275	54.3%	#23
UI component	1266	52.1%	#28
Website	1265	53.0%	#26

Precios

por 1M tokens

Entrada$0.435 /1M

Salida$0.87 /1M

Lectura de caché$0.004 /1M

Descripción del Modelo

Introduction

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens.

DeepSeek-V4 series incorporate several key upgrades in architecture and optimization:

Hybrid Attention Architecture: We design a hybrid attention mechanism combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to dramatically improve long-context efficiency. In the 1M-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2.
Manifold-Constrained Hyper-Connections (mHC): We incorporate mHC to strengthen conventional residual connections, enhancing stability of signal propagation across layers while preserving model expressivity.
Muon Optimizer: We employ the Muon optimizer for faster convergence and greater training stability.

We pre-train both models on more than 32T diverse and high-quality tokens, followed by a comprehensive post-training pipeline. The post-training features a two-stage paradigm: independent cultivation of domain-specific experts (through SFT and RL with GRPO), followed by unified model consolidation via on-policy distillation, integrating distinct proficiencies across diverse domains into a single model.

DeepSeek-V4-Pro-Max, the maximum reasoning effort mode of DeepSeek-V4-Pro, significantly advances the knowledge capabilities of open-source models, firmly establishing itself as the best open-source model available today. It achieves top-tier performance in coding benchmarks and significantly bridges the gap with leading closed-source models on reasoning and agentic tasks. Meanwhile, DeepSeek-V4-Flash-Max achieves comparable reasoning performance to the Pro version when given a larger thinking budget, though its smaller parameter scale naturally places it slightly behind on pure knowledge tasks and the most complex agentic workflows.

Model Downloads

*FP4 + FP8 Mixed: MoE expert parameters use FP4 precision; most other parameters use FP8.

Evaluation Results

Base Model

Instruct Model

DeepSeek-V4-Pro and DeepSeek-V4-Flash both support three reasoning effort modes:

Reasoning Mode	Characteristics	Typical Use Cases	Response Format
Non-think	Fast, intuitive responses	Routine daily tasks, low-risk decisions	`</think>` summary
Think High	Conscious logical analysis, slower but more accurate	Complex problem-solving, planning	`<think>` thinking `</think>` summary
Think Max	Push reasoning to its fullest extent	Exploring the boundary of model reasoning capability	Special system prompt + `<think>` thinking `</think>` summary

DeepSeek-V4-Pro-Max vs Frontier Models

Comparison across Modes

Chat Template

This release does not include a Jinja-format chat template. Instead, we provide a dedicated encoding folder with Python scripts and test cases demonstrating how to encode messages in OpenAI-compatible format into input strings for the model, and how to parse the model's text output. Please refer to the encoding folder for full documentation.

A brief example:

from encoding_dsv4 import encode_messages, parse_message_from_completion_text

messages = [
    {"role": "user", "content": "hello"},
    {"role": "assistant", "content": "Hello! I am DeepSeek.", "reasoning_content": "thinking..."},
    {"role": "user", "content": "1+1=?"}
]

# messages -> string
prompt = encode_messages(messages, thinking_mode="thinking")

# string -> tokens
import transformers
tokenizer = transformers.AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-V4-Pro")
tokens = tokenizer.encode(prompt)

How to Run Locally

Please refer to the inference folder for detailed instructions on running DeepSeek-V4 locally, including model weight conversion and interactive chat demos.

For local deployment, we recommend setting the sampling parameters to temperature = 1.0, top_p = 1.0. For the Think Max reasoning mode, we recommend setting the context window to at least 384K tokens.

License

This repository and the model weights are licensed under the MIT License.

Citation

@misc{deepseekai2026deepseekv4,
      title={DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence},
      author={DeepSeek-AI},
      year={2026},
}

Contact

If you have any questions, please raise an issue or contact us at service@deepseek.com.

Autor

DeepSeek

Organización · ✓

deepseek-ai

Detalles

Descargas1.2M

Me gusta5.1K

AccesoCódigo Abierto

Contexto1M tokens

Precio entrada$0.435 /1M

Precio salida$0.87 /1M

Tareatext-generation

Parámetros862B

Tendencia47

Licenciamit

Libreríatransformers

Creado22 abr 2026

Actualizado22 jun 2026

Ver en Hugging Face

Benchmarks

Inteligencia44.3

Código59.4

Agéntico36.4

Entiende todo el contexto.

Regístrate para leer casos de estudio completos, acceder a métricas detalladas y recibir todos los reportes.