M

MiniMax-M3

Multimodalpor MiniMax·Página del modelo

MiniMax-M3 es un modelo multimodal de mezcla de expertos de 427B parámetros de MiniMax que admite entradas de imagen, vídeo y texto con capacidades de codificación y agentes.

Rankings
#3en OpenRouter9.2%
16864.2B tokens
Ver rankings
Salida máxima

Máximo de tokens que el modelo puede devolver en una sola respuesta.

512Ktokens
Share:

Design Arena

Design Arena clasifica modelos en tareas reales de front-end y diseño —sitios web, componentes de UI, visualización de datos, SVG y más— mediante votos humanos cara a cara, puntuados como rating ELO.
CategoríaELOTasa de victoriaPuesto
Code129055.0%#16
Website128955.1%#17
3D128755.4%#22
UI component128453.3%#20
Game dev128051.5%#23

Precios

por 1M tokens
Entrada$0.3 /1M
Salida$1.2 /1M
Lectura de caché$0.06 /1M

Descripción del Modelo


MiniMax-M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters.

Highlights:

  • Native Multimodality: M3 undergoes mixed-modality training from the very first step, enabling deeper semantic fusion across text, image, and video.
  • Context Scaling via Sparse Attention: M3 introduces MiniMax Sparse Attention (MSA) to improve long context efficiency. M3 delivers 9× prefill and 15× decode speedups compared to M2 at 1M context, reducing per-token compute to 1/20.
  • Coding & Cowork Capability: M3 achieves frontier-level performance across long-horizon agentic benchmarks, excelling in both coding and cowork.

MiniMax Sparse Attention (MSA)

M3 is powered by MiniMax Sparse Attention (MSA), a high-performance sparse attention operator designed for million-token contexts. Compared with GQA, MSA dramatically reduces the attention compute and memory footprint while preserving model quality.

📄 Read the technical report: arXiv:2606.13392 · Hugging Face Papers

How to Use

M3 supports three reasoning modes through the thinking parameter:

  • enabled — Reasoning is always enabled.
  • adaptive — M3 automatically determines when additional reasoning is beneficial.
  • disabled — Reasoning is disabled to minimize latency and maximize throughput.

Local Deployment

Download the model:

hf download MiniMaxAI/MiniMax-M3 --local-dir MiniMax-M3

We recommend the following inference frameworks (listed alphabetically) to serve the model:

Inference Parameters

We recommend the following parameters for best performance: temperature=1.0, top_p=0.95.

Contact Us

Contact us at model@minimax.io.

Autor
M
MiniMax
Organización
MiniMaxAI
Detalles
Descargas196.4K
Me gusta1.3K
AccesoCódigo Abierto
Contexto1M tokens
Precio entrada$0.3 /1M
Precio salida$1.2 /1M
Tareaimage-text-to-text
Parámetros427B
Tendencia41
Licenciaother
Libreríatransformers
Creado2 jun 2026
Actualizado1 jul 2026
Ver en Hugging Face
Benchmarks
Inteligencia44.4
Código58.6
Agéntico35.4
Entiende todo el contexto.

Regístrate para leer casos de estudio completos, acceder a métricas detalladas y recibir todos los reportes.