X

MiMo-V2.5

Multimodalby Xiaomi·Model page

Xiaomi's MiMo V2.5 multimodal model with a 1M-token context for text, audio, image, and video inputs.

Rankings
#1on OpenRouter72.1%
24587.6B tokens
View rankings
Max output

Most tokens the model can return in a single response.

131Ktokens
Share:

Design Arena

Design Arena ranks models on real-world front-end and design tasks — websites, UI components, data viz, SVG and more — through head-to-head human votes, scored as an ELO rating.
CategoryELOWin rateRank
UI component129855.3%#17
Game dev129155.4%#19
Website129155.0%#15
Code128954.6%#18
3D128152.9%#23

Pricing

per 1M tokens
Input$0.14 /1M
Output$0.28 /1M
Cache read$0.003 /1M

Model Description

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...

Author
X
Xiaomi
Organization
xiaomi
Details
Downloads
Likes
AccessOpen Source
Context1M tokens
Input price$0.14 /1M
Output price$0.28 /1M
CreatedApr 22, 2026
Updated
View on Hugging Face
Benchmarks
Intelligence37.2
Coding56.8
Agentic23.7
Get the full context.

Sign up to read complete case studies, access detailed metrics, and unlock all use cases.

MiMo-V2.5 — AI Model Details | Applied