X

MiMo-V2.5

Multimodalby Xiaomi·Model page

Xiaomi's MiMo V2.5 multimodal model with a 1M-token context for text, audio, image, and video inputs.

Rankings
#5on OpenRouter74.3%
26476.7B tokens
View rankings
Max output

Most tokens the model can return in a single response.

131Ktokens
Share:

Design Arena

Design Arena ranks models on real-world front-end and design tasks — websites, UI components, data viz, SVG and more — through head-to-head human votes, scored as an ELO rating.
CategoryELOWin rateRank
Website128154.4%#24
UI component128054.9%#25
Code127854.3%#24
Game dev127254.7%#26
Data viz127055.0%#21

Pricing

per 1M tokens
Input$0.14 /1M
Output$0.28 /1M
Cache read$0.003 /1M

Model Description

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...

Author
X
Xiaomi
Organization
xiaomi
Details
Downloads
Likes
AccessOpen Source
Context1.1M tokens
Input price$0.14 /1M
Output price$0.28 /1M
CreatedApr 22, 2026
Updated
View on Hugging Face
Benchmarks
Coding56.8
Get the full context.

Sign up to read complete case studies, access detailed metrics, and unlock all use cases.

MiMo-V2.5 — AI Model Details | Applied