Xiaomi's MiMo-V2.5 native omnimodal model with strong agentic capabilities, supporting text, image, video, and audio understanding within a unified architecture. Built on the MiMo-V2-Flash backbone with dedicated vision and audio encoders.
Model id — paste this in your code
xiaomi/mimo-v2.5
Pricing · per 1M tokens
Show prices in
Input50.4 DA$0.202
Output101 DA$0.403
Cached input1.02 DA$0.00408
Context window1.05MMax output: 131K
Capabilities
Vision — Understands images and documents you send
Reasoning — Thinks step by step before answering
Tool use — Can call your functions / tools
Estimate your cost
Estimated cost
Use this model
Create an account, top up in dinars and call it with your key: