WayToClawEarn
High impact小米官方开放平台

Xiaomi MiMo-V2.5 API price cut by up to 99%: AI pricing war spreads, Agent developers usher in an era of ultra-low cost

Xiaomi MiMo-V2.5 has a permanent price reduction of up to 99% on all APIs, and the Token Plan quota has been increased by 5-8 times. The overseas output of mimo-v2.5-pro is only $0.87/M tokens, and the cache hit scenario input is as low as ¥0.025/M. The AI ​​pricing war between China and the United States is escalating again, and the API costs for Agent developers are falling off a cliff.

WayToClawEarn EditorialPublished May 27, 2026Updated Aug 8, 2026

Editorial review of public sources · AI-assisted drafting. How we work · Original source

Core conclusion

On May 27, 2026, Xiaomi MiMo Open Platform announced a permanent price reduction for the MiMo-V2.5 series API, with a maximum reduction of 99%. This is another price impact launched by Chinese AI model manufacturers after the permanent price reduction of DeepSeek V4 Pro by 75%.

With the Token Plan package, the usage is increased to 5-8 times, and the user quota of the existing validity period is fully reset. The overseas output price of MiMo-V2.5-pro is reduced to $0.87/M tokens (approximately ¥6.00/M), and the input in the cache hit scenario is only ¥0.025/M - this price is changing AI API calls from "cost consideration" to "casual use".

Key Points

  • Event Time: Effective at 00:00 Beijing time on 2026-05-27
  • Core changes: MiMo-V2.5 full-line API price reduction permanently, up to 99%, Token Plan quota doubled
  • Affected Audience: All developers using AI API, AI Agent automation operators, individual developers and small and medium-sized enterprises
  • Competitive Landscape: The price war between the Chinese camp (DeepSeek, MiMo) is intensifying, and the American camp (OpenAI, Anthropic) bucks the trend and increases prices

Background and Pricing Changes

Xiaomi will release the MiMo-V2.5 series of large models in the second half of 2025, including the mimo-v2.5-pro flagship version and the mimo-v2.5 standard version. After accumulating activities such as the "MiMo Orbit" Trillion Token Creator Incentive Program, Xiaomi's technical team said that it has optimized the underlying reasoning system before it dares to make "more thorough pricing adjustments."

Comparison of old and new prices (overseas pricing)

ModelScenarioOriginal price ($/M tokens)New price ($/M tokens)Decrease
mimo-v2.5-proOutput~$3.00 (v2-pro)$0.87~71%
mimo-v2.5-proinput (cache miss)~$1.00$0.435~57%
mimo-v2.5-proinput (cache hit)~$0.20$0.0036~98%
mimo-v2.5Output$0.28First price
mimo-v2.5Input$0.14Initial Price

For AI Agent users, the extremely low price in the cache hit scenario means that the cost of frequently calling tools with the same prompt prefix (such as Claude Code's system prompt) is almost negligible.

MiMo Opus

Inference technology optimization: confidence in price war

Xiaomi disclosed the technical route behind it in the announcement:

  • SWA (Sliding Window Attention): Based on the sliding window attention mechanism of SGLang HiCache, the data transfer volume of KV Cache between multi-level storage is reduced to about 1/7 of that before optimization.
  • Cache Capacity: The number of cacheable tokens has been increased to nearly 5 times that before optimization
  • Cluster throughput: Optimize expert parallel strategy and input length bucketing strategy to continuously reduce single Token service costs

These optimizations allow Xiaomi to significantly reduce inference costs while ensuring service quality, providing a technical foundation for price wars.

HN Community reaction and industry interpretation

This topic received 97 points and 101 comments on Hacker News (related discussion 60 points and 36 comments), and the community responded enthusiastically:

  • Performance Comparison: HN user irthomasthomas pointed out that MiMo is only 3 points lower than Opus on the Artificial Analysis benchmark, while the cost difference is a hundred times - MiMo is about $400/ months, Opus is about $5,000/ months at the old price
  • Geographic competition: Many commentators placed MiMo's price reduction in the context of Sino-US AI competition, believing that Chinese manufacturers strategically seize the market at low prices.
  • OpenRouter middleman problem: h4kunamata mentioned that the prices of third-party providers have not been reduced simultaneously, and questioned whether the middlemen have swallowed up the price reduction dividends.
  • Super large-scale Token Plan: User passive shared that his Token Plan was directly upgraded from 700 million tokens to 38 billion tokens/month

Practical impact on AI Agent automation

For WayToClawEarn readers – AI Agent users, automation operators – what this price reduction means:

  1. Agent cost drops off a cliff: If MiMo is used as the back-end model for programming Agent, the cost can be reduced by 70-90% in high-frequency calling scenarios, similar to the previously reported DeepSeek V4 Pro strategy
  2. The value of cache strategy is highlighted: MiMo’s SWA optimization is particularly effective in the Agent scenario (fixed system prompt + tool definition)
  3. Mixed models are more economical: MiMo is suitable for "lightweight Agent tasks" and can be used with Opus/Claude to handle complex tasks, which can significantly reduce the overall API cost.
  4. Third-party integration ready: Xiaomi open platform has supported the configuration integration of Claude Code, OpenClaw, Hermes Agent, OpenCode and other tools

Tool entry

Tool names that appear naturally in the text: DeepSeek, OpenAI, Claude, Claude Code, Hermes Agent, OpenClaw, OpenCode

Related reading

Update instructions

This article is based on Xiaomi MiMo open platform announcement on May 27, 2026 and Hacker News community discussion. Pricing information is subject to the latest official announcement from Xiaomi.

View source →

Disclaimer: this site shares educational insights only, for inspiration and reference. No outcome guarantee; external execution and decisions are your own responsibility.