Xiaomi MiMo-V2.5 Series API Gets Permanent Price Cut, Up to 99% Off
Amid the escalating AI model price wars, Xiaomi officially announced on May 27 that its MiMo large model would permanently reduce prices for the MiMo-V2.5 series API while simultaneously optimizing the billing system to further lower developers' calling costs through technological advancements.

I. Significant API Price Cuts — Up to 99% Off
The price change took effect globally at 00:00 Beijing Time on May 27. It applies to the two core versions, MiMo-V2.5 and MiMo-V2.5Pro, and no longer differentiates based on context window length, simplifying the pricing structure for greater transparency.
Model VersionInput Cache Hit PriceMaximum DiscountOutput PriceMaximum DiscountMiMo-V2.5Pro0.025 yuan per million tokens, up to 99% off; output: 6 yuan per million tokens, up to 86% offMiMo-V2.50.02 yuan per million tokens, up to 98% off; output: 2 yuan per million tokens, up to 93% offII. Billing System Upgrade — More Value at No Extra Cost
Beyond the direct API price cuts, Xiaomi has heavily optimized its Token Plan billing system:
Quadrupled Quota: Under the original pricing, the actual token usage quota has increased to 5 to 8 times the previous amount.
Simplified Rules: The introduction of Credits replaces the previous complex billing methods, making token consumption and cost calculation more intuitive for developers.

III. Technical Foundation — How Can It Keep Lowering Prices?
Xiaomi's official statement attributes these deep price cuts to technical breakthroughs in its underlying inference system architecture:
SWA Inference Optimization: By leveraging SGLang HiCache with full support for SWA (Sliding Window Attention Mechanism), the data transfer among GPU memory, CPU memory, and SSD has been reduced to one-seventh of the previous volume.
Improved Cache Efficiency: The number of cacheable tokens has increased nearly fivefold compared to the earlier optimized version, boosting cache hit rates and dramatically lowering per-inference cost.
Cluster Throughput Optimization: With the introduction of expert parallel (MoE) and input length bucketing strategies, the cluster's input throughput has seen a qualitative leap, maintaining high service quality while steadily reducing cost per token.
Xiaomi's move is seen as a proactive response to the current intense competition in large model commercialization. As price barriers continue to drop, the MiMo series' cost-effectiveness will become even more pronounced, accelerating the deep integration of AI capabilities across vertical industries and developer workflows.
Related article
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control
According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
Related Special Topic Recommendations
Comments (0)
0/500
Amid the escalating AI model price wars, Xiaomi officially announced on May 27 that its MiMo large model would permanently reduce prices for the MiMo-V2.5 series API while simultaneously optimizing the billing system to further lower developers' calling costs through technological advancements.

I. Significant API Price Cuts — Up to 99% Off
The price change took effect globally at 00:00 Beijing Time on May 27. It applies to the two core versions, MiMo-V2.5 and MiMo-V2.5Pro, and no longer differentiates based on context window length, simplifying the pricing structure for greater transparency.
Model VersionInput Cache Hit PriceMaximum DiscountOutput PriceMaximum DiscountMiMo-V2.5Pro0.025 yuan per million tokens, up to 99% off; output: 6 yuan per million tokens, up to 86% offMiMo-V2.50.02 yuan per million tokens, up to 98% off; output: 2 yuan per million tokens, up to 93% offII. Billing System Upgrade — More Value at No Extra Cost
Beyond the direct API price cuts, Xiaomi has heavily optimized its Token Plan billing system:
Quadrupled Quota: Under the original pricing, the actual token usage quota has increased to 5 to 8 times the previous amount.
Simplified Rules: The introduction of Credits replaces the previous complex billing methods, making token consumption and cost calculation more intuitive for developers.

III. Technical Foundation — How Can It Keep Lowering Prices?
Xiaomi's official statement attributes these deep price cuts to technical breakthroughs in its underlying inference system architecture:
SWA Inference Optimization: By leveraging SGLang HiCache with full support for SWA (Sliding Window Attention Mechanism), the data transfer among GPU memory, CPU memory, and SSD has been reduced to one-seventh of the previous volume.
Improved Cache Efficiency: The number of cacheable tokens has increased nearly fivefold compared to the earlier optimized version, boosting cache hit rates and dramatically lowering per-inference cost.
Cluster Throughput Optimization: With the introduction of expert parallel (MoE) and input length bucketing strategies, the cluster's input throughput has seen a qualitative leap, maintaining high service quality while steadily reducing cost per token.
Xiaomi's move is seen as a proactive response to the current intense competition in large model commercialization. As price barriers continue to drop, the MiMo series' cost-effectiveness will become even more pronounced, accelerating the deep integration of AI capabilities across vertical industries and developer workflows.
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur





Home






