Zhipu Unveils GLM-5.1 Speed Version: 400 Tokens/s Sets New Global API Record

On May 22, Zhipu (02513.HK) made waves across both capital markets and the tech industry. Its stock surged over 22% during trading hours, pushing its market value above 450 billion HKD. That same day, the company launched a major new product for enterprise customers: GLM-5.1 Highspeed API (GLM-5.1-highspeed).
In real-world tests, this model delivers an impressive 400 tokens/s (400 tokens per second), breaking the current speed ceiling for official APIs from major global large model providers. To put that into perspective: the volume of text a creator could produce after days of nonstop work can now be completed in just a minute. A system reengineering task that would take an engineer three days of typing can be fully executed in the time it takes to drink a cup of coffee.
Key Highlights:
Breaking conventions: Traditionally, the industry assumed that "fast" meant "small or lightweight." Zhipu is the first domestic large model provider to achieve a seamless combination of **"full-scale flagship capabilities" and "extreme low latency."**
Hardcore achievements: Output speed reaches 400 tokens/s, supports a 200K long context window, and allows a maximum single output of 128K tokens.
Underlying black tech: Developed through deep collaboration between Zhipu's GLM team and the TileRT team, this model restructures the entire system-level inference ecosystem.
Targeted pilot test: Now available to select enterprise customers through Zhipu's MaaS (Model as a Service) open platform.
What Does "Instant Response" Actually Feel Like? A Breakthrough in Speed-Sensitive Scenarios
Over the past year, coding (programming) and agent (intelligence) collaboration capabilities in domestic large models have made significant strides. But "speed" has remained the core bottleneck for long-chain, high-frequency interaction tasks. Zhipu notes that the shift from large models as "tools" to "real-time partners" becomes truly transformative at 400 tokens/s:
AI Programming (Coding Agent): Traditional intelligent agents often require dozens of cross-file calls and long-text alignment. A single round of response lagging by a few seconds can stretch the entire task to over ten minutes. With the high-speed version, writing code feels like running at 10x speed — functions, interfaces, and underlying call chains unfold instantly as the user types, eliminating any waiting time for large-scale engineering rework.
Real-time interaction and 3D games: Extreme low latency allows the model to handle real-time dynamic generation within game worlds, instant web UI construction, and immediate system state changes based on continuous user input — all without lag.
Business decision clusters: In multi-agent parallel simulation and real-time big data analysis, the high-speed version supports completing complex web agent cluster multi-personality parallel responses within 30 seconds, significantly raising the efficiency ceiling for high-frequency quantification and simulation.
Seamless real-time voice: In AI coaching and smart customer service scenarios, ultra-fast response brings latency from speech recognition (ASR) to synthesis (TTS) close to zero, delivering truly natural, equal-paced human-like conversation flow.
Decoding the Three Layers of Black Tech: How Did We Reach 400 Tokens/s?
This global speed record is primarily the result of system-level engineering optimization jointly developed by Zhipu's GLM team and the TileRT team. The 400 tokens/s is not a flashy "peak moment" — it's a stable, production-ready capability. The underlying optimization logic breaks down into three layers:
[Infrastructure Layer: Cluster/Load Balancing Collaboration] ─► [Scheduling System Layer: Dynamic Batching & KV Cache Scheduling] ─► [Inference Engine Layer: TileRT Architecture Rewriting Core Path] ─► 400 tokens/s Stable Output
Inference Engine Layer (TileRT Deep Customization): Based on the unique network architecture of GLM-5.1, the team completely rewrote the most critical inference path and underlying operators, allowing single GPU throughput and hardware execution efficiency to approach physical limits.
Scheduling System Layer (Intelligent Merging): Introduced highly aggressive dynamic batching, request merging technologies, and groundbreaking KV cache scheduling optimization, effectively solving the tail latency issues that traditional models face under high concurrency and multiple user requests.
Infrastructure Layer (Cluster Collaboration): Comprehensive hardware-level collaborative optimization around networking deployment, network link topology, and ultra-high-frequency load balancing of the inference cluster ensures computing power is transmitted without loss across the entire pipeline.
Industry Reassessment: The Second Half of AI Is About "Value and Time" Settlement
As top international analytical institutions like UBS emphasized at recent Hong Kong stock tech forums, this round of AI-driven industry reassessment is fundamentally different from the "traffic and time monetization" of the mobile internet era. AI's revenue and survival philosophy is not about keeping users trapped in software — it's about "helping users and enterprises save time, improve efficiency, and share the value created."
Zhipu's GLM-5.1 Highspeed Version directly addresses this pain point. By compressing the cost and time of producing each token to a fraction of the original, it allows enterprises to stop making painful trade-offs between "high intelligence (choosing a large model but slow)" and "speed (choosing a small model but dull)."
Related article
Musk's xAI Halts Recruitment of AI Mentors Amid HR Overload
Elon Musk’s AI venture, xAI, has recently introduced an unexpected operational shift. Citing informed sources, Bloomberg reports that the company has temporarily halted recruitment for specialists tasked with training and refining its Grok chatbot. T
How to fix Core Web Vitals for better SEO rankings
How to Build a Free Website: Free Domain, Hosting & AI Website BuilderTable of ContentsIntroductionAI Website Builder 1: Hookous AIStep 1: Sign UpStep 2: Choose Business Type/NicheStep 3: Select ServicesStep 4: Define Website GoalsStep 5: Enter Busin
Apple, Google Partner With Anthropic to Address 27-Year-Old Vulnerability via Glass Wing Protection
As artificial intelligence advances rapidly in code generation and logical reasoning, the cybersecurity landscape faces unprecedented challenges. Recently, the prominent AI startup Anthropic officially launched a cross-industry collaboration called *
Related Special Topic Recommendations
Comments (0)
0/500

On May 22, Zhipu (02513.HK) made waves across both capital markets and the tech industry. Its stock surged over 22% during trading hours, pushing its market value above 450 billion HKD. That same day, the company launched a major new product for enterprise customers: GLM-5.1 Highspeed API (GLM-5.1-highspeed).
In real-world tests, this model delivers an impressive 400 tokens/s (400 tokens per second), breaking the current speed ceiling for official APIs from major global large model providers. To put that into perspective: the volume of text a creator could produce after days of nonstop work can now be completed in just a minute. A system reengineering task that would take an engineer three days of typing can be fully executed in the time it takes to drink a cup of coffee.
Key Highlights:
Breaking conventions: Traditionally, the industry assumed that "fast" meant "small or lightweight." Zhipu is the first domestic large model provider to achieve a seamless combination of **"full-scale flagship capabilities" and "extreme low latency."**
Hardcore achievements: Output speed reaches 400 tokens/s, supports a 200K long context window, and allows a maximum single output of 128K tokens.
Underlying black tech: Developed through deep collaboration between Zhipu's GLM team and the TileRT team, this model restructures the entire system-level inference ecosystem.
Targeted pilot test: Now available to select enterprise customers through Zhipu's MaaS (Model as a Service) open platform.
What Does "Instant Response" Actually Feel Like? A Breakthrough in Speed-Sensitive Scenarios
Over the past year, coding (programming) and agent (intelligence) collaboration capabilities in domestic large models have made significant strides. But "speed" has remained the core bottleneck for long-chain, high-frequency interaction tasks. Zhipu notes that the shift from large models as "tools" to "real-time partners" becomes truly transformative at 400 tokens/s:
AI Programming (Coding Agent): Traditional intelligent agents often require dozens of cross-file calls and long-text alignment. A single round of response lagging by a few seconds can stretch the entire task to over ten minutes. With the high-speed version, writing code feels like running at 10x speed — functions, interfaces, and underlying call chains unfold instantly as the user types, eliminating any waiting time for large-scale engineering rework.
Real-time interaction and 3D games: Extreme low latency allows the model to handle real-time dynamic generation within game worlds, instant web UI construction, and immediate system state changes based on continuous user input — all without lag.
Business decision clusters: In multi-agent parallel simulation and real-time big data analysis, the high-speed version supports completing complex web agent cluster multi-personality parallel responses within 30 seconds, significantly raising the efficiency ceiling for high-frequency quantification and simulation.
Seamless real-time voice: In AI coaching and smart customer service scenarios, ultra-fast response brings latency from speech recognition (ASR) to synthesis (TTS) close to zero, delivering truly natural, equal-paced human-like conversation flow.
Decoding the Three Layers of Black Tech: How Did We Reach 400 Tokens/s?
This global speed record is primarily the result of system-level engineering optimization jointly developed by Zhipu's GLM team and the TileRT team. The 400 tokens/s is not a flashy "peak moment" — it's a stable, production-ready capability. The underlying optimization logic breaks down into three layers:
[Infrastructure Layer: Cluster/Load Balancing Collaboration] ─► [Scheduling System Layer: Dynamic Batching & KV Cache Scheduling] ─► [Inference Engine Layer: TileRT Architecture Rewriting Core Path] ─► 400 tokens/s Stable Output
Inference Engine Layer (TileRT Deep Customization): Based on the unique network architecture of GLM-5.1, the team completely rewrote the most critical inference path and underlying operators, allowing single GPU throughput and hardware execution efficiency to approach physical limits.
Scheduling System Layer (Intelligent Merging): Introduced highly aggressive dynamic batching, request merging technologies, and groundbreaking KV cache scheduling optimization, effectively solving the tail latency issues that traditional models face under high concurrency and multiple user requests.
Infrastructure Layer (Cluster Collaboration): Comprehensive hardware-level collaborative optimization around networking deployment, network link topology, and ultra-high-frequency load balancing of the inference cluster ensures computing power is transmitted without loss across the entire pipeline.
Industry Reassessment: The Second Half of AI Is About "Value and Time" Settlement
As top international analytical institutions like UBS emphasized at recent Hong Kong stock tech forums, this round of AI-driven industry reassessment is fundamentally different from the "traffic and time monetization" of the mobile internet era. AI's revenue and survival philosophy is not about keeping users trapped in software — it's about "helping users and enterprises save time, improve efficiency, and share the value created."
Zhipu's GLM-5.1 Highspeed Version directly addresses this pain point. By compressing the cost and time of producing each token to a fraction of the original, it allows enterprises to stop making painful trade-offs between "high intelligence (choosing a large model but slow)" and "speed (choosing a small model but dull)."
How to fix Core Web Vitals for better SEO rankings
How to Build a Free Website: Free Domain, Hosting & AI Website BuilderTable of ContentsIntroductionAI Website Builder 1: Hookous AIStep 1: Sign UpStep 2: Choose Business Type/NicheStep 3: Select ServicesStep 4: Define Website GoalsStep 5: Enter Busin
Apple, Google Partner With Anthropic to Address 27-Year-Old Vulnerability via Glass Wing Protection
As artificial intelligence advances rapidly in code generation and logical reasoning, the cybersecurity landscape faces unprecedented challenges. Recently, the prominent AI startup Anthropic officially launched a cross-industry collaboration called *





Home






