ZhiPu and TileRT Launch GLM-5.1 High-Speed API at 400 Tokens/s, Setting New Global Record
Zhipu has officially launched the GLM-5.1 Highspeed API (GLM-5.1-highspeed) for select enterprise customers today. This model achieves a remarkable output speed of 400 tokens/s, surpassing the current global speed record for large model APIs.
This challenges the industry's long-held belief that "high-performance models come with high latency" or "fast models are necessarily lightweight." The GLM-5.1 Highspeed version is the first domestic large model to deliver premium model capabilities and ultra-low latency together in production, freeing users from having to trade model quality for speed.

Transforming the User Experience for Speed-Critical Scenarios
In long-range tasks and complex production environments, the speed boost has fundamentally changed how products are designed:
AI Programming (Coding Agent): Leveraging the powerful capabilities of GLM-5.1, the new model delivers instant answers to queries. It understands the engineering context while continuously generating code and refining solutions. In reconstruction projects that require dozens of calls, it eliminates the cumulative waiting time of several minutes.
Real-Time Dynamic Modeling: In 3D map field tests, players control a character's movement and input text, and the model instantly completes modeling, dynamically updating the scene in real time.
Agent Swarm Parallel Scheduling: In long-range tasks, the model processes complex web pages within 30 seconds and instantly schedules 50 different personalities to answer in parallel, showcasing the prototype of a new operating system.
Core Technology Deep Dive: TileRT High-Performance Inference Engine
The stable production-level throughput of 400 TPS results from system-level optimization carried out by the Zhipu GLM Team and TileRT Team:
Inference Engine Layer (TileRT Compilation Period AOT Static Scheduling):
Traditional mainstream frameworks use operators (operator/kernel) as the basic scheduling unit. In single-token and small-batch scenarios, this amplifies scheduling, memory access, and synchronization overhead. TileRT completely eliminates dynamic scheduling at the runtime layer, instead statically scheduling the entire computation graph into a persistent GPU persistent Engine Kernel during compilation (AOT). Within a single GPU, computing, asynchronous I/O, and communication are decomposed into tile-level micro-tasks. The entire inference process launches only one kernel, with intermediate results transmitted directly through registers, shared memory, and L2 cache, without writing back to global memory.
Scheduling System Layer:
Through dynamic batching, request merging, and KV cache scheduling optimization, tail latency in high-concurrency scenarios is significantly reduced.
Infrastructure Layer:
At multi-card scale, TileRT extends the concept of Warp Specialization within a single SM to the entire 8-card NVL topology. Different GPU ranks are specialized into different workers based on computational density and data dependencies, combined with network links and load balancing for collaborative optimization, ensuring high-performance and stable operation.
Availability and Access
The GLM-5.1 Highspeed version is ideal for AI programming, real-time interaction, business decision-making, and real-time voice applications that demand extremely low response latency. The service is now officially available on the Zhipu MaaS Platform and is open to select enterprise customers
Related article
Suno to Watermark Songs Amid Legal Battles
Suno, the platform enabling users to generate AI-created music, has unveiled new features to label platform-produced tracks, restrict downloads, and update community standards to curb unauthorized replicas. These updates arrive as Suno confronts mult
Musk Admits Grok Build Leaked User Code, Promises to Erase All Historical Data
Elon Musk directly addressed the privacy controversy surrounding Grok Build, beginning with a simple "True" to confirm the incident's validity. He pledged that all user data previously uploaded to SpaceXAI would be permanently erased, stating, "not a
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Related Special Topic Recommendations
Comments (0)
0/500
Zhipu has officially launched the
This challenges the industry's long-held belief that "high-performance models come with high latency" or "fast models are necessarily lightweight." The GLM-5.1 Highspeed version is the first domestic large model to deliver premium model capabilities and ultra-low latency together in production, freeing users from having to trade model quality for speed.

Transforming the User Experience for Speed-Critical Scenarios
In long-range tasks and complex production environments, the speed boost has fundamentally changed how products are designed:
AI Programming (Coding Agent): Leveraging the powerful capabilities of GLM-5.1, the new model delivers instant answers to queries. It understands the engineering context while continuously generating code and refining solutions. In reconstruction projects that require dozens of calls, it eliminates the cumulative waiting time of several minutes.
Real-Time Dynamic Modeling: In 3D map field tests, players control a character's movement and input text, and the model instantly completes modeling, dynamically updating the scene in real time.
Agent Swarm Parallel Scheduling: In long-range tasks, the model processes complex web pages within 30 seconds and instantly schedules 50 different personalities to answer in parallel, showcasing the prototype of a new operating system.
Core Technology Deep Dive: TileRT High-Performance Inference Engine
The stable production-level throughput of 400 TPS results from system-level optimization carried out by the Zhipu GLM Team and TileRT Team:
Inference Engine Layer (TileRT Compilation Period AOT Static Scheduling):
Traditional mainstream frameworks use operators (operator/kernel) as the basic scheduling unit. In single-token and small-batch scenarios, this amplifies scheduling, memory access, and synchronization overhead. TileRT completely eliminates dynamic scheduling at the runtime layer, instead statically scheduling the entire computation graph into a persistent GPU persistent Engine Kernel during compilation (AOT). Within a single GPU, computing, asynchronous I/O, and communication are decomposed into tile-level micro-tasks. The entire inference process launches only one kernel, with intermediate results transmitted directly through registers, shared memory, and L2 cache, without writing back to global memory.
Scheduling System Layer:
Through dynamic batching, request merging, and KV cache scheduling optimization, tail latency in high-concurrency scenarios is significantly reduced.
Infrastructure Layer:
At multi-card scale, TileRT extends the concept of Warp Specialization within a single SM to the entire 8-card NVL topology. Different GPU ranks are specialized into different workers based on computational density and data dependencies, combined with network links and load balancing for collaborative optimization, ensuring high-performance and stable operation.
Availability and Access
The GLM-5.1 Highspeed version is ideal for AI programming, real-time interaction, business decision-making, and real-time voice applications that demand extremely low response latency. The service is now officially available on the
Suno to Watermark Songs Amid Legal Battles
Suno, the platform enabling users to generate AI-created music, has unveiled new features to label platform-produced tracks, restrict downloads, and update community standards to curb unauthorized replicas. These updates arrive as Suno confronts mult
Musk Admits Grok Build Leaked User Code, Promises to Erase All Historical Data
Elon Musk directly addressed the privacy controversy surrounding Grok Build, beginning with a simple "True" to confirm the incident's validity. He pledged that all user data previously uploaded to SpaceXAI would be permanently erased, stating, "not a
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation





Home






