Zhipu Shatters Global AI Speed Record with GLM-5.1 High-Speed Version
China's renowned AI team Zhipu officially announced today the launch of a new GLM-5.1 Highspeed API for select enterprise clients. Codenamed "GLM-5.1-highspeed," this model has stunned the industry since its release, delivering an output speed of 400 tokens per second.
This figure directly surpasses the current global API speed benchmark set by major large model providers, showcasing strong technical dominance. Previously, the AI industry assumed that model speed and size were mutually exclusive, with higher speeds typically requiring compromises in model capabilities.
Breaking Industry Norms with Flagship Performance
The GLM-5.1 Highspeed version, however, shattered the long-held industry belief that "fast means small." For the first time in domestic large models, it successfully brings flagship-level technical capabilities and ultra-low latency into real-world production environments.
According to reports, this model was developed jointly by Zhipu's GLM team and the TileRT team. Both teams conducted deep, comprehensive system-level optimizations across three layers: the inference engine, the scheduling system, and the underlying infrastructure, moving away from traditional dynamic scheduling.
Three-Level Optimization Ensures Stable Output
On the technical front, the development team not only rewrote the core inference path of the model architecture to boost single-card throughput, but also reduced latency in high-concurrency scenarios through techniques such as dynamic batching. Meanwhile, collaborative optimizations around the infrastructure ensured that 400 TPS became a stable and usable production-level capability.
This high-speed model has broad application prospects, particularly suited for scenarios with strict response latency requirements. Whether for AI programming, real-time voice interaction, or frequent business decisions, the model is now available on Zhipu's MaaS platform for select enterprises.
Related article
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control
According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
How to fix Core Web Vitals for better SEO rankings
Streamline Report Card Comments with AI ToolsIntroductionAI Tools for Generating Report Card CommentsMagic SchoolAlmanac AIChat GPTUsing Magic School to Generate Report Card CommentsLogging into Magic SchoolSelecting the Report Card Comments ToolCust
Related Special Topic Recommendations
Comments (0)
0/500
China's renowned AI team Zhipu officially announced today the launch of a new GLM-5.1 Highspeed API for select enterprise clients. Codenamed "GLM-5.1-highspeed," this model has stunned the industry since its release, delivering an output speed of 400 tokens per second.
This figure directly surpasses the current global API speed benchmark set by major large model providers, showcasing strong technical dominance. Previously, the AI industry assumed that model speed and size were mutually exclusive, with higher speeds typically requiring compromises in model capabilities.
Breaking Industry Norms with Flagship Performance
The GLM-5.1 Highspeed version, however, shattered the long-held industry belief that "fast means small." For the first time in domestic large models, it successfully brings flagship-level technical capabilities and ultra-low latency into real-world production environments.
According to reports, this model was developed jointly by Zhipu's GLM team and the TileRT team. Both teams conducted deep, comprehensive system-level optimizations across three layers: the inference engine, the scheduling system, and the underlying infrastructure, moving away from traditional dynamic scheduling.
Three-Level Optimization Ensures Stable Output
On the technical front, the development team not only rewrote the core inference path of the model architecture to boost single-card throughput, but also reduced latency in high-concurrency scenarios through techniques such as dynamic batching. Meanwhile, collaborative optimizations around the infrastructure ensured that 400 TPS became a stable and usable production-level capability.
This high-speed model has broad application prospects, particularly suited for scenarios with strict response latency requirements. Whether for AI programming, real-time voice interaction, or frequent business decisions, the model is now available on Zhipu's MaaS platform for select enterprises.
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
How to fix Core Web Vitals for better SEO rankings
Streamline Report Card Comments with AI ToolsIntroductionAI Tools for Generating Report Card CommentsMagic SchoolAlmanac AIChat GPTUsing Magic School to Generate Report Card CommentsLogging into Magic SchoolSelecting the Report Card Comments ToolCust





Home






