Zhipu Unveils GLM-5.3-Flash, Open-Sourcing Native Multimodal Model at 1/40th Price
Zhipu has officially launched and open-sourced GLM-5.3-Flash (320B-A18B), the first native multimodal model in the GLM-5 series, delivering cutting-edge AI performance at an unprecedentedly low cost.
Boasting global-leading capabilities, the model scored 57 on the Artificial Analysis Intelligence Index, matching Claude Opus4.8. It demonstrates superior comprehensive performance and coding proficiency compared to the larger GLM-5.2. Prior to its release, it anonymously topped usage charts on OpenCode and OpenRouter under the name "Ox-Alpha".

Cost efficiency is a key advantage. Priced at just one-tenth of the standard GLM-5.3—and as low as one-twentieth during limited-time promotions—GLM-5.3-Flash costs roughly one-fortieth of Opus4.8, drastically lowering the barrier to accessing state-of-the-art AI.
Architecturally, the model utilizes a hybrid sparse and linear attention mechanism, a first for open-source frontier models. Enhanced by manifold constraint super connections, it reduces attention computation by 3.01x and KV cache usage by 4.44x, optimizing the balance between long-context handling and inference costs. Its native visual integration enables independent execution of front-end development, 3D modeling, and document creation, excelling in professional domains like finance and law.

Notably, the model runs entirely on domestic chip clusters. Leveraging a custom inference engine with EPD separation and advanced memory optimizations, the team achieved a threefold boost in end-to-end inference speed, matching the hardware utilization efficiency of mainstream NVIDIA GPUs and validating domestic computing power for large-scale AI.
GLM-5.3-Flash is fully open-sourced, providing API access, online demos, and integration with Agent platforms like ZCode and AutoClaw. It also offers daily experience cards granting ten thousand uses, available to developers worldwide.
Online Experience
Z.ai: https://chat.z.ai
Zhipu Qingyan App: https://chatglm.cn
Related article
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control
According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
Related Special Topic Recommendations
Comments (0)
0/500
Zhipu has officially launched and open-sourced GLM-5.3-Flash (320B-A18B), the first native multimodal model in the GLM-5 series, delivering cutting-edge AI performance at an unprecedentedly low cost.
Boasting global-leading capabilities, the model scored 57 on the Artificial Analysis Intelligence Index, matching Claude Opus4.8. It demonstrates superior comprehensive performance and coding proficiency compared to the larger GLM-5.2. Prior to its release, it anonymously topped usage charts on OpenCode and OpenRouter under the name "Ox-Alpha".

Cost efficiency is a key advantage. Priced at just one-tenth of the standard GLM-5.3—and as low as one-twentieth during limited-time promotions—GLM-5.3-Flash costs roughly one-fortieth of Opus4.8, drastically lowering the barrier to accessing state-of-the-art AI.
Architecturally, the model utilizes a hybrid sparse and linear attention mechanism, a first for open-source frontier models. Enhanced by manifold constraint super connections, it reduces attention computation by 3.01x and KV cache usage by 4.44x, optimizing the balance between long-context handling and inference costs. Its native visual integration enables independent execution of front-end development, 3D modeling, and document creation, excelling in professional domains like finance and law.

Notably, the model runs entirely on domestic chip clusters. Leveraging a custom inference engine with EPD separation and advanced memory optimizations, the team achieved a threefold boost in end-to-end inference speed, matching the hardware utilization efficiency of mainstream NVIDIA GPUs and validating domestic computing power for large-scale AI.
GLM-5.3-Flash is fully open-sourced, providing API access, online demos, and integration with Agent platforms like ZCode and AutoClaw. It also offers daily experience cards granting ten thousand uses, available to developers worldwide.
Online Experience
Z.ai: https://chat.z.ai
Zhipu Qingyan App: https://chatglm.cn
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur





Home






