Moonshot AI and Tsinghua University Propose PrfaaS Architecture for Data Centers
As large language models (LLMs) demand more computational power during reasoning, conventional service architectures encounter bottlenecks. Moonshot AI, in collaboration with a research team from Tsinghua University, has introduced a novel architecture called Pre-filling as a Service (PrfaaS), aimed at overcoming the constraints of data centers and computing resources in LLM services.

The reasoning process of large language models is generally split into two phases: pre-filling and decoding. Pre-filling is computationally intensive, where the model processes input and creates a key-value cache (KVCache). Decoding is memory bandwidth-intensive, generating outputs sequentially. Traditional setups require both stages to occur within the same data center, leading to constraints in computation and bandwidth.
PrfaaS enables efficient cross-data-center service by offloading pre-filling tasks to dedicated high-compute clusters, using a standard Ethernet network to transfer the generated KVCache to local decoding clusters. Research indicates that this architecture boosts processing performance, achieving a 54% increase in service throughput over traditional models. Real-world case studies also show lower latency and higher efficiency.
The PrfaaS architecture decouples the three main subsystems—computing, networking, and storage—managing each independently. A precise routing mechanism ensures efficient transmission of long requests, preventing congestion from uneven resource allocation seen in traditional approaches. Additionally, a dual-timescale scheduling mechanism adapts to varying traffic patterns, further optimizing resource usage.
As the demand for cross-data-center reasoning grows and new hardware continues to emerge, PrfaaS offers a promising solution for future AI applications.
Related article
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control
According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
Related Special Topic Recommendations
Comments (10)
0/500
Интересная архитектура PrfaS от Moonshot и Цинхуа для решения проблем с вычислительной мощностью LLM. Надеюсь, это поможет оптимизировать центры обработки данных! 🚀
A arquitetura PrfaS proposta por Moonshot e Tsinghua é um passo interessante para resolver gargalos de computação em LLMs. Será que isso viabiliza data centers mais eficientes? 🤔
La arquitectura PrfaS de Moonshot y Tsinghua parece una gran solución para los cuellos de botella en centros de datos con LLMs. ¡Esperamos ver resultados concretos pronto! 📈
Moonshot AI und Tsinghua zeigen mit PrfaS einen vielversprechenden Ansatz gegen Rechenengpässe bei LLMs. Spannend zu sehen, wie diese Kooperation die Datenzukunft gestaltet! 🌐
As large language models (LLMs) demand more computational power during reasoning, conventional service architectures encounter bottlenecks. Moonshot AI, in collaboration with a research team from Tsinghua University, has introduced a novel architecture called Pre-filling as a Service (PrfaaS), aimed at overcoming the constraints of data centers and computing resources in LLM services.

The reasoning process of large language models is generally split into two phases: pre-filling and decoding. Pre-filling is computationally intensive, where the model processes input and creates a key-value cache (KVCache). Decoding is memory bandwidth-intensive, generating outputs sequentially. Traditional setups require both stages to occur within the same data center, leading to constraints in computation and bandwidth.
PrfaaS enables efficient cross-data-center service by offloading pre-filling tasks to dedicated high-compute clusters, using a standard Ethernet network to transfer the generated KVCache to local decoding clusters. Research indicates that this architecture boosts processing performance, achieving a 54% increase in service throughput over traditional models. Real-world case studies also show lower latency and higher efficiency.
The PrfaaS architecture decouples the three main subsystems—computing, networking, and storage—managing each independently. A precise routing mechanism ensures efficient transmission of long requests, preventing congestion from uneven resource allocation seen in traditional approaches. Additionally, a dual-timescale scheduling mechanism adapts to varying traffic patterns, further optimizing resource usage.
As the demand for cross-data-center reasoning grows and new hardware continues to emerge, PrfaaS offers a promising solution for future AI applications.
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Интересная архитектура PrfaS от Moonshot и Цинхуа для решения проблем с вычислительной мощностью LLM. Надеюсь, это поможет оптимизировать центры обработки данных! 🚀
A arquitetura PrfaS proposta por Moonshot e Tsinghua é um passo interessante para resolver gargalos de computação em LLMs. Será que isso viabiliza data centers mais eficientes? 🤔
La arquitectura PrfaS de Moonshot y Tsinghua parece una gran solución para los cuellos de botella en centros de datos con LLMs. ¡Esperamos ver resultados concretos pronto! 📈
Moonshot AI und Tsinghua zeigen mit PrfaS einen vielversprechenden Ansatz gegen Rechenengpässe bei LLMs. Spannend zu sehen, wie diese Kooperation die Datenzukunft gestaltet! 🌐





Home






