OpenAI, NVIDIA Forge MRC Protocol to Transform AI Training Networks
OpenAI has officially announced a collaboration with five industry leaders—AMD, Broadcom, Intel, Microsoft, and NVIDIA—to launch the Multipath Reliable Connection (MRC) protocol. This open-source protocol, released through the Open Compute Project (OCP), is designed to tackle the network latency and failures commonly encountered in large-scale AI training.

Eliminating the "Single Point of Failure": From Three-Tier to Two-Tier Architecture
In traditional AI model training, network congestion or a minor failure on a single link can cascade like dominoes, forcing tens of thousands of GPUs into idle states and leading to significant computational waste.
To fundamentally improve system resilience, the MRC protocol introduces a multi-plane network design. It intelligently divides a single 800Gb/s interface into multiple smaller links. This structural optimization enables the system to support massive clusters of up to approximately 131,000 GPUs using just two switch layers. Compared to traditional two- or four-tier architectures, this shift not only drastically reduces the number of physical components and energy consumption but also significantly cuts construction costs.
Advanced Traffic Management: Packet "Spraying" and Microsecond-Level Recovery
Beyond architectural simplification, MRC introduces a novel approach to traffic distribution. It employs adaptive packet spraying technology, moving away from traditional single-path transmission. This method breaks down task packets and distributes them across hundreds of parallel paths. Even if packets arrive out of order, the receiver can accurately reassemble them, effectively preventing localized congestion in the core network.
For network control, MRC replaces complex dynamic routing protocols (like BGP) with SRv6 source routing technology. This allows the sender to directly specify the path, while switches perform only simple static forwarding. This design slashes network fault recovery time from seconds to microseconds, enabling the system to achieve near "seamless self-healing" in the face of link instability.
Real-World Validation: The Supercomputer "Stabilizer"
The MRC protocol is already deployed in NVIDIA's GB200 supercomputer and Oracle's cloud infrastructure. Test data confirms that even during active training scenarios, MRC can automatically reroute around disruptions—such as sudden link jitter or switch reboots—ensuring complex training tasks continue without interruption.
Related article
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control
According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
Related Special Topic Recommendations
Comments (0)
0/500
OpenAI has officially announced a collaboration with five industry leaders—AMD, Broadcom, Intel, Microsoft, and NVIDIA—to launch the Multipath Reliable Connection (MRC) protocol. This open-source protocol, released through the Open Compute Project (OCP), is designed to tackle the network latency and failures commonly encountered in large-scale AI training.

Eliminating the "Single Point of Failure": From Three-Tier to Two-Tier Architecture
In traditional AI model training, network congestion or a minor failure on a single link can cascade like dominoes, forcing tens of thousands of GPUs into idle states and leading to significant computational waste.
To fundamentally improve system resilience, the MRC protocol introduces a multi-plane network design. It intelligently divides a single 800Gb/s interface into multiple smaller links. This structural optimization enables the system to support massive clusters of up to approximately 131,000 GPUs using just two switch layers. Compared to traditional two- or four-tier architectures, this shift not only drastically reduces the number of physical components and energy consumption but also significantly cuts construction costs.
Advanced Traffic Management: Packet "Spraying" and Microsecond-Level Recovery
Beyond architectural simplification, MRC introduces a novel approach to traffic distribution. It employs adaptive packet spraying technology, moving away from traditional single-path transmission. This method breaks down task packets and distributes them across hundreds of parallel paths. Even if packets arrive out of order, the receiver can accurately reassemble them, effectively preventing localized congestion in the core network.
For network control, MRC replaces complex dynamic routing protocols (like BGP) with SRv6 source routing technology. This allows the sender to directly specify the path, while switches perform only simple static forwarding. This design slashes network fault recovery time from seconds to microseconds, enabling the system to achieve near "seamless self-healing" in the face of link instability.
Real-World Validation: The Supercomputer "Stabilizer"
The MRC protocol is already deployed in NVIDIA's GB200 supercomputer and Oracle's cloud infrastructure. Test data confirms that even during active training scenarios, MRC can automatically reroute around disruptions—such as sudden link jitter or switch reboots—ensuring complex training tasks continue without interruption.
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur





Home






