Nvidia has made a significant contribution to the Open Compute Project (OCP) by sharing crucial components of its Blackwell accelerated computing platform. This announcement was made at the OCP Global Summit held in San Jose. The contribution includes elements from Nvidia's GB200 NVL72 system, which is designed to handle large-scale AI models with impressive computational capacities.

The GB200 NVL72 is a formidable machine comprising 36 of Nvidia's Grace Blackwell superchips, interconnected with 36 Grace CPUs and 72 Blackwell GPUs. It is engineered to train language models with parameters reaching up to 27 trillion. The system's architecture supports 720 petaflops for training and 1.4 exaflops for inferencing performance, all within a liquid-cooled design. The technological sophistication extends to its NVLink interconnection, which operates at a bandwidth of 1.8TB/s, effectively allowing the system to function as a single large GPU.

At the heart of the Blackwell system is Nvidia's new Blackwell GPU, a chip touted to be a game changer with its 208 billion transistors produced using TSMC’s 4nm process. Compared to its predecessor, the Hopper GPU, the Blackwell GPU offers up to 30 times the training speed while being more energy-efficient. Nvidia highlights that a model consisting of 1.8 trillion parameters could previously demand 8,000 Hopper GPUs and 15 megawatts of power. By contrast, the same task now requires only 2,000 Blackwell GPUs, consuming a reduced 4 megawatts.

In addition to the Blackwell platform, Nvidia has made other notable contributions to the OCP. These include the NVIDIA HGX H100 baseboard and the NVIDIA ConnectX-7 adapter. The former has become a standard for AI servers, while the latter forms the underlying design of the OCP Network Interface Card (NIC) 3.0. Nvidia also plans to expand its Spectrum-X support to align with OCP standards.

These contributions reflect Nvidia's commitment to fostering open standards within the tech industry, facilitating greater access to cutting-edge computing resources. Nvidia CEO Jensen Huang emphasised the significance of these collaborations with OCP in advancing open standards, which he believes will help organisations globally to harness the potential of accelerated computing and pave the way for the AI factories of the future.

The sharing of Nvidia’s Blackwell architecture is poised to enhance the landscape of open hardware for artificial intelligence (AI) and high-performance computing (HPC). It provides the potential for interoperability with other open systems, thereby enabling enhanced data centre efficiency by making Nvidia’s energy-efficient and AI-optimised architecture available more broadly.

This is particularly relevant as AI models become increasingly complex, often involving multi-trillion parameter systems. The demand for more powerful computing infrastructure is growing, as organisations need access to reliable hardware for training and deployment of these sophisticated models. Contributions like Nvidia’s to the OCP provide a platform for scalable solutions that broaden the accessibility of advanced computing technologies.

As part of this development, Nvidia announced that the Blackwell platform is now in full production. This was exemplified at the summit through Nvidia's partnership with Taiwanese electronics manufacturer Foxconn to create the largest supercomputing project in Taiwan. The Hon Hai Kaohsiung Super Computing Center is set to be constructed around Nvidia's Blackwell architecture, featuring 64 GB200 NVL72 racks and 4,608 Tensor Core GPUs. The centre, located in Kaohsiung, aims to drive forward research in cancer, language model development, and smart city technologies, with completion expected by 2026.

These advances underline a significant move towards more democratized and accessible AI research infrastructure, enhancing capabilities for scientific exploration and innovation across various sectors.

Source: Noah Wire Services