The debate over the computing capabilities categorised as 'Zettascale' and 'Exascale-class' AI supercomputers has been brought to the forefront by Doug Eadline, a prominent expert in high-performance computing, in his recent analysis on HPCWire. Eadline scrutinises the prevalent use of these terms, questioning their validity, especially when applied to AI workloads.
Traditionally, the term 'exascale' refers to computers capable of executing at least one quintillion floating-point operations per second (FLOPS). This benchmark has been a standard for assessing the potential of high-performance computing systems, ensuring they achieve 10^18 FLOPS in sustained, double-precision (64-bit) calculations. However, Eadline highlights a growing tendency to misuse or misrepresent these terms, particularly with claims linked to AI supercomputers.
Eadline's investigation reveals that recent assertions of 'exascale' or 'zettascale' performance frequently rely on speculative, rather than empirical, metrics. He articulates this issue with the question, "How do these 'snort your coffee' numbers arise from unbuilt systems?", underscoring the gap between experimental peak outputs and actual, tested results.
In the field of AI, variations in floating-point formats are pivotal. AI workloads typically employ lower-precision formats such as FP16, FP8, or even FP4, as opposed to traditional high-performance computing (HPC) systems, which necessitate higher precision for reliability. This dependence on lower precision contributes to exaggerated claims of achieving exaFLOP or zettaFLOP performance, which Eadline critiques, stating, "Calling it 'AI zetaFLOPS' is silly because no AI was run on this unfinished machine."
The expert underscores the significance of adhering to verified benchmarks such as HPLinpack. Established in 1993, HPLinpack remains the gold standard for measuring HPC performance. He argues that using theoretical peak figures in place of tested metrics can be misleading. Currently, only two supercomputers — Frontier at Oak Ridge National Laboratory and Aurora at Argonne National Laboratory — have been validated as part of the exascale club through rigorous testing with real applications.
To elucidate the distinction between varying floating-point formats, Eadline uses an automotive analogy: "The average double precision FP64 car weighs about 4,000 pounds (1,814 kilogrammes). It is great at navigating terrain, holds four people comfortably, and gets 30 miles per gallon. Now, consider the FP4 car, which has been stripped down to 250 pounds (113 kilogrammes) and gets an astounding 480 miles per gallon." This analogy illustrates the discrepancies in capacity and utility between higher precision FP64, which requires comprehensive features, and lower precision FP4, which is less capable of managing complex terrains—paralleling the demands of high-precision calculations against simpler inference tasks.
Eadline's examination emphasises the ongoing convergence of AI and HPC technologies, while also advocating for a clear distinction in performance measurement standards. By stressing the importance of validated systems that meet the rigorous demands for double-precision calculations, Eadline aims to clarify the benchmark for what constitutes true exascale or zettascale systems in both AI and traditional HPC domains.
Source: Noah Wire Services