Tag: GPU supercluster

  • WhiteFiber’s 83 km, 136 Tbps Link Bets Fiber Can Pool AI Power Across Sites

    WhiteFiber’s 83 km, 136 Tbps Link Bets Fiber Can Pool AI Power Across Sites

    TL;DR · 30-second read

    The Short Version

    WhiteFiber, a company that rents out computing power for artificial intelligence, is now selling a way to make two data centers about 50 miles apart work as one giant computer. The two buildings are joined by a bundle of fiber-optic cables.

    Why it matters: the computers that train artificial intelligence are often limited by how much electricity a single building can get. WhiteFiber’s idea is to combine spare electricity at several existing buildings instead of building one enormous new site.

    The top speed is still being tested, and no customers have been named yet.

    StorageReview reported on September 25 that WhiteFiber (WYFI) has made WhiteFiber Continuum commercially available. Continuum is the productized version of Project Redwood, a two-site GPU supercluster the company first detailed in July. It links two HITRUST-certified QTS data centers 83 kilometers apart over 12 Zayo dark fiber strands, with DriveNets supplying the Ethernet network fabric and WEKA’s NeuralMesh supplying shared storage and memory.

    WhiteFiber rates the architecture at 136 Tbps of aggregate bandwidth, which it describes as 170 wavelength channels of 800 gigabits per second each, with 0.9 milliseconds of guaranteed round-trip latency. The July research result measured 111.2 Tbps on part of the fiber spectrum. Full-spectrum testing to confirm the guaranteed commercial specifications is underway, and the company has submitted patent applications for the implementation.

    Executive Summary

    WhiteFiber is selling a way to make GPUs in two separate buildings behave like a single cluster. GPUs are the chips that train and run AI models, and large training jobs have normally been confined to one building because the chips have to exchange data constantly. Continuum stretches that boundary to 83 kilometers by carrying the traffic over leased, unlit fiber that WhiteFiber lights itself, which the industry calls dark fiber.

    The launch matters because of what WhiteFiber says the design is for. The company pitches Continuum as a way to get past the power and space ceiling of a single campus. Enterprises would pool GPUs across sites, and telecom and metro facilities with spare power could contribute capacity to one logical cluster. If that works at the stated performance, the useful unit of AI capacity is no longer one building with enough electricity. It is any set of buildings connected by enough fiber.

    The headline number still needs confirmation. The 136 Tbps figure is a design rating that depends on additional wavelengths coming online and on full-spectrum testing now in progress. The figure WhiteFiber has actually measured is 111.2 Tbps, taken in July on part of the spectrum.

    Fiber as a Way Around the Single-Site Power Ceiling

    AI capacity is increasingly set by power: how many megawatts a single building or campus can draw. A megawatt is a million watts, and a large GPU campus uses many of them. A company that has filled its power allocation at one site usually faces a choice between waiting for more grid capacity and building a new campus from scratch. WhiteFiber offers a third option. Continuum lets enterprises pool GPUs across sites and add overflow training and inference capacity without a separate greenfield deployment. It also lets telecom and metro facilities with spare power and fiber join the same logical cluster.

    The design choice that makes this plausible is what crosses the link. According to WhiteFiber’s product page, both sites carry live traffic at the same time and only model gradients cross the 83 km connection. Gradients are the incremental updates each group of GPUs computes during training. The bulk of the work and data stays local to each site, and the two halves synchronize their learning over the fiber. A single job can span both sites, or the cluster can be split into independent workloads that burst from one site to the other as demand shifts.

    Fiber does not create power. It lets power that already exists in two places be used by one workload. The groups most affected are operators whose sites have stranded capacity, such as telco central offices and metro facilities with power to spare but too little of it to host a full cluster. Fiber owners like Zayo, whose dark fiber becomes part of the compute fabric, stand to gain as well. Enterprises that have maxed out a campus are the third group. Whether this changes where AI capacity gets built will depend on how much of a real training job’s performance survives the split, and WhiteFiber has not yet published that comparison.

    What 136 Tbps Establishes, and What It Does Not Yet

    The internal arithmetic holds together. 170 wavelength channels at 800 gigabits per second each come to 136 terabits per second. Dense wavelength-division multiplexing sends many separate colors of light down the same fiber strands, each carrying its own data stream. Adding wavelengths is how the design scales from the 111.2 Tbps measured in July to the 136 Tbps now rated.

    The status of the number matters to buyers. WhiteFiber says the design reaches 136 Tbps after additional wavelengths come online, and that full-spectrum testing is underway to confirm the guaranteed commercial specifications. Enterprises evaluating Continuum should treat 111.2 Tbps as the demonstrated figure and 136 Tbps as the contracted target until test results are released.

    The latency figure tells a different story. WhiteFiber guarantees 0.9 milliseconds round trip, which it says is within 8% of the physical limit for light in fiber over 83 km. There is little left to optimize, so distance itself is the fixed cost of the architecture. Longer links between sites add latency in proportion to distance. The 83 km result shows the approach works at metro-to-regional range, and it says nothing yet about sites much farther apart.

    Resilience and Data Boundaries as the Enterprise Pitch

    Continuum is also sold as a resilience and compliance product. WhiteFiber says neither site is a single point of failure, so training runs are designed to continue if one location goes down. For regulated industries, it says sensitive workloads can stay within their originating jurisdiction while failover and pooled compute happen across locations. HITRUST is a security and privacy certification framework widely used in healthcare, and the certification of both QTS facilities is aimed at that audience.

    CEO Sam Tabar framed the product around both themes, saying it is for enterprises that need AI compute to “hold up under compliance requirements, and not break when a single site has a problem.” These are credible goals for a two-site design. They are also claims about behavior under failure, and no failure test results have been released.

    A Partner Stack With a Patent Claim on Top

    Continuum is assembled from named partners. QTS provides the facilities, Zayo the fiber, DriveNets the Ethernet-based network fabric, and WEKA the shared storage and memory layer. WEKA CEO Liran Zvibel said NeuralMesh lets GPUs that sit many kilometers apart “operate as if the data were local.” DriveNets said its fabric is built to maximize GPU utilization across scale-across superclusters. Using Ethernet rather than proprietary interconnects for the inter-site fabric keeps the design within a widely supported networking standard.

    WhiteFiber calls Continuum the first commercially available distributed GPU supercluster architecture and has submitted patent applications for the implementation. Submitted applications are not granted patents. Because most of the stack comes from third-party vendors, WhiteFiber’s lasting advantage will likely come from its integration and operating know-how, and from customer contracts once it signs them.

    Background

    WhiteFiber, which trades under the ticker WYFI, provides AI infrastructure. In July it detailed Project Redwood, a research effort that connected GPUs in two data centers 83 km apart and measured 111.2 Tbps across the link. Its partners are well-established infrastructure suppliers. QTS operates data centers, Zayo owns and leases fiber networks, DriveNets builds Ethernet networking for AI clusters, and WEKA makes high-performance data storage software.

    Large AI models have traditionally been trained inside a single building. The GPUs doing the work exchange data constantly, and longer distances add delay. As single campuses run up against limits on how much power and space they can secure, operators have looked for ways to spread one workload across several sites. Continuum is WhiteFiber’s commercial answer to that problem.

    Sources

    Source: WhiteFiber Continuum Goes Commercial, Scaling Its 83 km Two-Site GPU Supercluster Design to 136 Tbps (StorageReview.com): WhiteFiber makes its two-site, 83 km GPU supercluster architecture commercially available with a rated 136 Tbps of inter-site bandwidth.