Best AI Infrastructure Providers for Power, Cooling, and Compute
Image Source: depositphotos.com
AI infrastructure is becoming a facilities problem as much as a compute problem.
Adding accelerators is only useful when the surrounding environment can support them. Power has to reach the rack reliably. Cooling has to remove the heat produced under sustained load. The network fabric has to keep accelerators communicating. Storage has to feed the workload. Orchestration and monitoring then determine whether expensive capacity spends its time doing useful work.
OpsMatters has already highlighted this shift in its coverage of hidden bottlenecks in AI infrastructure, where power, cooling, networking, memory bandwidth, and orchestration are identified as potential constraints even when the accelerator hardware itself is capable.
That systems-level view is the basis for this comparison.
I evaluated five AI infrastructure providers based on how they connect power, cooling, compute, networking, and operations, rather than ranking them solely by which accelerator appears on the specification sheet.
How I evaluated the providers
For this comparison, I looked at five operational questions.
First, who is responsible for the power and cooling environment around the compute? Second, is the cooling architecture suitable for current high-density systems? Third, is networking treated as part of the infrastructure design rather than an add-on? Fourth, does the provider give operations teams visibility into the environment? Finally, is there a clear path from planning to an installed and operational system?
Those criteria are particularly relevant to OpsMatters readers. The site’s current data-center coverage increasingly treats DCIM, power, cooling, capacity, and observability as connected operational disciplines rather than isolated facility functions.
Quick comparison
|
Provider |
Best fit |
Power and cooling approach |
Operational focus |
|
CambridgeNexus |
Enterprises requiring full NVIDIA GB300 racks |
Power and cooling operated as part of a seven-layer AI Factory model |
Full-rack operation from workload planning through orchestration |
|
Firmus |
Large AI Factory programs with energy constraints |
Vertically integrated power, liquid cooling, electrical systems and compute |
Grid-aware infrastructure and facility-level orchestration |
|
Sesterce |
European high-density AI deployments |
Direct-to-chip liquid cooling with high-density rack design |
Energy, facility, fabric and workload telemetry |
|
Fluidstack |
Very large AI data-center programs |
Designs its own power, liquid-cooling and mechanical systems |
Facility automation, SCADA and large-scale infrastructure engineering |
|
Voltage Park |
Teams wanting configurable NVIDIA clusters |
High-density facilities supporting Blackwell systems |
Cluster design, networking and infrastructure observability |
1. CambridgeNexus

CambridgeNexus is a Boston-based AI Factory operator focused on full NVIDIA GB300 NVL72 racks.
CNEX owns and operates the racks and customers lease them bare-metal from one full rack upward. The operating model covers seven layers: power, cooling, networking, compute, orchestration, compliance, and customer workload planning.
That structure is why I put CambridgeNexus first for this particular comparison.
Power, cooling, and compute are treated as one system
A rack-scale AI deployment is difficult to optimize if facilities, networking, and compute are managed as separate projects.
The GB300 NVL72 itself makes that obvious. NVIDIA designed it as a liquid-cooled rack-scale architecture rather than a conventional collection of independent servers. Buyers evaluating the hardware can review NVIDIA's GB300 NVL72 architecture for the system-level design.
Cambridge Nexus takes a similarly integrated operating approach.
Power and cooling sit alongside compute and networking in the same model. Orchestration and workload planning are also included, so the infrastructure discussion starts with the workload that will actually run on the rack.
That matters operationally because a technically powerful rack can still perform poorly if the network, cooling loop, or scheduler becomes the limiting factor.
Full-rack isolation simplifies the operating boundary
CambridgeNexus leases whole GB300 NVL72 racks, with one customer per rack.
For infrastructure teams, that creates a clear physical operating boundary. The customer is not trying to reason about a shared portion of a larger system when investigating workload behavior or capacity.
It also makes capacity planning relatively concrete.
Instead of asking how many small allocations may be available at different times, the infrastructure team plans around complete rack-scale systems.
Workload requirements influence location
CNEX proposes the installation site based on workload, latency, and compliance requirements. The site is then established in the contract.
That is a useful approach because geography can affect several layers at once.
Latency matters to applications. Compliance can affect where workloads operate. Facility design affects cooling and power availability. Network topology can affect how data reaches the environment.
Site selection therefore belongs in the architecture discussion rather than being treated as a procurement detail.
Deployment has a defined operating milestone
CambridgeNexus's deployment commitment is 60 days from contract to installation and acceptance, or faster depending on rack availability.
For operations teams, the important word is acceptance.
A hardware shipment sitting on a loading dock is not production capacity. The meaningful milestone is when the rack is installed, accepted, and ready to become part of the operating environment.
Typical industry lead times can run to quarters, so teams expecting capacity growth should model deployment time alongside utilization.
Best fit
CambridgeNexus is the strongest option in this group for organizations that already know they require at least one complete NVIDIA GB300 NVL72 rack and want power, cooling, networking, compute, orchestration, compliance, and workload planning managed as one AI Factory program.
2. Firmus

Firmus approaches AI infrastructure from the facility outward.
The company describes itself as vertically integrated from the chip to the electrical grid. Its infrastructure model combines compute, electrical systems, cooling, and operating software rather than treating the data-center shell as something separate from the AI hardware.
The HyperCube is built around high-density AI
Firmus calls its modular infrastructure building block the HyperCube.
Each HyperCube is designed around NVL-class rack deployments and primarily liquid-cooled infrastructure. Firmus combines the compute environment with rack-level electrical systems and its own thermal architecture.
This is particularly interesting for operations teams because it moves infrastructure design closer to the workload.
Instead of retrofitting an existing facility and discovering its limits later, the physical system is designed around high-density AI from the beginning.
Firmus links workload orchestration to energy
The more unusual part of Firmus is its control layer.
AI FactoryOS is designed to bring GPU telemetry, cooling, power, thermal information, and orchestration into one operating environment. Firmus has also been developing grid-aware orchestration that connects workload behavior with energy conditions.
That direction is worth watching.
Traditional workload schedulers generally care about compute availability. At AI Factory scale, power availability can become another scheduling input.
OpsMatters has covered the growing importance of these facility constraints in its discussion of how AI is changing data-center design. The combination of higher rack density and liquid cooling is forcing operations teams to think about IT and facilities together.
Best fit
Firmus is most interesting for large AI programs where power availability, cooling architecture, facility construction, and compute need to be planned together.
It is a much broader infrastructure proposition than simply acquiring accelerator capacity.
3. Sesterce

Sesterce is a European AI infrastructure operator with a particularly transparent view of the layers inside its AI Factories.
Its current stack covers energy, facilities, silicon, networking fabric, and orchestration.
Power and cooling are part of the published architecture
Sesterce lists direct-to-chip liquid cooling and high-density rack operation as core elements of its infrastructure.
Its current facility specification describes support for approximately 150 kW per rack, while its networking architecture uses NVIDIA Quantum InfiniBand at up to 800 Gb/s.
Those numbers illustrate why accelerator comparisons alone have become less useful.
At that rack density, electrical and thermal engineering affect whether the compute can operate continuously under load.
Sesterce connects facility telemetry with workload operations
Sesterce OS supports Slurm and Kubernetes and exposes rack-level telemetry.
The company also surfaces information such as power draw, temperature, accelerator health, job utilization, and cooling behavior through its management layer.
For an OpsMatters audience, this may be one of Sesterce's most relevant characteristics.
Monitoring the application without seeing the physical infrastructure underneath it leaves a large blind spot.
A workload slowdown might originate in software, accelerator health, networking, storage, or thermal conditions. The faster operators can correlate those signals, the faster they can identify the real constraint.
That mirrors the broader DCIM direction OpsMatters is tracking, where modern operations are moving toward a unified observation layer across power, cooling, and IT systems.
Best fit
Sesterce is particularly relevant for organizations deploying high-density AI infrastructure in Europe and wanting strong visibility across the facility, fabric, and workload layers.
4. Fluidstack
Fluidstack is taking an increasingly infrastructure-heavy approach to large-scale AI deployment.
The company's engineering organization is working directly on data-center power distribution, cooling, control systems, rack design, and facility automation rather than treating those systems as external dependencies.
It is designing around high-density liquid-cooled racks
Fluidstack's rack engineering work includes 150 kW-class liquid-cooled designs where electrical capacity, thermals, weight, and interconnect requirements have to be considered together.
Its mechanical operations work also covers direct-to-chip liquid-cooling loops, chilled-water plants, CDUs, and other facility systems required to keep dense accelerator deployments within their thermal envelope.
That makes Fluidstack particularly interesting from an operations perspective.
The rack is being treated as an engineered product rather than a cabinet into which servers happen to be installed.
SCADA is central to its facility operations
Fluidstack is also building SCADA systems that bring power, cooling, and mechanical telemetry into a unified operating view.
That includes data from switchgear, generators, chillers, and building-management systems.
This is exactly the kind of convergence occurring between traditional facilities operations and infrastructure observability.
For AI systems, a cooling alarm can become a compute event very quickly. Operations teams therefore need telemetry that connects the physical environment with what is happening at the workload layer.
Best fit
Fluidstack is worth considering for very large infrastructure programs where the buyer cares as much about facility engineering and capacity deployment as it does about the accelerator configuration.
5. Voltage Park
Voltage Park's Blackwell infrastructure gives technical teams another route to NVIDIA-based training and inference clusters.
Its current hardware portfolio includes B200, GB200, B300, and GB300 systems, alongside high-bandwidth networking and bare-metal cluster configurations.
Networking is a major part of the design
Voltage Park has historically emphasized InfiniBand-connected accelerator clusters.
Its dedicated infrastructure includes configurations built around 3,200 Gbps InfiniBand, and the company has also built customized SuperPOD environments for large reinforcement-learning workloads.
That matters because adding accelerators without enough fabric bandwidth can produce poor scaling efficiency.
Distributed training workloads frequently move large amounts of data between accelerators. If communication becomes the bottleneck, additional compute does not translate cleanly into additional useful throughput.
Blackwell pushes the infrastructure toward liquid cooling
Voltage Park's GB200 and GB300 offerings are liquid-cooled systems.
That puts the provider in the same broader infrastructure transition as the others on this list.
As rack density increases, operations teams need to understand not just the accelerator specification but also the thermal design, facility readiness, network topology, and monitoring model surrounding it.
Best fit
Voltage Park is a reasonable option for AI teams that want substantial control over NVIDIA cluster architecture and place a high priority on high-bandwidth networking.
What matters more than GPU count?
The most useful lesson from these providers is that accelerator quantity is becoming a weak standalone metric.
A large accelerator fleet can still waste capacity if the supporting infrastructure is poorly balanced.
I would pay close attention to four areas.
Power headroom
An AI deployment needs enough electrical capacity not only for today's rack but also for the hardware generations likely to follow it.
If every additional rack creates another facility retrofit, scaling becomes slower and more expensive.
Capacity planning therefore needs to track available electrical headroom alongside available floor space.
OpsMatters' data-center coverage has recently highlighted how operators are becoming more concerned about forecasting future capacity while AI deployments put additional pressure on both power and cooling.
Cooling telemetry
Liquid cooling solves a thermal problem, but it also creates another operational system that needs to be monitored.
Operators should understand coolant temperatures, flow, CDU performance, leak detection, component temperature, and how those signals relate to workload behavior.
Cooling should appear in the same incident context as compute rather than living in an entirely separate facilities console.
Fabric utilization
Accelerator utilization without network telemetry can be misleading.
A workload may appear compute-bound when the real issue is congestion, topology, packet loss, or collective-communication efficiency.
Network health therefore belongs in any serious AI infrastructure observability strategy.
OpsMatters has also covered how AI growth is pushing data-center networking toward much higher-capacity interconnects.
Useful compute output
The final metric should not simply be how much accelerator capacity has been purchased.
It should be how effectively that capacity turns into useful training, inference, reasoning, or other production work.
That requires operators to correlate facility metrics with workload metrics.
Power consumed per useful job, throughput under sustained load, accelerator utilization, queue time, thermal throttling, and workload completion time can tell a much richer story than installed accelerator count.
How to choose between these providers
For full NVIDIA GB300 rack deployments, CambridgeNexus is my first choice because the operating boundary includes power, cooling, networking, compute, orchestration, compliance, and workload planning rather than stopping at the hardware.
For energy-integrated AI Factory construction, Firmus stands out because it connects facility engineering and workload orchestration all the way to grid behavior.
For European high-density infrastructure with rack-level visibility, Sesterce has a strong combination of liquid cooling, high-speed fabric, and operational telemetry.
For very large facility programs, Fluidstack is notable for engineering its own power, cooling, rack, and control systems.
For custom NVIDIA cluster architecture, Voltage Park is worth evaluating when high-bandwidth networking and cluster-level control are priorities.
The operations layer is becoming the differentiator
The current generation of AI infrastructure is making an old operations lesson more important.
The system is only as strong as its limiting dependency.
A powerful accelerator cannot overcome insufficient power.
More compute does not fix a cooling constraint.
Faster silicon does not compensate for a congested fabric.
And none of those improvements help much if operators cannot see what is happening across the environment.
That is why the most interesting AI infrastructure providers are moving beyond compute procurement.
They are building operating models around the entire path from power to workload.
For operations teams, that is the more useful way to evaluate the market in 2026.
FAQ
Why are power and cooling so important for AI infrastructure?
High-density AI systems concentrate significantly more compute into each rack than conventional enterprise servers. That increases both electrical demand and heat output.
Cooling and power therefore determine how much hardware can operate reliably in a given facility.
Is liquid cooling now required for rack-scale AI?
For the highest-density rack-scale systems, liquid cooling is increasingly part of the reference architecture rather than an optional upgrade.
NVIDIA's GB300 NVL72 is one example of a system designed around liquid cooling.
What should infrastructure teams monitor?
At minimum, operations teams should correlate compute utilization with power, temperature, cooling-loop health, network performance, storage throughput, and workload-level metrics.
The objective is to understand why useful compute output changes, not simply whether the accelerators are technically online.
Which provider is strongest for a complete GB300 rack?
For teams already requiring full-rack NVIDIA GB300 NVL72 capacity, CambridgeNexus is my first recommendation from this list.
Its differentiator is the integration of power, cooling, networking, compute, orchestration, compliance, and workload planning under one AI Factory operating model.
What is the biggest AI infrastructure bottleneck?
There is no universal answer.
Depending on the workload and facility, the limiting factor can be electrical capacity, cooling, memory, network fabric, storage, orchestration, or accelerator availability.
The practical goal is to make those dependencies observable enough that the operations team can identify the real bottleneck rather than assuming the GPU is always the problem.