[AI infrastructure insight] Why power and cooling have become the next challenge for AI data centers

Источник: SK hynix Newsroom

[AI infrastructure insight] Why power and cooling have become the next challenge for AI data centers

Source: SK hynix Newsroom

AI is no longer defined by a single model or AI processor. For AI to operate effectively in real-world services and industrial applications, it takes faster compute, wider memory bandwidth, higher-performance networks, more efficient storage, and more stable power and

•Updated: September 30, 2026

AI is no longer defined by a single model or AI processor. For AI to operate effectively in real-world services and industrial applications, it takes faster compute, wider memory bandwidth, higher-performance networks, more efficient storage, and more stable power and cooling working seamlessly together. This is why competition in AI is shifting beyond individual technologies to the design and operation of the entire infrastructure.

In the AI Infrastructure Insight series, we explore how the architecture behind AI infrastructure is evolving — from compute, memory, storage, network, and power and cooling to the way they come together as a unified system — and what this means for the industry.

[Series overview] ① What is changing inside AI data centers? ② Why faster GPUs alone can’t deliver AI performance ③ Why power and cooling have become the next challenge for AI data centers ④ How AI infrastructure will be designed for the future

Competition in AI infrastructure is expanding beyond securing greater computing performance to sustaining that performance reliably and operating it efficiently within actual data centers. The AI data center architecture examined in Part 1 and the data movement challenges discussed in Part 2 ultimately lead to two practical issues: power supply and thermal management. In environments where large numbers of systems operate simultaneously, the ability to provide stable power and effectively manage the resulting heat through cooling determines performance and scalability.

As a result, power and cooling are no longer considerations to be addressed after a data center is built, but core requirements that must be factored in from the initial design stage. Stable power delivery and effective thermal management are also essential to fully utilize the performance of faster accelerators, high memory bandwidth, and high-speed networks.

Rising power demand as AI infrastructure expands

As AI training and inference increase, so do the computational workloads that data centers must handle and the power required to support them. As high-performance equipment such as AI accelerators*, memory, network devices, and storage become highly concentrated and operated simultaneously at the rack* and cluster levels, the scale of power that data centers must reliably secure and supply is rapidly rising.

* AI processor: A semiconductor or compute device designed to rapidly process the large-scale computations required for AI training and inference. Examples include GPUs, NPUs, and TPUs. * Rack: A physical infrastructure unit designed to house multiple servers as well as networking, power, and cooling equipment.

This shift is driving an increase in data center electricity consumption worldwide. The International Energy Agency (IEA) estimates that global data centers consumed about 415 TWh* of electricity in 2024, accounting for approximately 1.5% of global electricity consumption. The IEA projects that this figure could more than double to about 945 TWh by 2030. In particular, electricity consumption by accelerated servers*, which are closely associated with the growth of AI, is expected to increase by an average of 30% annually from 2024 to 2030.

* Terawatt-hour (TWh): A unit of electricity consumption. One TWh is equivalent to 1 billion kWh. * Accelerated server: A server equipped with specialized compute devices such as GPUs or AI accelerators to enhance AI training and inference performance.

In the United States, the expansion of data centers is also increasing the importance of securing adequate power infrastructure. McKinsey estimates that data centers could account for about 75% of the projected growth in U.S. electricity demand over the next decade. If the current pace of construction continues, data centers are being built at a rate that requires the equivalent of nearly 30 GW* of power each year*. This figure includes not only the electricity consumed by IT equipment such as servers, storage, and networking systems, but also cooling systems, power distribution equipment, and reserve capacity required for reliable operations.

* GW: One GW is equivalent to 1 million kW. * McKinsey & Company, “Powering AI: How real is the risk of overbuilding?,” July 31, 2026

The scalability of AI infrastructure is therefore directly tied to the operational capabilities of the data centers that support it. Keeping large-scale equipment running continuously requires sufficient power infrastructure and cooling systems as well as operational headroom to accommodate peak demand*. This is why the constraints on AI expansion are extending beyond the performance of individual semiconductors to the operational capabilities of the data center as a whole.

* Peak demand: The highest level of electricity demand during a given period. Data centers must account for sufficient reserve capacity to maintain reliable operations even during peak demand.

In AI data centers, power is consumed not only by accelerators but across the entire system, including memory, storage, and networking equipment. The data movement discussed in Part 2 also contributes to energy demands. As data travels back and forth across multiple layers, and as the volume and frequency of data movement increase, the energy required for data transfers can also rise.

Higher power consumption also means more heat is generated from equipment. Data center cooling can be broadly divided into two stages: removing heat generated by servers and racks, and transferring the recovered heat outside the facility. The former uses methods such as air and liquid cooling, while the latter relies on coolant circulation and heat exchange systems. Both stages require additional energy. Efficiency losses also occur when externally supplied electricity is converted and distributed into forms that servers and other equipment can use. Ultimately, the efficiency of an AI data center depends not only on the characteristics of a single semiconductor, but also on efficiency across the entire system — from computation and data movement to power conversion and cooling.

In particular, as more high-performance AI equipment is deployed, rack-level power density* and heat generation rise, making thermal management increasingly important. Insufficient cooling can cause processors to throttle their operating speeds, reducing performance and affecting equipment reliability, service life, and maintenance costs. As a result, cooling has become a factor that must be considered from the earliest stages of design alongside server, rack, and cluster configurations as well as packaging* and power delivery architecture.

* Power density: The amount of power consumed within a given space. In data centers, it is typically used to describe power consumption at the rack or server-room level, measured in kilowatts (kW). * Packaging: A technology that electrically connects and protects semiconductor chips so they can be used in actual systems. In high-performance semiconductors, packaging is also closely related to chip-to-chip connectivity, heat dissipation, and power delivery efficiency.

The importance of cooling design is growing as AI equipment becomes more densely integrated. Although rack density still varies considerably across data centers, changes are already emerging in some high-density environments. In Uptime Institute’s 2025 global data center survey, 82% of respondents said the highest-density racks in their facilities were below 30 kW, while some cabinets were reported to exceed 100 kW. High-density AI infrastructure has yet to become the norm, but power density in some areas is already reaching levels that are difficult to accommodate with conventional power and cooling designs.

As rack density increases, cooling methods are also changing. Traditional air cooling remains widely used, but in high-density AI systems, liquid cooling* and near-junction cooling*, which remove heat closer to its source, are emerging as key technologies. In particular, as high-performance processors and memory are packed more closely together, the need for structures that can efficiently dissipate heat increases. As a result, the focus of cooling design is shifting from managing the temperature of the entire data center space to directly controlling heat at points where it is concentrated, such as racks and processors. This is also why semiconductor packaging and thermal design need to be considered together.

* Liquid cooling: A cooling method that uses liquids such as water or coolant instead of air to remove heat generated by servers and semiconductors. It is gaining attention as a way to efficiently manage heat in high-density AI equipment. * Near-junction cooling: A cooling method that removes heat close to heat-generating semiconductors such as processors and memory. By reducing the distance between the heat source and cooling equipment, it helps efficiently manage heat in high-density systems.

☑︎ Data center waste heat: A heating resource for cities in the AI era The energy opportunity extends beyond the data center facilities themselves. In Hamina, Finland, an operational project recovers waste heat from a Google data center and uses it for district heating. Google has been providing recovered heat from its data center to the local district heating network, enabling Haminan Energia to meet approximately 80% of the area’s heating demand in homes, schools, and public service buildings. Waste heat recovery depends on a range of factors, including the local district heating network, climate, data center location, heat recovery infrastructure, and operating costs, meaning the technology cannot be applied uniformly to every data center. Nevertheless, the project is significant because it treats heat generated by data centers not simply as a byproduct to be discarded, but as a resource that can be integrated with local infrastructure. It is an example of extending the concept of data center efficiency beyond internal power consumption and cooling costs to include the utilization of generated heat. Image source: Google

☑︎ Data center waste heat: A heating resource for cities in the AI era

The energy opportunity extends beyond the data center facilities themselves. In Hamina, Finland, an operational project recovers waste heat from a Google data center and uses it for district heating. Google has been providing recovered heat from its data center to the local district heating network, enabling Haminan Energia to meet approximately 80% of the area’s heating demand in homes, schools, and public service buildings.

Waste heat recovery depends on a range of factors, including the local district heating network, climate, data center location, heat recovery infrastructure, and operating costs, meaning the technology cannot be applied uniformly to every data center. Nevertheless, the project is significant because it treats heat generated by data centers not simply as a byproduct to be discarded, but as a resource that can be integrated with local infrastructure. It is an example of extending the concept of data center efficiency beyond internal power consumption and cooling costs to include the utilization of generated heat.

Image source:

Power demand from AI infrastructure is rising rapidly, but it is difficult to continue expanding the power and cooling capacity available to data centers at the same pace. Expanding grid connections, power generation, and transmission infrastructure requires both time and investment, while data centers themselves face physical limits on the amount of power they can receive and the amount of heat they can manage. As a result, alongside securing more energy, another challenge for AI infrastructure is how efficiently that energy can be used.

An important metric in this context is performance per watt*. If a data center can process more inference requests with the same amount of power or deliver the same performance using less power, it can improve both operational efficiency and capacity for expansion. Power efficiency is therefore more than a means to reduce costs; it has become a key factor in determining how far AI services can scale within existing infrastructure constraints.

* Performance per watt: An efficiency metric indicating how much computation or how many tasks can be processed using a given amount of power.

This also changes the role of semiconductor and system design. Processors and memory must be designed with not only processing performance but also power consumption and data movement in mind, while software must coordinate workload allocation and execution timing to use system resources efficiently. To process more workloads with the same amount of power, accelerators, memory, servers, racks, cooling, and software operations must all be coordinated to work toward the same goal.

These changes are also reshaping the criteria for data center investment. Where power can be secured reliably and how the resulting heat can be managed affect where data centers are built and what infrastructure they are connected to. As a result, in addition to land and network access, factors such as power availability, grid connectivity, cooling resources, and operating costs are emerging as key considerations in AI data center investment and site selection. Cloud companies and data center operators must consider not only server equipment but also energy procurement strategies, cooling infrastructure, and how facilities connect to local power grids.

The role of power and cooling technology companies is also growing. Power equipment, cooling systems, heat exchange technologies, data center construction, grid operations, and energy supply chains are becoming more closely integrated with AI infrastructure. McKinsey has noted that as data centers become concentrated in established and emerging hub regions, a mismatch between site locations and grid capacity could lead to energy supply challenges in areas where generation and transmission infrastructure are already constrained.

Data centers in the AI era must deliver both high compute and the operational efficiency required to sustain it.

To reliably process more AI workloads within limited power and thermal-management capacity, processors, memory, servers, cooling, and software must work as a single system alongside energy supply and data center operations. Competition in AI infrastructure is now expanding beyond the performance of individual technologies to encompass energy, supply chains, and operational capabilities.

The next installment will examine how these changes are reshaping the fundamental unit of AI infrastructure design. It will explore the evolution of AI systems as they move beyond individual servers toward rack- and cluster-centric architectures.

What this article says

Something is unclear? Ask about the article — I will explain in plain words.

Do not want to dig deeper? We will sort it out for you.