Research Article算电协同
Rida Fatima, Xingpeng Li
Published 2026-10-04 · arXiv · Credibility S
The rapid growth of AI data centers is creating concentrated electricity demands that can exceed available distribution network headroom and delay interconnection. This paper investigates whether capacity can be used effectively by coordinating flexibility on both sides of the interconnection. A day-ahead mixed integer second order cone programming framework maximizes feasible AI data center IT capacity using utilit…
Abstract, interpretation and reference
Abstract
The rapid growth of AI data centers is creating concentrated electricity demands that can exceed available distribution network headroom and delay interconnection. This paper investigates whether capacity can be used effectively by coordinating flexibility on both sides of the interconnection. A day-ahead mixed integer second order cone programming framework maximizes feasible AI data center IT capacity using utility side conservation voltage reduction and network topology reconfiguration, together with data center workload shifting and reserve constrained uninterruptible power supply storage. Background feeder demand uses voltage dependent ZIP models, while the data center is modeled as constant power demand using Training, Inference, and Mixed workload profiles. A two-stage solution first maximizes interconnection capacity and then minimizes feeder losses for the retained capacity, with numerical relaxation screening. Studies on a 116-bus model derived from the IEEE 123-node feeder show that, at the constrained bus-60 connection point, grid side flexibility, driven almost entirely by network reconfiguration, increases feasible IT capacity by approximately 42.5-45.0%, while data center flexibility alone provides approximately 2.1-5.4% gains. Coordinated operation increases feasible IT capacity from approximately 10.5-10.6 MW under BASE operation to 15.3-16.0 MW, corresponding to gains of approximately 45.9-50.8%; the 16.0 MW Inference result is limited by the adopted UPS reserve requirement. Benefits are strongly location dependent and influenced by workload flexibility, deferral duration, PUE, UPS energy and reserve requirements, and feeder loading. Under a combined adverse operating condition, FULL capacity decreases by only 3.4% relative to the nominal Mixed case while retaining substantial additional capacity over BASE.
Research Article算电协同
Michael Chertkov
Published 2026-10-02 · arXiv · Credibility S
Power grids are beginning to host AI data centers whose demand can change abruptly and in a correlated way, and operators and planners must decide whether existing controllers can absorb the resulting transients and where new flexibility is worth installing. We cast this as a question in stochastic control: how far from optimal is an implementable controller? For controlled diffusions with affine, state-independent …
Abstract, interpretation and reference
Abstract
Power grids are beginning to host AI data centers whose demand can change abruptly and in a correlated way, and operators and planners must decide whether existing controllers can absorb the resulting transients and where new flexibility is worth installing. We cast this as a question in stochastic control: how far from optimal is an implementable controller? For controlled diffusions with affine, state-independent actuation, additive and possibly degenerate noise and quadratic control cost, any Hamilton-Jacobi-Bellman (HJB) subsolution bounds the optimal cost from below and simulation bounds the deployed cost from above, so their gap certifies the permissible suboptimality. We construct subsolutions from path-integral control by completing the control geometry: enlarging the control Gramian until it matches the physical noise makes the problem linearly solvable, and its Feynman-Kac value is an automatic lower bound whose HJB residual is exactly the energy of the fictitious control. A dual noise-deflation construction can be tighter but requires a curvature condition. The geometry yields planning rules: price control authority in proportion to local noise variance, and use shadow values to guide sparse reinforcement. For nonlinear stochastic swing dynamics of the IEEE 118-bus system after a severe load loss, a simple generator controller is certified within 2.8% of optimal under homogeneous forcing; under heterogeneous forcing the gap is 39% with uniform prices and 1.2% once the same total authority is repriced, before any hardware is added.
Research Article算电协同
Suntao Su, Liang Du, Shengyi Wang
Published 2026-09-29 · arXiv · Credibility S
The rapid growth of artificial intelligence (AI) data centers has introduced new challenges to power system operation. As their power demand becomes larger and more variable, quantitatively characterizing their demand flexibility is increasingly important for effective power system coordination. However, heterogeneous workload characteristics and resource requirements make this flexibility difficult to characterize …
Abstract, interpretation and reference
Abstract
The rapid growth of artificial intelligence (AI) data centers has introduced new challenges to power system operation. As their power demand becomes larger and more variable, quantitatively characterizing their demand flexibility is increasingly important for effective power system coordination. However, heterogeneous workload characteristics and resource requirements make this flexibility difficult to characterize directly. This paper proposes a framework for assessing the grid-compatible demand flexibility of AI data centers via batch workload temporal shifting. An averaging-based resource usage processing method is developed to map fine-resolution CPU, GPU and memory usage into unified time intervals compatible with power system operation. A workload temporal scheduling model is then formulated to shift batch workloads while preserving execution continuity, delay constraints, and server resource capacities, and is coupled with a utilization-dependent server power model to translate workload scheduling decisions into server power demand. Two complementary flexibility metrics are evaluated: short-term peak demand shaving and the maximum duration of sustained power reduction. Numerical results based on real GPU cluster traces demonstrate that workload temporal shifting can provide quantifiable and grid-compatible demand flexibility for AI data centers with limited disruption to computing workloads.
Research Article芯片与算力
Yixing Li, Mark Fenton, Matthew Kaufeler, Ka Ming Leung, Xin Ai, Zhiyu Zeng
Published 2026-10-01 · arXiv · Credibility S
The emerging of large language models (LLMs) has posed significant challenges to the thermal management of data center. Intense GPU computation for LLMs results in localized hotspots. Moreover, spiking thermal loads during training and inference bursts make real-time cooling response more difficult to predict and control. Thermal-aware capacity planning of data center requires massive expensive high-fidelity CFD sim…
Abstract, interpretation and reference
Abstract
The emerging of large language models (LLMs) has posed significant challenges to the thermal management of data center. Intense GPU computation for LLMs results in localized hotspots. Moreover, spiking thermal loads during training and inference bursts make real-time cooling response more difficult to predict and control. Thermal-aware capacity planning of data center requires massive expensive high-fidelity CFD simulations. AI models can perform real-time prediction for unseen designs. However, existing works either have large prediction error, or have over-simplified assumptions for data center operations. This work presents an AI-driven framework that can perform thermal-aware capacity planning for a real-world data center in seconds. The embedded AI model learns from numerous key parameters (rack power, server power, server placement, HVAC settings etc.), and provides temperature prediction within milliseconds. This AI model is tested against high-fidelity CFD simulations, and results show that for unseen data center designs, model can achieve high accuracy with 10000X speedup. Driven by the AI model, the authors design the thermal-aware capacity planning framework. This framework can help data center designers and operators instantaneously optimize both workload distribution and HVAC cooling efficiency.
Research Article热管理与液冷
Davide Bray, Francesca Dominici
Published 2026-10-02 · arXiv · Credibility S
The rapid build-out of hyperscale computing for artificial-intelligence workloads has made data-center noise a recurring source of community concern in the United States, yet calibrated, standards-grade acoustic data remain scarce. This narrative review synthesizes direct measurements, regulatory and legal records, investigative journalism, and relevant engineering literature to assess what is currently known and id…
Abstract, interpretation and reference
Abstract
The rapid build-out of hyperscale computing for artificial-intelligence workloads has made data-center noise a recurring source of community concern in the United States, yet calibrated, standards-grade acoustic data remain scarce. This narrative review synthesizes direct measurements, regulatory and legal records, investigative journalism, and relevant engineering literature to assess what is currently known and identify the measurements needed for credible health inference. Available sources report heterogeneous acoustic values at residences, property lines, and near-source positions, with individual readings from 39 to above 70 dB(A). Because location, metric, duration, instrumentation, and provenance differ, these values are not directly comparable community exposures, do not define a pooled distribution, and do not support an aggregate comparison with nighttime limits. Among cases with spectral detail, all but one reported continuous tonal noise associated with complaints, most often low-frequency content near 70-140 Hz, that A-weighted metrics alone may under-represent. Engineering evidence further suggests that cooling architecture may be a major determinant of facility sound emission; the cited 15-30 dB span is an equipment-level sound-power estimate, not a measured facility- or residential-receptor difference, and remains a hypothesis pending comparative field measurements. We propose a standards-based field-measurement and modeling protocol designed to generate frequency-resolved exposure data at named facilities. Such exposure characterization is a necessary foundation for credible assessment of community health effects and environmental justice implications.
Research Article芯片与算力
Eric Masanet
Published 2026-10-03 · arXiv · Credibility S
Concerns are growing about the water use of data centers. To quantify the scales, drivers, and possible trajectories of this water use, analytical estimation is required. However, current estimation methods and results vary widely due, in part, to a lack of established best practices. This variance is illustrated with a U.S. case study demonstrating a roughly 24-fold difference in possible outcomes based on observed…
Abstract, interpretation and reference
Abstract
Concerns are growing about the water use of data centers. To quantify the scales, drivers, and possible trajectories of this water use, analytical estimation is required. However, current estimation methods and results vary widely due, in part, to a lack of established best practices. This variance is illustrated with a U.S. case study demonstrating a roughly 24-fold difference in possible outcomes based on observed U.S. study design variations. To address these challenges, this review synthesizes 14 best practices for advancing the science and utility of data center water use estimation. These proposed best practices can also be used as critical questions for stakeholders when evaluating the quality of any study. Using a structured rubric, it further assesses the literature to date (31 studies) against these best practices. The assessment reveals important opportunities for improvement and standardization in future study scopes, clarity, methods, and scientific knowledge building utility. It further finds that, while many studies provide reasonable transparency, scope definitions, replicability, and uncertainty treatment, many fall short in scientific justification, data representativeness, transparency on hydropower assumptions, terminological consistency, and meaningful discussions of limitations and future work. Recommendations are proposed for analysts, stakeholders, and policymakers to leverage these findings to help build a vibrant data center water use research community. Finally, this review can also be used as a basic primer for understanding data center cooling configurations and water use among analysts, policymakers, the media, and the public.
Research ArticleAI 运维优化
Junyu Lin, Wenjie Liu, Shunbo Lei, Wentian Lu, Jianhui Wang, Junhong Liu
Published 2026-09-30 · arXiv · Credibility S
The rapid proliferation of data centers (DCs), driven by cloud computing and artificial intelligence (AI), has led to massive energy demand and carbon emissions, posing significant sustainability challenges. Carbon-aware optimization in geographically distributed data centers has been widely studied. Most existing approaches mainly focus on operational carbon emissions from server usage. However, existing literature…
Abstract, interpretation and reference
Abstract
The rapid proliferation of data centers (DCs), driven by cloud computing and artificial intelligence (AI), has led to massive energy demand and carbon emissions, posing significant sustainability challenges. Carbon-aware optimization in geographically distributed data centers has been widely studied. Most existing approaches mainly focus on operational carbon emissions from server usage. However, existing literature often ignores workload-induced thermal stress, which accelerates nonlinear hardware degradation. This leads to more frequent server replacements and ultimately increases embodied carbon emissions. To address these limitations, we propose a comprehensive carbon life-cycle modeling framework for distributed data centers. Apart from operational carbon emissions, this work combines workload scheduling with a utilization-dependent exponential aging model to evaluate long-term carbon costs from server degradation. In order to solve the proposed optimization model in an online and privacy-preserving manner, an enhanced Lyapunov framework with time-varying queue shifting (TVQS) is first introduced to handle system uncertainties. Then, a zero-sum perturbation-based alternating direction method of multipliers (ZSP-ADMM) framework is developed to enable distributed coordination across geographically separated data centers while protecting locally exchanged workload information. Simulation results demonstrate that the proposed approach achieves up to 13.0% lower carbon emissions and 12.6% lower operational costs compared with benchmarks.
Research Article算电协同
Prabhat Ranjan Bana, Novan Zakkia, Jean-Philippe Hasler, Christer Danielsson
Published 2026-09-27 · arXiv · Credibility S
The rapid expansion of large-scale AI data centers (AIDC) is introducing new stability challenges, particularly in weak or low-inertia networks characterized by fast, step-like demand variations and strict requirements on voltage and dynamic performance. This paper investigates the use of grid-forming (GFM) Enhanced STATCOMs (E-STATCOMs) to support reliable integration of such facilities. A power-admittance-based li…
Abstract, interpretation and reference
Abstract
The rapid expansion of large-scale AI data centers (AIDC) is introducing new stability challenges, particularly in weak or low-inertia networks characterized by fast, step-like demand variations and strict requirements on voltage and dynamic performance. This paper investigates the use of grid-forming (GFM) Enhanced STATCOMs (E-STATCOMs) to support reliable integration of such facilities. A power-admittance-based linear modelling framework is developed to capture system interactions and is validated through detailed EMT simulations. The results demonstrate that E-STATCOMs provide fast, well-damped responses to abrupt load changes while effectively mitigating low-frequency oscillations and interactions with network resonances. By enabling tunable dynamic behavior via a load balancer, virtual impedance, and coordinated active-reactive power support, the proposed approach allows precise shaping of system response and improved regulation at the point of connection. These features make E-STATCOMs a flexible and scalable solution for integrating large data centers into weak grids and long transmission systems, supported by a design-oriented framework that facilitates parameter selection and performance assessment without extensive reliance on EMT studies to meet grid codes and AIDC interconnection requirements.
Research Article算电协同
Farnaz Farid, Tashfia Towkee, Sania Nasreen, Sami bin Azad
Published 2026-09-25 · arXiv · Credibility S
As artificial intelligence (AI) becomes embedded in everyday life, its environmental footprint, particularly water consumption remains largely invisible. While energy and carbon impacts are widely recognized, the substantial freshwater demands of data center cooling and electricity generation receive little attention. To address this gap, we introduce SustainAI, a water-aware, closed-loop framework incorporating env…
Abstract, interpretation and reference
Abstract
As artificial intelligence (AI) becomes embedded in everyday life, its environmental footprint, particularly water consumption remains largely invisible. While energy and carbon impacts are widely recognized, the substantial freshwater demands of data center cooling and electricity generation receive little attention. To address this gap, we introduce SustainAI, a water-aware, closed-loop framework incorporating environmental accountability into AI deployment. SustainAI integrates real-time water metering, a hallucination-aware penalty model, and a water-aware routing algorithm that accounts for regional water stress. Evaluated via Small Language Models (SLMs) extracting health misinformation, results reveal an 11-fold variation in water footprint across geographically distributed data centers (0.0477 mL to 0.5360 mL per inference). Across 1,335 inference runs, the system consumed approximately 399 mL of water but produced only 240 correct outputs, demonstrating that substantial resources are spent on inaccurate responses. Crucially, SustainAI extends beyond technical optimization through a Care by Design lens, framing AI sustainability around relational ethics, regional equity, and ecological stewardship. By combining water monitoring, adaptive accountability, and Care by Design principles, SustainAI provides a practical foundation for integrating ethical care and environmental responsibility into AI infrastructure design and lifecycle management.
Research Article算电协同
Ziang Liu, Ruizhang Yang, Xin Cui, Francis Yunhe Hou
Published 2026-09-22 · arXiv · Credibility S
The rapid growth of large language model training and serving is driving AI data centers (AIDCs) toward gigawatt scale. Unlike conventional commercial loads, AIDCs possess significant operational flexibility through dynamic voltage and frequency scaling (DVFS) of training and inference workloads, while periodic model checkpointing can induce abrupt power drops and rebounds that erode operating reserves and increase …
Abstract, interpretation and reference
Abstract
The rapid growth of large language model training and serving is driving AI data centers (AIDCs) toward gigawatt scale. Unlike conventional commercial loads, AIDCs possess significant operational flexibility through dynamic voltage and frequency scaling (DVFS) of training and inference workloads, while periodic model checkpointing can induce abrupt power drops and rebounds that erode operating reserves and increase transmission congestion risks. Coordinating AIDC operation with grid scheduling under these unique operational characteristics is challenging because grid and AIDC operators are generally unwilling to share proprietary data and decision-making authority. This paper proposes a hierarchical privacy-preserving coordinated operation scheme between the power grid and AIDCs to address this gap. The proposed scheme contains three phases. In Phase I, the grid operator computes a certified inner approximation of the AIDCs security region for subsequent coordination. In Phase II, the AIDC operator coordinates training and inference AIDCs to optimize workload allocation within the certified security region and generate power schedules and checkpoint alerts. In Phase III, the grid operator solves a checkpoint-aware two-stage robust optimal power flow (OPF) considering renewable generation and checkpoint uncertainties. By exchanging only compact interface information, the framework preserves the privacy of both grid and AIDCs, avoids frequent iterative communication, and enables secure coordination with guaranteed feasibility. Numerical studies on a modified IEEE 14-bus system and a modified NYISO system demonstrate the effectiveness, robustness, and security of the proposed framework.
Research Article芯片与算力
Wanqun Yang, Jun Chen
Published 2026-09-28 · arXiv · Credibility S
This paper develops an integrated modeling and nonlinear model predictive control (NMPC) framework for coordinating thermal management, flexible workload scheduling, wave-power utilization, and battery operation in a wave-powered subsea data center. Realistic data center workloads are constructed from job-level CPU, memory, and GPU measurements from the MIT Supercloud dataset and divided into interactive and delay-t…
Abstract, interpretation and reference
Abstract
This paper develops an integrated modeling and nonlinear model predictive control (NMPC) framework for coordinating thermal management, flexible workload scheduling, wave-power utilization, and battery operation in a wave-powered subsea data center. Realistic data center workloads are constructed from job-level CPU, memory, and GPU measurements from the MIT Supercloud dataset and divided into interactive and delay-tolerant flexible jobs. Thermal behavior is represented by a three-node lumped model of the IT equipment, recirculating nitrogen, and pressure hull with surrounding seawater as the thermal boundary. The NMPC jointly optimizes the flexible workload power budget and cooling command subject to thermal, battery, and workload constraints. Closed-loop simulations under different workload, thermal, battery, and renewable-generation conditions demonstrate that the proposed framework maintains thermal safety while adapting cooling operation and flexible workload execution to wave-power availability and battery state-of-charge. The parametric studies show that battery capacity and wave-generation capacity strongly affect battery availability and flexible-workload queue accumulation, while excessive renewable generation capacity may lead to increased energy curtailment. Monte Carlo and distance-correlation analyses further show that flexible-job delay is relatively insensitive to the investigated system parameters, whereas terminal battery state-of-charge is primarily influenced by battery energy capacity and wave generation capacity.
Research Article算电协同
Zhanhua Pan, Xiao Liu, Zhilong Cao, Jianhong Wang, Dawei Qiu
Published 2026-09-28 · arXiv · Credibility S
Power system operation is a safety-critical sequential decision-making problem, making it a natural testbed for reinforcement learning (RL). However, existing RL environments for power systems are often narrow in scope and computationally limited by CPU-based simulation workflows, making large-scale evaluation difficult. We introduce PowerZooJax, a JAX-based benchmark suite for RL in power system operation. It provi…
Abstract, interpretation and reference
Abstract
Power system operation is a safety-critical sequential decision-making problem, making it a natural testbed for reinforcement learning (RL). However, existing RL environments for power systems are often narrow in scope and computationally limited by CPU-based simulation workflows, making large-scale evaluation difficult. We introduce PowerZooJax, a JAX-based benchmark suite for RL in power system operation. It provides five constrained Markov decision process tasks spanning generation, transmission, distribution, distributed energy resources, and data center microgrid. By rewriting power flow, economic dispatch, market clearing, and device dynamics as JAX computation graphs, PowerZooJax keeps the entire training and evaluation loop on the GPU. Experiments show substantial speedups over CPU-based simulations and demonstrate standardized evaluation of policy returns, safety violations, and out-of-distribution stress conditions. Our open-source benchmark is available at: https://github.com/powerzoojax/PowerZooJax.
Research Article热管理与液冷
Zixu Han, Peng Zhang
Published 2026-09-11 · arXiv · Credibility S
The rapid development of liquid-cooled data centers has imposed imperative demands on the performance of liquid cooling plate. The density-based topology optimization (TO) is an effective approach to resolving the growing thermal-hydraulic performance requirements of liquid cooling plate. However, existing TO methods can hardly optimize convective heat transfer directly which is the intrinsic heat transfer mechanism…
Abstract, interpretation and reference
Abstract
The rapid development of liquid-cooled data centers has imposed imperative demands on the performance of liquid cooling plate. The density-based topology optimization (TO) is an effective approach to resolving the growing thermal-hydraulic performance requirements of liquid cooling plate. However, existing TO methods can hardly optimize convective heat transfer directly which is the intrinsic heat transfer mechanism, due to the highly complex and evolving structural topologies, varying flow and temperature fields, making it extremely challenging to explicitly describe the heat transfer coefficient and heat transfer area during TO process. A convective heat transfer topology optimization (CTO) method is proposed in this study, where the iteratively evolving heat transfer coefficient is explicitly depicted by the field synergy theory in the thermal objective, and directly described by the velocity and temperature fields without relying on specific geometry. Combined with the explicit depiction of heat transfer area by the fractal geometry theory, a CTO framework is built for a direct optimization of convective heat transfer under both the laminar and turbulent flow conditions. The CTO tends to generate more hierarchical and directional structural topologies in optimization results, which is conducive to reducing low-velocity stagnation zones and improving flow direction in branched channels, achieving enhanced synergy and thermal-hydraulic performance in the optimized liquid cooling plates. Compared with the TO results without incorporation of field synergy theory, the CTO can reduce average temperature rise by 20% while improving the Nusselt number by 15% under laminar flow conditions, and reduce maximum temperature rise by 10.2% and pressure drop by 25% under turbulent flow conditions.
Research Article芯片与算力
Rui Lu, Rui Ge, Huanghuang Liang, Xiaobo Zhou, Dan Wang
Published 2026-09-14 · arXiv · Credibility S
Large language model (LLM) inference in AI datacenters creates a coupled control problem between GPU serving and facility cooling. Raising ambient temperature setpoints can reduce cooling energy and carbon, but also shrinks thermal headroom, induces GPU throttling, and leads to Service-Level-Objective (SLO) violations. In this paper, we study joint cooling--computing control for LLM inference: minimizing per-job GPU…
Abstract, interpretation and reference
Abstract
Large language model (LLM) inference in AI datacenters creates a coupled control problem between GPU serving and facility cooling. Raising ambient temperature setpoints can reduce cooling energy and carbon, but also shrinks thermal headroom, induces GPU throttling, and leads to Service-Level-Objective (SLO) violations. In this paper, we study joint cooling--computing control for LLM inference: minimizing per-job GPU-plus-cooling energy while satisfying thermal safety and latency SLO constraints. We present ETCInfer, an energy-efficient, thermal-aware scheduler that selects a pre-job Computer Room Air Conditioner (CRAC) setpoint and adapts per-GPU frequency and micro-batch size during execution. ETCInfer builds compact physics-informed control models by calibrating GPU heat generation, chassis heat dissipation, CRAC power, and prefill/decode latency relations from telemetry. These models estimate hidden thermal states and time-to-throttle, enabling the scheduler to evaluate energy, temperature, and latency before applying an action. We formulate this joint setpoint--frequency--micro-batch control problem as a partially observable Markov decision process and design ETCAdapter, a learning-based controller that minimizes per-job energy under thermal safety and SLO constraints. We implement ETCInfer as a coordination layer over typical inference and cluster management stacks. Evaluation across real-trace simulation and validation experiments shows that ETCInfer reduces total job energy by up to 33.1%, thermal throttle exposure by up to 92.9%, and keeps SLO violation rates below 0.7% even at ambient temperatures up to $48^{\circ}\mathrm{C}$.
Research Article算电协同
Krishna Chaitanya Sunkara
Published 2026-09-25 · arXiv · Credibility S
GPU-dense AI data centers need to run on liquid cooling as air simply cannot shed the heat at these power densities. Yet the cooling loops themselves are blind to what workloads are about to run; they crank up flow only after a sensor catches a temperature climb, which can take 30 to 50 seconds. We built Job-Class Thermal Intent (JCTI) to close that window. The scheduler already knows a job is coming and what class …
Abstract, interpretation and reference
Abstract
GPU-dense AI data centers need to run on liquid cooling as air simply cannot shed the heat at these power densities. Yet the cooling loops themselves are blind to what workloads are about to run; they crank up flow only after a sensor catches a temperature climb, which can take 30 to 50 seconds. We built Job-Class Thermal Intent (JCTI) to close that window. The scheduler already knows a job is coming and what class it belongs to; JCTI feeds that information straight to the cooling controller so it can stage coolant before the heat shows up. We pulled the thermal signatures for each job class out of MLPerf GPU power traces and tuned arrival patterns against Alibaba cluster data. Over 120 paired Monte Carlo trials the numbers come out to 56.4% fewer thermal violations and 60.2% less cumulative overshoot than a straight PI loop. As AI data centers evolving towards gigawatt grid loads with highly fluctuating power swings, thermally-aware scheduling reduces sudden demand and improves load prediction in grid side. Cooling and scheduling have been running as two separate systems for years despite each one knowing something the other needs, JCTI wires them together.
Research Article算电协同
Garrett Alston, Nancy Love, Rabab Haider
Published 2026-09-21 · arXiv · Credibility S
Data centers are being developed at an unprecedented pace, yet their energy and water impacts, and the spatial and temporal distribution of these impacts, remain poorly characterized. Data centers consume water for cooling (direct) and through electricity generation (indirect). Decisions on siting and cooling technology result in water-energy trade-offs that extend impacts beyond the facility's location. Existing as…
Abstract, interpretation and reference
Abstract
Data centers are being developed at an unprecedented pace, yet their energy and water impacts, and the spatial and temporal distribution of these impacts, remain poorly characterized. Data centers consume water for cooling (direct) and through electricity generation (indirect). Decisions on siting and cooling technology result in water-energy trade-offs that extend impacts beyond the facility's location. Existing assessment frameworks rely on facility efficiency metrics and average grid water intensity factors, suppressing the temporal impacts of data center load and generation availability. They also attribute indirect consumption to the facility's location rather than to the generators (and corresponding hydrologic regions) that respond to the added load, misattributing spatial impacts. To close this gap, we develop a computational model of the data center-energy-water nexus that links facility cooling and electricity demand with hourly economic dispatch, generator-level water consumption, and monthly subbasin depletion. Built on open-source data, the model resolves where and when water is consumed, and where this consumption compounds existing water risk or creates new risk. Using the model, we study different cooling configurations and proposed developments in the state of Michigan. Air-cooled data centers halve total water consumption relative to evaporative cooling, but increase electricity demand and raise indirect water consumption by one-third, shifting the water footprint from the facility to generators. Mapping these changes to subbasins reveals depletion increases beyond the data center sites, in regions that facility-level reporting may overlook. These results show that data center water and energy impacts cannot be assessed in isolation, motivating the need for integrated modeling to inform siting, design, and reporting practices.
Research Article能效优化
Sharifa Sultana, Syed Ishtiaque Ahmed
Published 2026-09-20 · arXiv · Credibility S
Standard data center sustainability metrics, including Power Usage Effectiveness (PUE), Water Usage Effectiveness (WUE), and Carbon Usage Effectiveness (CUE), measure a facility's resource use and emissions intensity, normalized to IT energy use, without directly representing local resource scarcity, infrastructure capacity, or social footprint. This gap has become politically consequential. In the first quarter of …
Abstract, interpretation and reference
Abstract
Standard data center sustainability metrics, including Power Usage Effectiveness (PUE), Water Usage Effectiveness (WUE), and Carbon Usage Effectiveness (CUE), measure a facility's resource use and emissions intensity, normalized to IT energy use, without directly representing local resource scarcity, infrastructure capacity, or social footprint. This gap has become politically consequential. In the first quarter of 2026 alone, local opposition delayed or canceled roughly $130 billion in projects across the United States, driven overwhelmingly by recurring concerns over water use, power demand, infrastructure capacity, and transparency rather than internal efficiency, matching the total for all of 2025 [11]. We propose a five-category local impact audit framework covering efficiency, water stewardship, carbon and renewables, regulatory compliance, and local disclosure. The framework is designed for recurring quarterly assessment and independent verification against public records. We illustrate its application using publicly available data from three Illinois facilities that are currently at the center of local policy disputes, and we examine the data-access barriers that constrain independent verification. We position this framework as both a research contribution and a practical instrument for county-level policymakers evaluating data center permitting and moratorium decisions.
Research Article算电协同
Bojun Du, Hongyang Jia, Tonghui Li, Qingchun Hou, Ze Wang, Ershun Du, Ning Zhang
Published 2026-09-09 · arXiv · Credibility S
AI data centers may face power supply shortages during certain periods, requiring operators to shift large language model (LLM) inference workloads spatially to maintain service rates. However, existing workload-shifting methods typically assume that any data center with sufficient computing resources can immediately serve shifted requests, which may lead to infeasible transfers and unserved demand. This letter prop…
Abstract, interpretation and reference
Abstract
AI data centers may face power supply shortages during certain periods, requiring operators to shift large language model (LLM) inference workloads spatially to maintain service rates. However, existing workload-shifting methods typically assume that any data center with sufficient computing resources can immediately serve shifted requests, which may lead to infeasible transfers and unserved demand. This letter proposes model commitment (MC), a mixed-integer linear programming framework that jointly schedules model deployment and cross-site request routing under power constraints and electricity-price signals. First, MC formulates the intertemporal coupling introduced by model replica loading. Second, it translates prefill and decode latency requirements into the amount of demand that each replica can serve. Case studies based on real-world data show that MC enables AI data center operators to achieve a 100% service rate under time-varying grid conditions and reduce total operating cost by 29.0%.
Research Article热管理与液冷
James Teague, Ashmita Rajmohan, Yannick Muehlhaeuser
Published 2026-09-16 · arXiv · Credibility S
Proposals for international agreements that limit frontier AI development depend on verification, and a central challenge is detecting undeclared compute facilities used to evade restrictions. Underwater data centers (UDCs) have been suggested as one such evasion vector, but their feasibility at frontier scale and their detectability have not been seriously assessed. We examine current UDC deployments, evaluate cons…
Abstract, interpretation and reference
Abstract
Proposals for international agreements that limit frontier AI development depend on verification, and a central challenge is detecting undeclared compute facilities used to evade restrictions. Underwater data centers (UDCs) have been suggested as one such evasion vector, but their feasibility at frontier scale and their detectability have not been seriously assessed. We examine current UDC deployments, evaluate construction and maintenance complexity relative to land-based facilities, and analyse the feasibility of a 100,000 H100-equivalent training run underwater. We find that power delivery and cooling are tractable, but interconnect and the hands-on maintenance that large training runs require are severe obstacles - surmountable only by a well-resourced state actor accepting large cost and schedule penalties, and only where concealment, rather than efficiency, is the objective. We then assess detectability through thermal, acoustic, optical and synthetic-aperture-radar (SAR) surveillance. Thermal detection of an operational pod is unlikely outside shallow, calm water; acoustic detection is marginally more effective, but faces limitations in attribution; and optical/SAR monitoring is most powerful during construction and maintenance, when the pressure-vessel fabrication base and the cable-laying fleet create distinctive signatures for AIS-tracking. We conclude that UDCs are a comparatively unlikely evasion route relative to underground or industrially disguised land-based facilities, but the residual risk is non-zero and warrants operationalising the detection modalities discussed.
Research Article算电协同
Dayuan Chen, Ziliang Zong
Published 2026-09-13 · arXiv · Credibility S
The rapid rise of generative AI is accelerating cloud data center expansion, with electricity demand projected to double by 2026. Because carbon-intensity varies by more than 5.5x across grids and times of day, where and when inference tasks execute significantly affects operational emissions. We address this issue with three aspects in this paper. First, we compile a global alignment dataset unifying 140 operationa…
Abstract, interpretation and reference
Abstract
The rapid rise of generative AI is accelerating cloud data center expansion, with electricity demand projected to double by 2026. Because carbon-intensity varies by more than 5.5x across grids and times of day, where and when inference tasks execute significantly affects operational emissions. We address this issue with three aspects in this paper. First, we compile a global alignment dataset unifying 140 operational and planned cloud regions across 8 major providers with five-minute carbon-intensity traces for 145 grid regions from 2022 to 2024, revealing that 50% of current sites lie in medium-to-high carbon-intensity grids, indicating a siting-carbon mismatch and unrealized carbon reduction potential. Second, we develop CATS (Carbon-Aware Task Simulator), a flexible trace-driven framework that profiles six AI inference tasks across multiple GPU types, synthesizes realistic diurnal curve, geographical and task mixes, and SLA constraints, and evaluates spatial and temporal schedulers against two baselines while reporting comprehensive metrics including carbon emissions, energy consumption, runtime, queue delay, and hardware utilization. Third, we quantify achievable CO2 savings under realistic constraints: in a 24-hour trace with 600,000 tasks at fleet utilization of 0.37, spatial shifting reduces CO2 by 38.4% versus speed-first baseline, while temporal shifting yields 16% savings with bounded SLA violations at 3.27%. These results advocate locating future data centers in low carbon-intensity grids and demonstrate that carbon-aware scheduling on today's fleets can achieve substantial operational emissions reduction.
Research Article算电协同
Yubo Song, Rui Kong, Takuro Umihara, Pooya Davari, Frede Blaabjerg, Subham Sahoo
Published 2026-09-10 · arXiv · Credibility S
The rapid growth of artificial intelligence (AI) computing is transforming data centers into large, dynamic electrical loads. Their deployment is primarily constrained by energy availability and grid-connection capacity, which is further aggravated by the ability of power-delivery architectures, control systems, and computing workloads to operate reliably during fast grid disturbances. This article presents a techno…
Abstract, interpretation and reference
Abstract
The rapid growth of artificial intelligence (AI) computing is transforming data centers into large, dynamic electrical loads. Their deployment is primarily constrained by energy availability and grid-connection capacity, which is further aggravated by the ability of power-delivery architectures, control systems, and computing workloads to operate reliably during fast grid disturbances. This article presents a technological perspective on AI data centers as grid-interactive computing systems. First, it reviews grid-integration bottlenecks, evolving connection policies, grid-code requirements, which has fostered new technological trends via spatio-temporal flexibility available through workload orchestration, cooling systems, on-site resources, and energy storage. Second, it maps the evolution of power-delivery architectures from medium-voltage grid interfaces to chip-level, discussing higher-voltage DC distribution, solid-state transformers, wide-bandgap devices, advanced chip-level power delivery, and liquid cooling. Third, it establishes a three-level stability framework spanning rack-level DC-bus dynamics, facility-level converter interactions, and system-level grid-coupled behavior. The framework connects dominant instability mechanisms, including constant power load effects, impedance interactions, forced oscillations, and operating-mode transitions, with suitable modeling, assessment, and mitigation approaches. Synthesizing these topics, this article highlights grid-to-chip co-design as a central requirement for scalable AI infrastructure, linking computing workloads, power-delivery systems, energy buffers, and grid operation.