Research Article算电协同
Garrett Alston, Nancy Love, Rabab Haider
Published 2026-09-21 · arXiv · Credibility S
Data centers are being developed at an unprecedented pace, yet their energy and water impacts, and the spatial and temporal distribution of these impacts, remain poorly characterized. Data centers consume water for cooling (direct) and through electricity generation (indirect). Decisions on siting and cooling technology result in water-energy trade-offs that extend impacts beyond the facility's location. Existing as…
Abstract, interpretation and reference
Abstract
Data centers are being developed at an unprecedented pace, yet their energy and water impacts, and the spatial and temporal distribution of these impacts, remain poorly characterized. Data centers consume water for cooling (direct) and through electricity generation (indirect). Decisions on siting and cooling technology result in water-energy trade-offs that extend impacts beyond the facility's location. Existing assessment frameworks rely on facility efficiency metrics and average grid water intensity factors, suppressing the temporal impacts of data center load and generation availability. They also attribute indirect consumption to the facility's location rather than to the generators (and corresponding hydrologic regions) that respond to the added load, misattributing spatial impacts. To close this gap, we develop a computational model of the data center-energy-water nexus that links facility cooling and electricity demand with hourly economic dispatch, generator-level water consumption, and monthly subbasin depletion. Built on open-source data, the model resolves where and when water is consumed, and where this consumption compounds existing water risk or creates new risk. Using the model, we study different cooling configurations and proposed developments in the state of Michigan. Air-cooled data centers halve total water consumption relative to evaporative cooling, but increase electricity demand and raise indirect water consumption by one-third, shifting the water footprint from the facility to generators. Mapping these changes to subbasins reveals depletion increases beyond the data center sites, in regions that facility-level reporting may overlook. These results show that data center water and energy impacts cannot be assessed in isolation, motivating the need for integrated modeling to inform siting, design, and reporting practices.
Research Article能效优化
Sharifa Sultana, Syed Ishtiaque Ahmed
Published 2026-09-20 · arXiv · Credibility S
Standard data center sustainability metrics, including Power Usage Effectiveness (PUE), Water Usage Effectiveness (WUE), and Carbon Usage Effectiveness (CUE), measure a facility's resource use and emissions intensity, normalized to IT energy use, without directly representing local resource scarcity, infrastructure capacity, or social footprint. This gap has become politically consequential. In the first quarter of …
Abstract, interpretation and reference
Abstract
Standard data center sustainability metrics, including Power Usage Effectiveness (PUE), Water Usage Effectiveness (WUE), and Carbon Usage Effectiveness (CUE), measure a facility's resource use and emissions intensity, normalized to IT energy use, without directly representing local resource scarcity, infrastructure capacity, or social footprint. This gap has become politically consequential. In the first quarter of 2026 alone, local opposition delayed or canceled roughly $130 billion in projects across the United States, driven overwhelmingly by recurring concerns over water use, power demand, infrastructure capacity, and transparency rather than internal efficiency, matching the total for all of 2025 [11]. We propose a five-category local impact audit framework covering efficiency, water stewardship, carbon and renewables, regulatory compliance, and local disclosure. The framework is designed for recurring quarterly assessment and independent verification against public records. We illustrate its application using publicly available data from three Illinois facilities that are currently at the center of local policy disputes, and we examine the data-access barriers that constrain independent verification. We position this framework as both a research contribution and a practical instrument for county-level policymakers evaluating data center permitting and moratorium decisions.
Research Article算电协同
Ziang Liu, Ruizhang Yang, Xin Cui, Francis Yunhe Hou
Published 2026-09-22 · arXiv · Credibility S
The rapid growth of large language model training and serving is driving AI data centers (AIDCs) toward gigawatt scale. Unlike conventional commercial loads, AIDCs possess significant operational flexibility through dynamic voltage and frequency scaling (DVFS) of training and inference workloads, while periodic model checkpointing can induce abrupt power drops and rebounds that erode operating reserves and increase …
Abstract, interpretation and reference
Abstract
The rapid growth of large language model training and serving is driving AI data centers (AIDCs) toward gigawatt scale. Unlike conventional commercial loads, AIDCs possess significant operational flexibility through dynamic voltage and frequency scaling (DVFS) of training and inference workloads, while periodic model checkpointing can induce abrupt power drops and rebounds that erode operating reserves and increase transmission congestion risks. Coordinating AIDC operation with grid scheduling under these unique operational characteristics is challenging because grid and AIDC operators are generally unwilling to share proprietary data and decision-making authority. This paper proposes a hierarchical privacy-preserving coordinated operation scheme between the power grid and AIDCs to address this gap. The proposed scheme contains three phases. In Phase I, the grid operator computes a certified inner approximation of the AIDCs security region for subsequent coordination. In Phase II, the AIDC operator coordinates training and inference AIDCs to optimize workload allocation within the certified security region and generate power schedules and checkpoint alerts. In Phase III, the grid operator solves a checkpoint-aware two-stage robust optimal power flow (OPF) considering renewable generation and checkpoint uncertainties. By exchanging only compact interface information, the framework preserves the privacy of both grid and AIDCs, avoids frequent iterative communication, and enables secure coordination with guaranteed feasibility. Numerical studies on a modified IEEE 14-bus system and a modified NYISO system demonstrate the effectiveness, robustness, and security of the proposed framework.
Research Article算电协同
Dayuan Chen, Ziliang Zong
Published 2026-09-13 · arXiv · Credibility S
The rapid rise of generative AI is accelerating cloud data center expansion, with electricity demand projected to double by 2026. Because carbon-intensity varies by more than 5.5x across grids and times of day, where and when inference tasks execute significantly affects operational emissions. We address this issue with three aspects in this paper. First, we compile a global alignment dataset unifying 140 operationa…
Abstract, interpretation and reference
Abstract
The rapid rise of generative AI is accelerating cloud data center expansion, with electricity demand projected to double by 2026. Because carbon-intensity varies by more than 5.5x across grids and times of day, where and when inference tasks execute significantly affects operational emissions. We address this issue with three aspects in this paper. First, we compile a global alignment dataset unifying 140 operational and planned cloud regions across 8 major providers with five-minute carbon-intensity traces for 145 grid regions from 2022 to 2024, revealing that 50% of current sites lie in medium-to-high carbon-intensity grids, indicating a siting-carbon mismatch and unrealized carbon reduction potential. Second, we develop CATS (Carbon-Aware Task Simulator), a flexible trace-driven framework that profiles six AI inference tasks across multiple GPU types, synthesizes realistic diurnal curve, geographical and task mixes, and SLA constraints, and evaluates spatial and temporal schedulers against two baselines while reporting comprehensive metrics including carbon emissions, energy consumption, runtime, queue delay, and hardware utilization. Third, we quantify achievable CO2 savings under realistic constraints: in a 24-hour trace with 600,000 tasks at fleet utilization of 0.37, spatial shifting reduces CO2 by 38.4% versus speed-first baseline, while temporal shifting yields 16% savings with bounded SLA violations at 3.27%. These results advocate locating future data centers in low carbon-intensity grids and demonstrate that carbon-aware scheduling on today's fleets can achieve substantial operational emissions reduction.
Research Article芯片与算力
Rui Lu, Rui Ge, Huanghuang Liang, Xiaobo Zhou, Dan Wang
Published 2026-09-14 · arXiv · Credibility S
Large language model (LLM) inference in AI datacenters creates a coupled control problem between GPU serving and facility cooling. Raising ambient temperature setpoints can reduce cooling energy and carbon, but also shrinks thermal headroom, induces GPU throttling, and leads to Service-Level-Objective (SLO) violations. In this paper, we study joint cooling--computing control for LLM inference: minimizing per-job GPU…
Abstract, interpretation and reference
Abstract
Large language model (LLM) inference in AI datacenters creates a coupled control problem between GPU serving and facility cooling. Raising ambient temperature setpoints can reduce cooling energy and carbon, but also shrinks thermal headroom, induces GPU throttling, and leads to Service-Level-Objective (SLO) violations. In this paper, we study joint cooling--computing control for LLM inference: minimizing per-job GPU-plus-cooling energy while satisfying thermal safety and latency SLO constraints. We present ETCInfer, an energy-efficient, thermal-aware scheduler that selects a pre-job Computer Room Air Conditioner (CRAC) setpoint and adapts per-GPU frequency and micro-batch size during execution. ETCInfer builds compact physics-informed control models by calibrating GPU heat generation, chassis heat dissipation, CRAC power, and prefill/decode latency relations from telemetry. These models estimate hidden thermal states and time-to-throttle, enabling the scheduler to evaluate energy, temperature, and latency before applying an action. We formulate this joint setpoint--frequency--micro-batch control problem as a partially observable Markov decision process and design ETCAdapter, a learning-based controller that minimizes per-job energy under thermal safety and SLO constraints. We implement ETCInfer as a coordination layer over typical inference and cluster management stacks. Evaluation across real-trace simulation and validation experiments shows that ETCInfer reduces total job energy by up to 33.1%, thermal throttle exposure by up to 92.9%, and keeps SLO violation rates below 0.7% even at ambient temperatures up to $48^{\circ}\mathrm{C}$.
Research Article热管理与液冷
James Teague, Ashmita Rajmohan, Yannick Muehlhaeuser
Published 2026-09-16 · arXiv · Credibility S
Proposals for international agreements that limit frontier AI development depend on verification, and a central challenge is detecting undeclared compute facilities used to evade restrictions. Underwater data centers (UDCs) have been suggested as one such evasion vector, but their feasibility at frontier scale and their detectability have not been seriously assessed. We examine current UDC deployments, evaluate cons…
Abstract, interpretation and reference
Abstract
Proposals for international agreements that limit frontier AI development depend on verification, and a central challenge is detecting undeclared compute facilities used to evade restrictions. Underwater data centers (UDCs) have been suggested as one such evasion vector, but their feasibility at frontier scale and their detectability have not been seriously assessed. We examine current UDC deployments, evaluate construction and maintenance complexity relative to land-based facilities, and analyse the feasibility of a 100,000 H100-equivalent training run underwater. We find that power delivery and cooling are tractable, but interconnect and the hands-on maintenance that large training runs require are severe obstacles - surmountable only by a well-resourced state actor accepting large cost and schedule penalties, and only where concealment, rather than efficiency, is the objective. We then assess detectability through thermal, acoustic, optical and synthetic-aperture-radar (SAR) surveillance. Thermal detection of an operational pod is unlikely outside shallow, calm water; acoustic detection is marginally more effective, but faces limitations in attribution; and optical/SAR monitoring is most powerful during construction and maintenance, when the pressure-vessel fabrication base and the cable-laying fleet create distinctive signatures for AIS-tracking. We conclude that UDCs are a comparatively unlikely evasion route relative to underground or industrially disguised land-based facilities, but the residual risk is non-zero and warrants operationalising the detection modalities discussed.
Research Article算电协同
Yubo Song, Rui Kong, Takuro Umihara, Pooya Davari, Frede Blaabjerg, Subham Sahoo
Published 2026-09-10 · arXiv · Credibility S
The rapid growth of artificial intelligence (AI) computing is transforming data centers into large, dynamic electrical loads. Their deployment is primarily constrained by energy availability and grid-connection capacity, which is further aggravated by the ability of power-delivery architectures, control systems, and computing workloads to operate reliably during fast grid disturbances. This article presents a techno…
Abstract, interpretation and reference
Abstract
The rapid growth of artificial intelligence (AI) computing is transforming data centers into large, dynamic electrical loads. Their deployment is primarily constrained by energy availability and grid-connection capacity, which is further aggravated by the ability of power-delivery architectures, control systems, and computing workloads to operate reliably during fast grid disturbances. This article presents a technological perspective on AI data centers as grid-interactive computing systems. First, it reviews grid-integration bottlenecks, evolving connection policies, grid-code requirements, which has fostered new technological trends via spatio-temporal flexibility available through workload orchestration, cooling systems, on-site resources, and energy storage. Second, it maps the evolution of power-delivery architectures from medium-voltage grid interfaces to chip-level, discussing higher-voltage DC distribution, solid-state transformers, wide-bandgap devices, advanced chip-level power delivery, and liquid cooling. Third, it establishes a three-level stability framework spanning rack-level DC-bus dynamics, facility-level converter interactions, and system-level grid-coupled behavior. The framework connects dominant instability mechanisms, including constant power load effects, impedance interactions, forced oscillations, and operating-mode transitions, with suitable modeling, assessment, and mitigation approaches. Synthesizing these topics, this article highlights grid-to-chip co-design as a central requirement for scalable AI infrastructure, linking computing workloads, power-delivery systems, energy buffers, and grid operation.
Research Article热管理与液冷
Zixu Han, Peng Zhang
Published 2026-09-11 · arXiv · Credibility S
The rapid development of liquid-cooled data centers has imposed imperative demands on the performance of liquid cooling plate. The density-based topology optimization (TO) is an effective approach to resolving the growing thermal-hydraulic performance requirements of liquid cooling plate. However, existing TO methods can hardly optimize convective heat transfer directly which is the intrinsic heat transfer mechanism…
Abstract, interpretation and reference
Abstract
The rapid development of liquid-cooled data centers has imposed imperative demands on the performance of liquid cooling plate. The density-based topology optimization (TO) is an effective approach to resolving the growing thermal-hydraulic performance requirements of liquid cooling plate. However, existing TO methods can hardly optimize convective heat transfer directly which is the intrinsic heat transfer mechanism, due to the highly complex and evolving structural topologies, varying flow and temperature fields, making it extremely challenging to explicitly describe the heat transfer coefficient and heat transfer area during TO process. A convective heat transfer topology optimization (CTO) method is proposed in this study, where the iteratively evolving heat transfer coefficient is explicitly depicted by the field synergy theory in the thermal objective, and directly described by the velocity and temperature fields without relying on specific geometry. Combined with the explicit depiction of heat transfer area by the fractal geometry theory, a CTO framework is built for a direct optimization of convective heat transfer under both the laminar and turbulent flow conditions. The CTO tends to generate more hierarchical and directional structural topologies in optimization results, which is conducive to reducing low-velocity stagnation zones and improving flow direction in branched channels, achieving enhanced synergy and thermal-hydraulic performance in the optimized liquid cooling plates. Compared with the TO results without incorporation of field synergy theory, the CTO can reduce average temperature rise by 20% while improving the Nusselt number by 15% under laminar flow conditions, and reduce maximum temperature rise by 10.2% and pressure drop by 25% under turbulent flow conditions.
Research Article算电协同
Bojun Du, Hongyang Jia, Tonghui Li, Qingchun Hou, Ze Wang, Ershun Du, Ning Zhang
Published 2026-09-09 · arXiv · Credibility S
AI data centers may face power supply shortages during certain periods, requiring operators to shift large language model (LLM) inference workloads spatially to maintain service rates. However, existing workload-shifting methods typically assume that any data center with sufficient computing resources can immediately serve shifted requests, which may lead to infeasible transfers and unserved demand. This letter prop…
Abstract, interpretation and reference
Abstract
AI data centers may face power supply shortages during certain periods, requiring operators to shift large language model (LLM) inference workloads spatially to maintain service rates. However, existing workload-shifting methods typically assume that any data center with sufficient computing resources can immediately serve shifted requests, which may lead to infeasible transfers and unserved demand. This letter proposes model commitment (MC), a mixed-integer linear programming framework that jointly schedules model deployment and cross-site request routing under power constraints and electricity-price signals. First, MC formulates the intertemporal coupling introduced by model replica loading. Second, it translates prefill and decode latency requirements into the amount of demand that each replica can serve. Case studies based on real-world data show that MC enables AI data center operators to achieve a 100% service rate under time-varying grid conditions and reduce total operating cost by 29.0%.
Research Article热管理与液冷
Seung Hyeon Ham, Yu Min Kim, In Hyeok Choi, Jeong Woo Han
Published 2026-08-26 · arXiv · Credibility S
The rapid growth of generative AI has intensified the need for efficient heat dissipation in large-scale data centers. To control heat flow, thermal metamaterials with layered structures have been widely used, which impart the anisotropic properties of thermal conductivities. However, the conventional effective medium approximation (EMA) often fails to provide accurate predictions in systems with a high thermal cond…
Abstract, interpretation and reference
Abstract
The rapid growth of generative AI has intensified the need for efficient heat dissipation in large-scale data centers. To control heat flow, thermal metamaterials with layered structures have been widely used, which impart the anisotropic properties of thermal conductivities. However, the conventional effective medium approximation (EMA) often fails to provide accurate predictions in systems with a high thermal conductivity contrast between adjacent layers embedded in a background medium. Here, we generalize the EMA by introducing two corrective coefficients that extend its validity to regimes where the conventional EMA was previously inapplicable, i.e., high-contrast thermal metamaterials with the background medium. Notably, one of these coefficients that we proposed has the same mathematical form as the Fresnel reflection coefficient in optics. This allows us to interpret the "reflection-like" behavior of heat flow as it penetrates adjacent layers with high thermal contrast. Our findings suggest that heat diffusion, traditionally viewed as a purely dissipative process, can be understood intuitively through the framework of ray optics.
Research Article算电协同
Arya Joshi, Hamed Haggi, Chinmay Morankar
Published 2026-08-31 · arXiv · Credibility S
The accelerated growth in data center projects has introduced a demand-driven bottleneck throughout power grids and contributed to a substantial increase in carbon emissions. These concerns are fueling discussions on methods to use existing energy assets to drive operational efficiency. To this end, this paper explores the benefits of Vehicle-to-Building (V2B) applications to support peak shaving of data center cool…
Abstract, interpretation and reference
Abstract
The accelerated growth in data center projects has introduced a demand-driven bottleneck throughout power grids and contributed to a substantial increase in carbon emissions. These concerns are fueling discussions on methods to use existing energy assets to drive operational efficiency. To this end, this paper explores the benefits of Vehicle-to-Building (V2B) applications to support peak shaving of data center cooling loads. Initially, a literature review was conducted considering V2B constraints and optimization methods including SoC limitations, EV participation, tariffs, and building loads. This analysis was then used to develop a conceptual case study of a 10 MW data center in Loudoun County, VA by simulating a temperature-dependent load profile and adjusting the V2B participation of 40 commercial and passenger EVs. Simulation results indicate that, depending on seasonal variations in cooling load demands, strategic deployment of V2B assets between 12-5pm can offset gross cooling loads by 13-36%.
Research ArticleAI 运维优化
Slava G. Turyshev
Published 2026-08-27 · arXiv · Credibility S
Megawatt-class orbital data centers require continuous maintenance, replacement, inventory, and service capacity in addition to spacecraft power/thermal systems. We formulate an analytical lifecycle framework for permanent/transient failures, modular orbital replacement units, robotic servicing, spare inventory, scheduled technology refresh, correlated faults, cybersecurity, optional human support. The model combine…
Abstract, interpretation and reference
Abstract
Megawatt-class orbital data centers require continuous maintenance, replacement, inventory, and service capacity in addition to spacecraft power/thermal systems. We formulate an analytical lifecycle framework for permanent/transient failures, modular orbital replacement units, robotic servicing, spare inventory, scheduled technology refresh, correlated faults, cybersecurity, optional human support. The model combines nonhomogeneous component hazards, capacity-weighted availability, multiclass robotic-service capacity, Poisson base-stock inventory, replacement-flow accounting, human-support break-even relations. For a 1 MW cluster with 10 active 100 kW nodes, 1 reserve node, ~200 5 kW compute cartridges, low, nominal, high deployed-mass allocations span ~50-75 kg/kW. Assumptions yield 70.2 random or life-limited interventions and 323-349 planned refresh operations/(MW-year), for a total of 393-419 standardized operations/(MW-year). Analysis gives a first-generation logistics of 5.3-9.0 t/(MW-year), with a nominal case of ~ 6.6 t/(MW year), 560-700 productive robot-hors/(MW-year). Planned refresh exceeds random replacement under the stated component populations, hazards, 3-15-year intervals. At ~400 standardized operations/(MW-year), the post-internal-recovery exception probability is <$10^{-3}$, with an objective near $10^{-4}$ at large scale; terminal non-recovery $p_U$ requires a smaller mission-level allocation. The target catastrophic-loss hazard for a 100 kW node is 0.01-0.03 1/yr. Parametric workload and cost cases place contingency visits at 10s of MWs, periodic campaigns at 10-100s of MWs, dedicated personnel at several 100 MWs to GWs. The reference first deployment is uncrewed, autonomously fault-managed, robotically maintainable, supported by specific inventory based on a 6-month replenishment horizon, compatible with later human access without permanent habitation.
Research Article算电协同
Hassan Zahid Butt, Rida Fatima, Xingpeng Li
Published 2026-08-29 · arXiv · Credibility S
Securing grid interconnection capacity has become a bottleneck for AI data center projects and can take longer than constructing the facilities themselves. This mismatch can delay deployment for years, making early interconnection planning essential. This paper develops ICP-AI, an interconnection capacity planning framework from a data center developer's perspective. The framework minimizes grid import capacity unde…
Abstract, interpretation and reference
Abstract
Securing grid interconnection capacity has become a bottleneck for AI data center projects and can take longer than constructing the facilities themselves. This mismatch can delay deployment for years, making early interconnection planning essential. This paper develops ICP-AI, an interconnection capacity planning framework from a data center developer's perspective. The framework minimizes grid import capacity under a prescribed onsite investment budget while jointly sizing photovoltaic (PV) and battery energy storage system (BESS) resources and scheduling deadline constrained workload flexibility. A secondary refinement fixes the minimum grid capacity and selects the minimum-investment PV-BESS portfolio among solutions that achieve that capacity. The framework is evaluated using monthly composite stress profiles across varying temporal assumptions, load shapes, flexible load fractions, and deferral windows. Results show that interconnection capacity reduction depends strongly on the planning environment: at a $100M budget, it is about 6% for the high load factor baseline, exceeds 10% under monthly average solar availability, and reaches 13.3% for a more diurnal load. At a $10M budget, 5% flexible load with a 1 h workload deferral window reduces BESS capacity from 15.30 to 4.87 MWh while increasing capacity reduction from 4.43% to 4.84%. To test sensitivity to temporal compression, the model is also solved over the full 8,760 h chronology, which preserves the main capacity and flexibility trends. Overall, ICP-AI quantifies the interconnection capacity and infrastructure substitution value of workload flexibility, providing an investment-interconnection frontier to support capital allocation and early project planning in constrained grid environments.
Research Article算电协同
Jae-Kyeong Kim
Published 2026-08-31 · arXiv · Credibility S
The rapid expansion of large-scale artificial intelligence (AI) data centers is adding substantial, concentrated, and rapidly varying loads to transmission-constrained power systems. Although such load variations are generally regarded as operational challenges, this paper presents an alternative perspective in which the upward load flexibility of AI data centers could be coordinated for transient-stability support.…
Abstract, interpretation and reference
Abstract
The rapid expansion of large-scale artificial intelligence (AI) data centers is adding substantial, concentrated, and rapidly varying loads to transmission-constrained power systems. Although such load variations are generally regarded as operational challenges, this paper presents an alternative perspective in which the upward load flexibility of AI data centers could be coordinated for transient-stability support. To this end, this paper proposes training-induced load surge (TILS), a fast demand-side strategy that initiates or resumes flexible AI training workloads after fault clearing to increase active-power demand at electrically effective locations. The resulting load increase allows accelerating generators to supply additional electrical power, thereby reducing the accelerating-power imbalance and limiting the first-swing rotor-angle excursion. The underlying mechanism is first clarified in a single-machine infinite-bus (SMIB) system and then evaluated in the IEEE 39-bus system and a large-scale Korean power system. Results across all three systems demonstrate that TILS can increase the transient-stability-constrained generation limit. Larger responses, earlier activation, and siting at buses with a stronger electrical influence on the critical generators provide greater generation-limit increases. These results suggest that the upward load-response capability of AI data centers can provide complementary transient-stability support when sufficient electrical headroom, flexible workloads, and reliable grid-triggered activation are available.
Research Article算电协同
Pengyu Ren, Wei Sun, Fei Teng
Published 2026-09-02 · arXiv · Credibility S
Large data centers are emerging as concentrated, power-electronic grid loads whose abrupt disconnection or transfer to on-site backup supply during voltage disturbances can remove large demand from the power system, and may create a system-level stability problem. Their interconnection feasibility therefore depends not only on steady-state thermal and voltage limits, but also on whether internal power-conditioning s…
Abstract, interpretation and reference
Abstract
Large data centers are emerging as concentrated, power-electronic grid loads whose abrupt disconnection or transfer to on-site backup supply during voltage disturbances can remove large demand from the power system, and may create a system-level stability problem. Their interconnection feasibility therefore depends not only on steady-state thermal and voltage limits, but also on whether internal power-conditioning systems can maintain IT service while limiting customer-initiated load reduction. This paper presents a voltage ride-through (VRT)-aware data center and grid co-planning framework that couples transmission-level fault simulation with an internal data center ride-through model. Python-based dynamic simulations generate point-of-interconnection (POI) voltage trajectories under selected network faults, and the resulting waveforms drive an internal model incorporating IT and cooling-load dynamics, DC-link, Uninterruptible Power Supply (UPS) response, and converter apparent power limits. The IEEE 118-bus case study shows that internal VRT capability can become a binding interconnection constraint: steady-state planning alone can overestimate feasible data center capacity, whereas increased UPS converter headroom progressively restores hosting capacity. Under the reduced-order response models studied, the grid-forming mode provides greater ride-through margin than the current-limited grid-following mode under the same network fault conditions. The results further show that VRT constraints can materially change both the total hosting capacity of data centers and its spatial allocation across candidate interconnection buses.
Research Article算电协同
Xin Chen
Published 2026-09-03 · arXiv · Credibility S
To facilitate the grid-friendly integration of highly variable AI data center loads, this paper proposes a grid-mode-aware model predictive control (G-MPC) framework for managing a hybrid energy storage system (HESS) to smooth grid-side power demand. The framework optimally coordinates a battery energy storage system (BESS) and a supercapacitor (SC) by solving a multi-step optimization problem in a receding-horizon …
Abstract, interpretation and reference
Abstract
To facilitate the grid-friendly integration of highly variable AI data center loads, this paper proposes a grid-mode-aware model predictive control (G-MPC) framework for managing a hybrid energy storage system (HESS) to smooth grid-side power demand. The framework optimally coordinates a battery energy storage system (BESS) and a supercapacitor (SC) by solving a multi-step optimization problem in a receding-horizon manner. In particular, band-pass filter dynamics are directly embedded in the G-MPC formulation to extract and suppress grid-side power components associated with vulnerable grid oscillatory modes, thus mitigating load-induced grid oscillations. The resulting G-MPC optimization jointly minimizes violations of grid-side power-envelope, ramp-rate, and modal-power requirements and the degradation and power-ramping costs of the BESS and SC, while satisfying power limits, state-of-charge limits, and other operational constraints. To enable real-time implementation, a fix-and-re-optimize algorithm is developed to solve each G-MPC problem efficiently while preventing simultaneous charging and discharging. Extensive simulations demonstrate the effectiveness, flexibility, and computational efficiency of the proposed framework. The results also highlight the importance of explicitly suppressing power components associated with vulnerable grid modes, rather than merely reducing overall load variations, to effectively mitigate grid oscillations.