智算中心论文专站

AIDC Research Papers

Liquid Cooling AI Data Center Power & Thermal Systems
Current Issue

Volume 2026 · Issue 07-21

按期刊卷期页方式整理本期论文。每条仅使用日报已列出的可追溯公开来源,不新增未经核验事实。

Research Article算电协同

A Predict-then-Schedule framework for Power Distribution Networks with AI Data Centers

Siqi Yan, Jiebao Zhang, Xi Yao, Juan Huang, Ye Shi

Published 2026-07-20 · arXiv · Credibility S

The surge of GPU-intensive workloads in artificial intelligence (AI) data centers drives massive energy demands, leading to soaring costs and significant stress on local power distribution networks. Coordinating delay-tolerant workload scheduling with power grid conditions via precise workload prediction can mitigate these issues. However, a critical gap remains in conventional approaches, i.e., minimizing predictio…

Abstract, interpretation and reference

Abstract

The surge of GPU-intensive workloads in artificial intelligence (AI) data centers drives massive energy demands, leading to soaring costs and significant stress on local power distribution networks. Coordinating delay-tolerant workload scheduling with power grid conditions via precise workload prediction can mitigate these issues. However, a critical gap remains in conventional approaches, i.e., minimizing prediction error does not necessarily lead to minimized downstream operational loss. Hence, this paper proposes an end-to-end Predict-Then-Schedule (PTS) framework that integrates upstream workload prediction with downstream scheduling optimization. By leveraging differentiable convex optimization, the PTS framework maps input features directly to optimal scheduling and enables gradient-based training. Furthermore, to respect the data center's capacity, a workload over-shifted loss combining electricity cost with a penalty for load-shedding is introduced to evaluate scheduling quality. Experiments demonstrate that the proposed framework significantly reduces operational cost and enhances system security compared to the conventional two-stage baseline.

Full text 中文海报
算电协同 论文图示
Research Article热管理与液冷

Assessing Risks of Hydro-Generator Shaft Fatigue from Data Center Load Oscillations

Kaustav Chatterjee, Meghana Ramesh, Shuchismita Biswas, Brett A. Ross, Antos C. Varghese, Sameer Nekkalapu, Slaven Kincic

Published 2026-07-15 · arXiv · Credibility S

Large AI data center loads can introduce persistent sub-synchronous active-power oscillations that may impact nearby generators by exciting torsional modes and increasing shaft stress. This paper presents a model-based framework for evaluating hydro-generator shaft fatigue risk under oscillatory loading. An electromagnetic transient simulation model is developed using a two-mass turbine-generator shaft representatio…

Abstract, interpretation and reference

Abstract

Large AI data center loads can introduce persistent sub-synchronous active-power oscillations that may impact nearby generators by exciting torsional modes and increasing shaft stress. This paper presents a model-based framework for evaluating hydro-generator shaft fatigue risk under oscillatory loading. An electromagnetic transient simulation model is developed using a two-mass turbine-generator shaft representation with parameters from real-world generation units and a configurable AI data center load. The risk assessment is performed in two stages. First, a network transfer function quantifies the propagation of load oscillations from the data center point of interconnection to the hydro-generator terminal. A plant transfer function then characterizes the resulting shaft torque amplification. A frequency-scan approach identifies resonance regions and evaluates torque amplification at individual forcing frequencies. Parametric studies show that amplification is strongly affected by generator-to-turbine inertia ratio and torsional damping. Lower inertia ratios shift torsional modes to lower frequencies and increase amplification, indicating that some Kaplan-type units may be more susceptible than comparable Francis or Pelton units. Reduced damping further increases resonant response and fatigue exposure. A simplified fatigue assessment based on S--N curves and the Goodman diagram relates simulated torque response to mechanical integrity. The resulting Goodman safety factor provides a practical metric for evaluating the impact of persistent AI data center oscillations on hydro-generator service life and supports interconnection studies, oscillation limits, and plant-level monitoring strategies.

Full text 中文海报
热管理与液冷 论文图示
Research Article热管理与液冷

The Cost and Network Limits of Space-Based AI Compute

Kees van Berkel

Published 2026-07-15 · arXiv · Credibility S

This paper evaluates whether large-scale AI data centers deployed in low-Earth orbit (LEO) could become a cost-effective alternative to terrestrial facilities. The analysis compares orbital and ground-based systems across launch cost, power generation, cooling, radiation exposure, and atmospheric reentry, as well as compute-network performance. A key distinction is the shift from terrestrial Clos networks to space-b…

Abstract, interpretation and reference

Abstract

This paper evaluates whether large-scale AI data centers deployed in low-Earth orbit (LEO) could become a cost-effective alternative to terrestrial facilities. The analysis compares orbital and ground-based systems across launch cost, power generation, cooling, radiation exposure, and atmospheric reentry, as well as compute-network performance. A key distinction is the shift from terrestrial Clos networks to space-based mesh networks using laser inter-satellite links. Using bisection bandwidth, bisection intensity, and roofline-style models, we show that while LEO-based inference may be feasible, training frontier-scale LLMs in orbit is unlikely to be competitive with terrestrial data centers.

Full text 中文海报
热管理与液冷 论文图示
Research Article算电协同

A Phased Development Framework Enabling Islanded Operation of Sustainable AI Data Centers With Onsite Grid-Following and Grid-Forming Energy Architectures

Soham Ghosh, Nabil Mohammed, Mohammad Ashraf Hossain Sadi

Published 2026-07-19 · arXiv · Credibility S

As hyperscale and colocation AI data centers continue to expand, the electric grid is increasingly required to support large, concentrated loads, with individual facilities ranging from 500 MW to 2 GW. Current projections estimate that approximately 50 GW of AI data center capacity will require grid connectivity in the United States by 2030. While prior research has extensively examined the environmental and operati…

Abstract, interpretation and reference

Abstract

As hyperscale and colocation AI data centers continue to expand, the electric grid is increasingly required to support large, concentrated loads, with individual facilities ranging from 500 MW to 2 GW. Current projections estimate that approximately 50 GW of AI data center capacity will require grid connectivity in the United States by 2030. While prior research has extensively examined the environmental and operational impacts of AI data centers, as well as their potential role as grid-interactive assets, limited attention has been given to the challenges associated with their scalable deployment through engineering, procurement, and construction (EPC) processes. This manuscript addresses this gap by proposing a phased development framework for AI data center expansion. The approach is designed to enable developers to meet aggressive time-to-market objectives while navigating multi-year constraints associated with interconnection approvals and lead times associated with the procurement of component equipment. A modular construction architecture is presented, along with a detailed analysis of integrated energy systems and the role of hybrid on-site generation in supporting incremental capacity growth. Electromagnetic transient simulations (EMT) are used to evaluate system performance, demonstrating that a combination of on-site natural gas generation and grid-forming energy storage can reliably support data center operations during early and intermediate deployment phases. The study further examines the transition to full grid interconnection, including the capability of the data center to operate in islanded mode during grid disturbances. Finally, the manuscript compares grid-forming control strategies for system reconnection and restoration under varying conditions.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

The Environmental Cost of Digital Sovereignty: Water, Energy, and Emissions Impacts of Sovereign AI Infrastructure in the Global South

Muntaser Syed, Marius C. Silaghi, Sheikh Abujar, Sharun Akter Khushbu, Amal El Ahmad

Published 2026-07-15 · arXiv · Credibility S

Sovereign AI has become a strategic priority across the Global South, with over \$200 billion in state-led commitments announced between 2024 and 2026. Yet the physical infrastructure that compute sovereignty demands, above all data centers, imposes water, energy, and carbon costs that fall hardest on countries least equipped to absorb them. This paper presents a comparative environmental stress analysis across four…

Abstract, interpretation and reference

Abstract

Sovereign AI has become a strategic priority across the Global South, with over \$200 billion in state-led commitments announced between 2024 and 2026. Yet the physical infrastructure that compute sovereignty demands, above all data centers, imposes water, energy, and carbon costs that fall hardest on countries least equipped to absorb them. This paper presents a comparative environmental stress analysis across four cases: the United Arab Emirates, Bangladesh, India, and Africa (with a focus on Kenya). Using publicly available water stress data, grid carbon intensity factors, and GPU power specifications, we model the water consumption, energy demand, and carbon emissions of hypothetical sovereign AI deployments under multiple cooling technology scenarios. We find that a 1,024-GPU cluster using evaporative cooling in the UAE would consume over 30 million liters of water annually in a country classified as ``extremely high'' water stress. In Bangladesh, sovereign AI policy documents call for centralized GPU procurement but do not address where to site data centers in a country where more than a fifth of the land floods in an average year and the power grid struggles to deliver reliable supply. We identify a sovereignty-sustainability trilemma in which no country can simultaneously maximize AI sovereignty, minimize environmental impact, and maintain affordable resource access for citizens. We propose design principles for environmentally responsible sovereign AI, including mandatory water usage effectiveness reporting, climate-vulnerability siting assessments, and a preference for frugal small language models over frontier pre-training in resource-constrained settings.

Full text 中文海报
算电协同 论文图示
Research Article热管理与液冷

Classical Reversible Computation by Quantum Coherence

Daniel Loss

Published 2026-07-07 · arXiv · Credibility S

Rising energy demand from data-center and AI applications has renewed interest in reversible computation, where logic need not dissipate heat at every step if information is uncomputed. Implementations have so far been classical: adiabatic CMOS reduces dissipation by slowing charge motion but is still limited by the threshold physics of transistors. Here we propose classical reversible logic implemented by coherent …

Abstract, interpretation and reference

Abstract

Rising energy demand from data-center and AI applications has renewed interest in reversible computation, where logic need not dissipate heat at every step if information is uncomputed. Implementations have so far been classical: adiabatic CMOS reduces dissipation by slowing charge motion but is still limited by the threshold physics of transistors. Here we propose classical reversible logic implemented by coherent spin dynamics in a spin quantum-dot array, with inputs and outputs in classical basis states and no algorithmic use of superposition. The same spin stores, transports, and computes, with unitary rotation replacing irreversible switching. The universal building block is an iToffoli gate driven by DC voltage pulses and anisotropic exchange in Ge/Si hole spins. Simulations with experimental parameters reproduce the Toffoli truth table and yield a testable error landscape. Because shuttling transports the bit without measurement, logic and data movement remain reversible until readout. Millivolt pulses on femtofarad gates yield a gate energy below the 4 K Landauer scale, about five (eight) orders of magnitude below a room-temperature CMOS Toffoli with (without) 4 K cooling overhead. The same semiconductor hardware is therefore dual-use, supporting quantum algorithms when superposition is used and classical reversible logic otherwise.

Full text 中文海报
热管理与液冷 论文图示
Research Article算电协同

A Hierarchical Semi-Markov Load Model for AI Data Centers Coupling Job Scheduling with Bulk-Synchronous-Parallel Power Dynamics

Chandan Chaudhary, Atri Bera, Cody Newlun, Mohammed Ben-Idris, Joydeep Mitra

Published 2026-07-13 · arXiv · Credibility S

AI data centers are emerging as a dominant new load class with their power dynamics fundamentally from conventional industrial loads. Inside a training job, the bulk-synchronous-parallel algorithm moves each node through compute, sync, and checkpoint steps, which swings power between full load and near idle within seconds. Across the whole facility, jobs arrive, take blocks of nodes for hours to days, then leave, so…

Abstract, interpretation and reference

Abstract

AI data centers are emerging as a dominant new load class with their power dynamics fundamentally from conventional industrial loads. Inside a training job, the bulk-synchronous-parallel algorithm moves each node through compute, sync, and checkpoint steps, which swings power between full load and near idle within seconds. Across the whole facility, jobs arrive, take blocks of nodes for hours to days, then leave, so the number of busy nodes changes daily, weekly, and yearly. This slower shift drives facility-wide swings and the peak demand that sets the size of the grid link. A model that looks only at within-job behavior, and treats the facility as a fixed set of busy nodes, smooths out these swings and misses the true peak-to-average ratio. This paper develops a hierarchical semi-Markov Data-Center (HSM-DC) load model that couples two layers across two timescales. A job-scheduling layer creates jobs through a non-homogeneous compound-Poisson process shaped by daily, weekly, and seasonal patterns, gives each job a heavy-tailed node count and length, and places jobs on a fixed pool of nodes on a first-come basis. A within-job layer moves each busy node through a five-state semi-Markov chain for the BSP steps, with state-based Ornstein-Uhlenbeck noise. Facility power comes from this changing node count and the per-node power, set to match measured node data and the facility's straight-line power-versus-load curve. Configured to the reference facility at the same scale, the model matches mean power, its spread, and the peak-to-average ratio across load levels, with fit scores of 0.9997, 0.92, and 0.82. It also matches the share of queued jobs to within one point at high load. Facility-wide swings and peak demand come from how jobs arrive and get scheduled, so grid planning must model that process, not just scale up a single node's power curve.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

Large-Load Demand Flexibility as Virtual Storage

Chandan Chaudhary, Mohammed Ben-Idris, Joydeep Mitra

Published 2026-07-06 · arXiv · Credibility S

Water electrolysis plants, hyperscale data centers, and aluminum potlines represent gigawatts of demand-side flexibility for bulk power system balancing, operational planning, and procurement services. Such loads are scheduled through per-interval power bounds and horizon energy windows, whereas co-located battery energy storage systems (BESS) operate under state-of-charge dynamics. The two formulations share no com…

Abstract, interpretation and reference

Abstract

Water electrolysis plants, hyperscale data centers, and aluminum potlines represent gigawatts of demand-side flexibility for bulk power system balancing, operational planning, and procurement services. Such loads are scheduled through per-interval power bounds and horizon energy windows, whereas co-located battery energy storage systems (BESS) operate under state-of-charge dynamics. The two formulations share no common mathematical structure, and the joint procurement value of co-located loads and storage goes unrealized as a result. This paper establishes the connection between the two formulations through a virtual storage (VS) equivalence. Every feasible large-load trajectory under power-bound and energy-window constraints is a valid charge trajectory of a VS device that operates at unity accounting efficiency in the grid power balance. Production and service-level costs lie outside this abstraction and enter the dispatch through curtailment opportunity costs. For a portfolio co-located with a BESS, aggregation reduces the constraint count from O(NT) to O(T) and yields a co-dispatch price for both resources. Validation on the IEEE RTS-GMLC with three representative load classes shows that virtual storage delivers the dominant share of joint procurement savings. In the tested case, savings are additive because the two resources dispatch to non-overlapping intervals, and the curtailment shadow price tracks the peak-price band onset rather than the daily peak price.

Full text 中文海报
算电协同 论文图示
Research Article芯片与算力

Hierarchical Multi-Agent Reinforcement Learning for Carbon-Aware AI Data Centers in Power Distribution Systems

Hyunsoo Lee, Panggah Prabawa, Dae-Hyun Choi, Joongheon Kim

Published 2026-07-03 · arXiv · Credibility S

Eco-friendly energy management for artificial intelligence data centers (AIDCs) is crucial because of the significant increase in energy consumption-induced carbon emissions from AIDCs resulting from the rapid expansion of AI applications. This paper proposes a hierarchical carbon-aware multi-agent reinforcement learning (CA-MARL) framework for robust and efficient operations of AIDCs under uncertainties while ensur…

Abstract, interpretation and reference

Abstract

Eco-friendly energy management for artificial intelligence data centers (AIDCs) is crucial because of the significant increase in energy consumption-induced carbon emissions from AIDCs resulting from the rapid expansion of AI applications. This paper proposes a hierarchical carbon-aware multi-agent reinforcement learning (CA-MARL) framework for robust and efficient operations of AIDCs under uncertainties while ensuring low-carbon operation of power distribution systems. The framework comprises a workload manager (WM) agent and multiple local AIDC agents trained using a multi-agent transformer method, corresponding to a global AIDC aggregator and a local AIDC operator, respectively. Leveraging AIDC operation data along with nodal carbon intensity (NCI) calculated from the carbon emission flow-integrated distribution system operator problem, the WM agent spatially allocates AI training and inference jobs among all AIDCs. Based on the jobs allocated from the WM agent and NCI information, each AIDC agent schedules economical and eco-friendly operations of the AIDC by performing the following tasks: i) temporal shifting of training jobs, ii) spatial allocation of training graphics processing unit (GPU) blocks and inference GPUs within the AIDC, and iii) control of the supply air temperature of the cooling system. The effectiveness of the proposed framework was assessed using an IEEE 33-node power distribution system.

Full text 中文海报
芯片与算力 论文图示
Research Article算电协同

How Data Shapes RoPE Frequency Usage: From Positional Scale Matching to Length Generalization

Xinyi Wu, Siyuan Liu, Ali Jadbabaie

Published 2026-07-08 · arXiv · Credibility S

Rotary Position Embeddings (RoPE) provide transformers with a fixed grid of positional frequencies, yet trained models use these frequencies highly non-uniformly. We study what determines this frequency usage and propose a data-centered explanation: RoPE frequencies are selected to match the relative-distance structure of the training data. Viewing each frequency as a positional lens, we formalize a field-resolution…

Abstract, interpretation and reference

Abstract

Rotary Position Embeddings (RoPE) provide transformers with a fixed grid of positional frequencies, yet trained models use these frequencies highly non-uniformly. We study what determines this frequency usage and propose a data-centered explanation: RoPE frequencies are selected to match the relative-distance structure of the training data. Viewing each frequency as a positional lens, we formalize a field-resolution tradeoff and show that, for a data-induced dependency profile of width $W$, the optimal frequency scales as $1/W$. This frequency-matching principle explains controlled observations on synthetic and text-based data, and suggests that the mid-low frequency bands observed in language models arise from the multi-scale dependency structure of natural language. We further connect frequency selection to position-interpolation-based length generalization: scaling frequencies down expands the effective field while reducing resolution. This helps when longer-context dependencies are approximate dilations of those seen during training, but can fail when relevant dependencies do not scale with context length. Empirically, we show that natural language exhibits approximate self-similarity across positional scales, explaining why test-time frequency scaling can support long-context generalization. Overall, our results identify a data-driven mechanism behind emergent RoPE frequency usage and show that long-context generalization depends on two forms of scale matching: between learned frequencies and training-time dependencies, and between frequency scaling and how those dependencies extend to longer contexts.

Full text 中文海报
算电协同 论文图示
Research ArticleAI 运维优化

Storage as a Transmission Asset (SATA) for Large-Load Congestion Relief

Abanish Tiwari, Chandan Chaudhary, Yansong Pei, Mohammed Ben-Idris, Joydeep Mitra

Published 2026-07-05 · arXiv · Credibility S

Hyperscale data centers and other large concentrated loads can impose substantial new demand on existing transmission networks. If import corridors lack sufficient transfer capability, operators may need to curtail load, delay interconnection, or reinforce the network to maintain reliable service. An energy storage system (ESS) deployed as a storage-as-transmission asset (SATA) offers a non-wires alternative by prov…

Abstract, interpretation and reference

Abstract

Hyperscale data centers and other large concentrated loads can impose substantial new demand on existing transmission networks. If import corridors lack sufficient transfer capability, operators may need to curtail load, delay interconnection, or reinforce the network to maintain reliable service. An energy storage system (ESS) deployed as a storage-as-transmission asset (SATA) offers a non-wires alternative by providing operator-directed support to constrained import corridors. However, the operating-level reliability value of SATA dispatch remains insufficiently quantified. This paper evaluates operator-directed SATA using a day-ahead DC optimal power flow that co-optimizes generation, ESS dispatch, and load curtailment across Monte Carlo scenarios of demand and generator availability. Operating reliability is assessed using expected energy not served (EENS), loss-of-load hours (LOLH), and the conditional value at risk (CVaR) of daily unserved energy. Congestion-price and flow-sensitivity metrics are used to identify the limiting corridor and storage location. The interconnection is then screened to determine whether SATA is suitable, reinforcement is required, or storage would provide little transmission value. Results show that operator-directed SATA reduces average unserved energy, loss-of-load exposure, and tail risk compared with deploying the same ESS for pure arbitrage. These results demonstrate that the operating designation of storage is a primary driver of its transmission value.

Full text 中文海报
AI 运维优化 论文图示
Research Article热管理与液冷

AI-Driven Thermal Mapping and Management in 3D Integrated Photonic Circuits

Liton Kumar Biswas, Katayoon Yahyaei, Shajib Ghosh, M Shafkat M Khan, Himanandhan Reddy Kottur, Rayhane Ghane-Motlagh, Mahdi Nikdast, Navid Asadizanjani

Published 2026-06-25 · arXiv · Credibility S

Photonic Integrated Circuits (PICs) are advancing high-performance computing, data centers, and sensing, yet three-dimensional (3D) PICs introduce critical thermal management challenges due to high-density bonding and heterogeneous materials. Traditional methods like thermal microscopes and in-package sensors yield sparse data, limiting full thermal profile visibility. This paper presents a dual-method solution comb…

Abstract, interpretation and reference

Abstract

Photonic Integrated Circuits (PICs) are advancing high-performance computing, data centers, and sensing, yet three-dimensional (3D) PICs introduce critical thermal management challenges due to high-density bonding and heterogeneous materials. Traditional methods like thermal microscopes and in-package sensors yield sparse data, limiting full thermal profile visibility. This paper presents a dual-method solution combining an AI-driven thermal modeling framework with a design-based heuristic approach. The AI method integrates sparse sensor data with design layer and density information to predict multilayer temperature variations, while the heuristic approach uses localized material properties, design layout, component geometries, and sensor coordinates to refine thermal estimations in specific regions. A 2D thermal map of a 3D PIC is generated by interpolating sensor data and adjusting for local thermal resistivity using comparative analysis between design regions. The heuristic method complements the AI model, improving estimation accuracy without extensive training data. Together, these methods offer a scalable, accurate solution for real-time thermal mapping and design-time simulation, enabling reliable thermal management in next-generation 3D photonic systems.

Full text 中文海报
热管理与液冷 论文图示
Research Article芯片与算力

WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs

Mauricio Fadel Argerich, Jonathan Fürst, Marta Patiño-Martínez

Published 2026-07-02 · arXiv · Credibility S

Large Language Model (LLM) inference workloads are a rapidly growing contributor to data center energy consumption. Optimizing these deployments requires matching specific LLMs to the most efficient GPUs, but operators currently lack the tools to do so without exhaustively profiling each combination. While some predictive models exist, they still require profiling data and struggle to generalize to hardware unseen d…

Abstract, interpretation and reference

Abstract

Large Language Model (LLM) inference workloads are a rapidly growing contributor to data center energy consumption. Optimizing these deployments requires matching specific LLMs to the most efficient GPUs, but operators currently lack the tools to do so without exhaustively profiling each combination. While some predictive models exist, they still require profiling data and struggle to generalize to hardware unseen during training. To address this, we introduce \textit{WattGPU}, featuring two predictive models for mean GPU power draw and Inter-Token Latency (ITL). Our approach leverages only publicly available LLM metadata and GPU specifications, eliminating the need for hardware access or profiling while enabling generalization to unseen NVIDIA server-grade GPUs and LLMs. We evaluate our models using rigorous leave-one-GPU-out and leave-one-LLM-out cross-validation on a dataset of 42 open-source LLMs (0.1B--27B parameters) and 8 GPUs under both offline and server scenarios. The mean power draw model achieves a median absolute percentage error of $\leq3.4\%$ for offline and $\leq13.5\%$ for server scenarios on unseen GPUs, while the latency model achieves $\leq8.5\%$ in server mode, both maintaining strong GPU ranking correlations for server scenarios (Kendall $τ\geq0.76$). Compared to standard physically grounded baselines -- Load-Scaled Thermal Design Power (TDP) for power draw and roofline for latency -- our models reduce median absolute percentage error by approximately 4$\times$ on unseen LLM-GPU combinations for server scenarios or approximately 2$\times$ for completely unseen GPUs. WattGPU's data and code are publicly available at https://github.com/maufadel/wattgpu.

Full text 中文海报
芯片与算力 论文图示
Research Article算电协同

Grid-Interactive Thermal Management of AI Data Centers via Contextual Distributionally Robust Optimization

Jiachen Shen, Jian Shi, Yijie Yang, Chenye Wu, Dan Wang, Ju Bin Song, Zhu Han

Published 2026-06-30 · arXiv · Credibility S

Thermal management in AI data centers is increasingly challenged by bursty workloads and uncertain heat generation. To prevent thermal violations, existing cooling strategies either enforce conservative, rigid bounds that severely limit grid responsiveness, or rely on forecast-driven controllers that perform poorly under AI workload uncertainty and distribution shifts. To overcome the above challenges, this paper pr…

Abstract, interpretation and reference

Abstract

Thermal management in AI data centers is increasingly challenged by bursty workloads and uncertain heat generation. To prevent thermal violations, existing cooling strategies either enforce conservative, rigid bounds that severely limit grid responsiveness, or rely on forecast-driven controllers that perform poorly under AI workload uncertainty and distribution shifts. To overcome the above challenges, this paper proposes a Contextual Distributionally Robust Optimization (CDRO) framework for grid-interactive cooling control. Unlike standard DRO with fixed ambiguity sets, the proposed approach dynamically adapts the Wasserstein radius using real-time AI and grid context. This safely shrinks uncertainty bounds during stable regimes, unlocking deep demand-side flexibility. Theoretically, we formulate the control as an infinite-dimensional inf-sup problem, derive an exact tractable reformulation for the Wasserstein worst-case expected-cost term, and then derive a tractable conservative deterministic counterpart for the Distributionally Robust Conditional Value at Risk (DR-CVaR) thermal safety constraint. Solved via a scalable nested Alternating Direction Method of Multipliers (ADMM) algorithm, the CDRO controller achieves near-zero thermal violations under extreme workload spikes in high-fidelity EnergyPlus co-simulations. Simultaneously, it reduces the operational cost premium of robustness by approximately 13.7 percentage points relative to standard Min-Max Model Predictive Control (MPC).

Full text 中文海报
算电协同 论文图示
Research Article热管理与液冷

Financing Artificial Intelligence Infrastructure: Mapping AI Infrastructure Investment and Compute Governance Across Africa

Kai-Hsin Hung, Sumaya Nur Adan, Krupa Suchak, Armita Sadeghian Barzoki, Kofi Yeboah, Mohammad Amir Anwar

Published 2026-06-24 · arXiv · Credibility S

Artificial intelligence depends on large-scale compute resources and their supporting infrastructure. However, AI governance debates treat compute primarily as a technical input rather than as an outcome of investment, ownership, and financial control. This paper examines AI infrastructure investment flows across Africa through a systematic analysis of 46 publicly announced projects totalling USD $12.7 billion betwe…

Abstract, interpretation and reference

Abstract

Artificial intelligence depends on large-scale compute resources and their supporting infrastructure. However, AI governance debates treat compute primarily as a technical input rather than as an outcome of investment, ownership, and financial control. This paper examines AI infrastructure investment flows across Africa through a systematic analysis of 46 publicly announced projects totalling USD $12.7 billion between 2019 and 2025. Using a value chain framework, we analyze who invests in AI-relevant infrastructure and where investments concentrate. Our findings reveal a highly concentrated landscape dominated by global data center operators, hyperscale technology firms, and development finance institutions, clustering in South Africa, Kenya, Nigeria, and Egypt. We introduce asymmetrical interdependence to describe a structural condition in which capital and physical infrastructure account for 73% of total funding while control remains concentrated in the compute layer among a small number of global technology firms. We argue that compute governance must account for capital flows, ownership, and control, not only geographic access, because these dynamics shape AI compute equity. Infrastructure presence is necessary but insufficient for meaningful governance capacity.

Full text 中文海报
热管理与液冷 论文图示
Research Article算电协同

A Bilevel Framework for Data Center-Grid Coordination with DLMPs in Unbalanced Three-Phase Distribution Systems

Arash Baharvandi, Duong Tung Nguyen

Published 2026-06-24 · arXiv · Credibility S

This paper proposes a grid-aware coordination framework between data centers and distribution grids using a DLMP-based bilevel optimization model. The data center aggregator (DCA) determines active power demand in response to distribution locational marginal prices (DLMPs), while the distribution system operator (DSO) solves a network-constrained optimal power flow problem to determine DLMPs in an unbalanced three-p…

Abstract, interpretation and reference

Abstract

This paper proposes a grid-aware coordination framework between data centers and distribution grids using a DLMP-based bilevel optimization model. The data center aggregator (DCA) determines active power demand in response to distribution locational marginal prices (DLMPs), while the distribution system operator (DSO) solves a network-constrained optimal power flow problem to determine DLMPs in an unbalanced three-phase system. The model incorporates both active and reactive power consumption of data centers to evaluate their impacts on voltage regulation and phase imbalance. To mitigate adverse network effects, two operating cases are analyzed: without reactive power compensation and with static var generator (SVG)-based compensation. The proposed approach is validated on the IEEE 37-bus unbalanced distribution test system. Simulation results show that DLMP-based coordination captures economically efficient data center operation, and phase- and location-dependent network conditions, while SVG-based compensation improves voltage profiles and reduces phase unbalance.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

GaN Power Devices and Converter Architectures for AI Data Centers: Efficiency, Reliability, and Deployment Pathways

Donald Intal, Abasifreke Ebong

Published 2026-06-24 · arXiv · Credibility S

The growth of artificial-intelligence workloads is increasing the electrical and thermal demands on data-center power-delivery systems, making conversion efficiency, power density, and reliability critical design priorities. This review examines how gallium-nitride (GaN) power devices can be matched to specific stages of the grid-to-load conversion chain, including power-factor correction, isolated DC/DC conversion,…

Abstract, interpretation and reference

Abstract

The growth of artificial-intelligence workloads is increasing the electrical and thermal demands on data-center power-delivery systems, making conversion efficiency, power density, and reliability critical design priorities. This review examines how gallium-nitride (GaN) power devices can be matched to specific stages of the grid-to-load conversion chain, including power-factor correction, isolated DC/DC conversion, 48-V intermediate-bus conversion, and point-of-load regulation. Si, SiC, and GaN are compared using converter-relevant metrics, and lateral, vertical, and specialized GaN architectures are evaluated in terms of voltage scalability, switching behavior, reverse conduction, thermal pathways, gate control, and technology maturity. The analysis shows that GaN provides a stage-dependent rather than universal advantage. Commercial lateral GaN HEMTs are particularly effective in high-frequency, low-to-mid-voltage stages, while specialized and hybrid devices support bidirectional operation, normally-off control, extreme conversion ratios, and integration. Vertical GaN remains an emerging option for higher-voltage and higher-power conversion. A quantitative framework links cascaded converter efficiency to electrical-loss reduction, cooling demand, annual facility energy use, and operational carbon emissions. Broad deployment further requires low-parasitic packaging, disciplined gate-drive and EMI co-design, mission-profile reliability qualification, scalable manufacturing, and supply-chain resilience. GaN is therefore best treated as a stage-specific system lever whose value depends on coordinated device, topology, package, and thermal co-design.

Full text 中文海报
算电协同 论文图示
Research Article芯片与算力

Node-Level Performance and Energy Characterization of Flagship Science Applications on SuperMUC-NG Phase 2

Salvatore Cielo, Elmira Birang, Alexander Pöppl, Sajad Azizi, Plamen Dobrev, Margarita Egelhofer, Ivan Pribec, Gerald Mathias

Published 2026-06-22 · arXiv · Credibility S

We present a systematic performance and energy-efficiency characterization of five flagship scientific workloads on SuperMUC-NG phase 2, the 28 PetaFLOPs system at the Leibniz Supercomputing Center (LRZ) equipped with Intel Xeon Platinum 8480+ and Intel Data Center GPU Max 1550 (Ponte Vecchio, PVC) accelerators. The selected codes span molecular dynamics (gromacs, lammps), astrophysics and cosmology (OpenGadget3, At…

Abstract, interpretation and reference

Abstract

We present a systematic performance and energy-efficiency characterization of five flagship scientific workloads on SuperMUC-NG phase 2, the 28 PetaFLOPs system at the Leibniz Supercomputing Center (LRZ) equipped with Intel Xeon Platinum 8480+ and Intel Data Center GPU Max 1550 (Ponte Vecchio, PVC) accelerators. The selected codes span molecular dynamics (gromacs, lammps), astrophysics and cosmology (OpenGadget3, AthenaK), and finite-element PDE solvers from the dealii-X Center of Excellence. For each code we measure throughput and energy efficiency expressed as compute-elements per wall-clock second (or per Joule of consumed energy) on a single compute node, comparing CPU-only (SPR) against combined CPU+GPU (SPR+PVC) configurations where available. Energy measurements rely on lightweight code instrumentation with p3em, or the Energy Aware Runtime (EAR) present on the system. Our results show that GPU offload yields $4-12\times$ higher throughput and up to $15\times$ better energy efficiency compared to CPU-only execution, with lammps and AthenaK benefiting most. However, both throughput and energy gains are sensitive to problem granularity: insufficient work per GPU tile erodes the accelerator advantage, as clearly observed in AthenaK at small mesh-block sizes. The power-budget utilization is systematically lower for CPUs than it is for GPUs, indicating that even at peak useful-work rate, most applications running on CPUs leave a significant fraction of the node's thermal envelope unused.

Full text 中文海报
芯片与算力 论文图示
Research ArticleAI 运维优化

Hot AI in Cold Space: Thermal-Crosstalk-Aware Scheduling for Sustainable Orbital AI Clusters

Shuyi Chen, Zhengchang Hua, Nikos Tziritas, Georgios Theodoropoulos

Published 2026-06-23 · arXiv · Credibility S

Terrestrial AI training faces an unsustainable energy and water crisis, positioning Orbital Data Centers (ODCs) as a "zero operational carbon" alternative. However, the sub-$10μ\text{s}$ communication latency required for synchronized scientific workloads, such as distributed Large Language Model (LLM) training, forces ODCs into extreme physical density, triggering a critical "Proximity-Thermal Paradox." As these hi…

Abstract, interpretation and reference

Abstract

Terrestrial AI training faces an unsustainable energy and water crisis, positioning Orbital Data Centers (ODCs) as a "zero operational carbon" alternative. However, the sub-$10μ\text{s}$ communication latency required for synchronized scientific workloads, such as distributed Large Language Model (LLM) training, forces ODCs into extreme physical density, triggering a critical "Proximity-Thermal Paradox." As these high-density systems scale into Monolithic Structures or Proximity Swarms, they suffer from intense thermal-fluid crosstalk (heat traps in shared cooling loops) and thermal-radiative crosstalk (mutual heating that blocks deep-space cooling radiators). If left unmitigated, this persistent heat stagnation not only triggers severe thermal throttling that degrades training throughput, but also induces severe thermal fatigue, drastically shortening hardware lifespans and generating premature space e-waste. To make orbital AI truly sustainable, this position paper challenges traditional uniform load-sharing. We propose the Thermal-Aware Heterogeneity Thesis, which treats spatial cooling variances as a primary resource management dimension. Building on this, we introduce Thermal-Load Balancing (TLB), a software framework that dynamically migrates these intensive workloads to the coolest available units based on instantaneous fluid temperatures or absorbed radiation. Our analysis demonstrates that TLB resolves thermal bottlenecks to restore Model Flops Utilization (MFU), while simultaneously reducing physical thermal stress. Extending the operational lifespan of orbital hardware is crucial to amortize the massive embodied carbon of rocket launches, outlining a necessary pathway to scale orbital AI without accelerating e-waste.

Full text 中文海报
AI 运维优化 论文图示
Research Article热管理与液冷

Toward Next-Generation AI Data Centers: Power Delivery Architecture Shifts, Emerging Technologies, and Challenges

Sangwhee Lee, Rafal P. Wojda, Cheol-Hee Jo, Shuntaro Inoue, Pedro Ribeiro, Gui-Jia Su, Mostak Mohammad, Himel Barua, Nishanth Gadiyar, Praveen Kumar, Spencer Cochran, Subho Mukherjee, Whit Vinson, Vandana Rallabandi, Shajjad Chowdhury, Burak Ozpineci

Published 2026-06-23 · arXiv · Credibility S

The rapid growth of AI workloads is driving unprecedented increases in data center power demand, current transients, and thermal stress, exposing fundamental limitations in traditional 48 V rack architectures, low-voltage AC distribution, and line-frequency transformer interfaces. This paper reviews the three stages of architectural shifts required to support next-generation AI data centers and identifies three enab…

Abstract, interpretation and reference

Abstract

The rapid growth of AI workloads is driving unprecedented increases in data center power demand, current transients, and thermal stress, exposing fundamental limitations in traditional 48 V rack architectures, low-voltage AC distribution, and line-frequency transformer interfaces. This paper reviews the three stages of architectural shifts required to support next-generation AI data centers and identifies three enabling technological building blocks: high-voltage conversion-ratio DC/DC converters, facility-level low-voltage DC distribution, and medium-voltage solid-state transformers. The advantages, technical challenges, and potential solutions associated with each building block are reviewed. Finally, future research directions and open challenges are discussed.

Full text 中文海报
热管理与液冷 论文图示
Research Article算电协同

Power-Flexible AI Data Centers: A New Paradigm for Grid-Responsive Compute

Chris Williams, Philip Colangelo, Ayse Coskun, Ethan Levine, Andy Neale, Ciaran Roberts, Shayan Sengupta, Nikhil Shirolkar, Varun Sivaram, Sarah Soares, Ethan Tiao, Scott Underwood, Daniel Wilson, Frank Sharp, Luke Wainwright, Harry Petty, Scott Wallace, Brandon Records

Published 2026-06-23 · arXiv · Credibility S

The rapid expansion of artificial intelligence (AI) infrastructure is driving unprecedented growth in electricity demand from data centers. Traditional power-system planning treats large computing facilities as inflexible peak loads, leading to costly infrastructure upgrades and long delays in grid interconnection. Recent work has shown that AI clusters can reduce electricity consumption during peak demand through s…

Abstract, interpretation and reference

Abstract

The rapid expansion of artificial intelligence (AI) infrastructure is driving unprecedented growth in electricity demand from data centers. Traditional power-system planning treats large computing facilities as inflexible peak loads, leading to costly infrastructure upgrades and long delays in grid interconnection. Recent work has shown that AI clusters can reduce electricity consumption during peak demand through software-based workload orchestration. This article explores how modern GPU-based AI data centers can operate as grid-interactive assets that respond dynamically to power system conditions. We describe an architecture integrating grid signals, workload scheduling, and power telemetry for fine-grained cluster power control. Experimental results from a real-world deployment on a 130 kW GPU cluster demonstrate multiple forms of flexibility, including rapid load reduction, sustained curtailment, and carbon-aware operation while preserving service levels for priority jobs. We further demonstrate performance-aware load shifting across geographically distributed clusters, enabling workloads to migrate toward regions with lower grid stress. Together, these capabilities transform AI infrastructure from static electricity consumers into flexible resources that support grid reliability, accelerate interconnection, and improve computing sustainability.

Full text 中文海报
算电协同 论文图示