智算中心论文专站

AIDC Research Papers

Liquid Cooling AI Data Center Power & Thermal Systems
Current Issue

Volume 2026 · Issue 08-11

按期刊卷期页方式整理本期论文。每条仅使用日报已列出的可追溯公开来源,不新增未经核验事实。

Research Article热管理与液冷

Predictive Failure Detection in Network Hardware Using Thermal Imaging and Deep Learning with Sensor Fusion

Ashly Joseph

Published 2026-08-05 · arXiv · Credibility S

Unplanned network hardware malfunctions can interrupt services and result in expensive downtime in data centers. A deep learning-based predictive maintenance strategy is presented that utilizes thermal imaging and power sensor data to detect early indicators of equipment breakdown in routers, switches, and servers. A simulated dataset was generated comprising annotated thermal pictures and power readings indicative …

Abstract, interpretation and reference

Abstract

Unplanned network hardware malfunctions can interrupt services and result in expensive downtime in data centers. A deep learning-based predictive maintenance strategy is presented that utilizes thermal imaging and power sensor data to detect early indicators of equipment breakdown in routers, switches, and servers. A simulated dataset was generated comprising annotated thermal pictures and power readings indicative of three operating states: Normal, Warning, and Critical. Three ImageNet-pretrained convolutional neural network (CNN) models ResNet-50, InceptionV3, and VGG16 were assessed together with a multi-modal CNN-LSTM fusion model that integrates visual and sensor time-series information. Experiments were performed with and without pre-processing procedures, including region-of-interest (ROI) extraction and normalization. In the absence of pre-processing, CNNs attained moderate accuracy (e.g., ResNet-50 at 52%), but ROI-based pre-processing significantly enhanced performance (ResNet-50 accuracy reaching 91%). The CNN-LSTM model attained the greatest accuracy of 94%, with precision and recall approaching 95%, illustrating the effectiveness of multi-modal fusion. The results validate that domain-specific pre-processing and sensor fusion substantially improve early failure prediction, providing a potential foundation for proactive maintenance of network hardware through non-intrusive monitoring.

Full text 中文海报
热管理与液冷 论文图示
Research Article算电协同

Beyond the Grid: Cost, Carbon, and Capital Requirements of On-Site Power Technologies for AI Data Centers

Eliseo Curcio

Published 2026-08-08 · arXiv · Credibility S

Interconnection queues, not electricity prices, now govern where data centers can be built, and the standard levelized-cost comparison answers a question no developer faces: it assumes a load profile, freezes the grid price while modeling the demand that moves it, and quotes busbar costs a facility cannot buy. This paper evaluates nine on-site supply technologies against a delivered grid whose price is endogenous to…

Abstract, interpretation and reference

Abstract

Interconnection queues, not electricity prices, now govern where data centers can be built, and the standard levelized-cost comparison answers a question no developer faces: it assumes a load profile, freezes the grid price while modeling the demand that moves it, and quotes busbar costs a facility cannot buy. This paper evaluates nine on-site supply technologies against a delivered grid whose price is endogenous to projected data-center demand, on a complete-site basis that retains standby charges, with measured GPU training load, delivered fuel prices, production-pathway carbon, and statutory 45V and 48E incentive mechanics. Nothing beats the wire: gas combined cycle produces at 47 USD/MWh but costs about 114 USD per megawatt-hour of complete site energy against a 92 USD grid; four-hour storage is physically capped near 18 percent of annual energy and, charged at the margin, dirtier than the grid; hydrogen from grid-priced power fails on cost and carbon together. An investment inversion converts these findings into capital terms: conversion-hardware learning buys nothing, because free hardware still exceeds the grid for every low-carbon arm, while global electrolyser deployment on sited sub-20 USD/MWh power brings PEM hydrogen power to about 2.2 times the grid at 300 billion USD and 1.9 times at 1 trillion USD (2.7 and 2.3 for the hydrogen engine), with a carbon reduction of roughly 85 percent (6.8-fold) against grid-power production. Grid parity is not purchasable at any budget. On-site supply is an access and depth product; most current investment targets the wrong term.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

Environmental and Economic Implications of Artificial Intelligence Data Centers in the United States

Johanna Bolaños-Zuñiga, Alberto J. Lamadrid

Published 2026-08-10 · arXiv · Credibility S

In this study, we use electricity demand growth, cooling requirements, and backup system operation to evaluate the environmental and economic implications of artificial intelligence data centers in the United States. Our results indicate that impacts are not determined solely by facility design, but by the broader electricity, water, and land-use systems in which these facilities operate. Emissions are primarily dri…

Abstract, interpretation and reference

Abstract

In this study, we use electricity demand growth, cooling requirements, and backup system operation to evaluate the environmental and economic implications of artificial intelligence data centers in the United States. Our results indicate that impacts are not determined solely by facility design, but by the broader electricity, water, and land-use systems in which these facilities operate. Emissions are primarily driven by electricity consumption and therefore depend on marginal generation mixes, transmission constraints, and the spatial and temporal distribution of demand. Analysis further shows that local effects include pressures on water resources, increased noise exposure, and land-use changes, with outcomes varying across regions and infrastructure conditions. The assessment of technological and operational measures shows that improvements in energy efficiency, cooling configurations, and operational strategies can reduce these impacts, although their effectiveness depends on system-level conditions. Evaluation of regulatory and market structures suggests that existing frameworks may not fully account for location- and time-specific externalities. These findings support the need for integrated policy approaches that align data center deployment and operation with electricity system characteristics, water availability, and land-use planning to improve overall environmental and economic performance.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

A Hierarchical Semi-Markov Load Model for AI Data Centers Coupling Job Scheduling with Bulk-Synchronous-Parallel Power Dynamics

Chandan Chaudhary, Atri Bera, Cody Newlun, Mohammed Ben-Idris, Joydeep Mitra

Published 2026-07-13 · arXiv · Credibility S

AI data centers are emerging as a dominant new load class with their power dynamics fundamentally from conventional industrial loads. Inside a training job, the bulk-synchronous-parallel algorithm moves each node through compute, sync, and checkpoint steps, which swings power between full load and near idle within seconds. Across the whole facility, jobs arrive, take blocks of nodes for hours to days, then leave, so…

Abstract, interpretation and reference

Abstract

AI data centers are emerging as a dominant new load class with their power dynamics fundamentally from conventional industrial loads. Inside a training job, the bulk-synchronous-parallel algorithm moves each node through compute, sync, and checkpoint steps, which swings power between full load and near idle within seconds. Across the whole facility, jobs arrive, take blocks of nodes for hours to days, then leave, so the number of busy nodes changes daily, weekly, and yearly. This slower shift drives facility-wide swings and the peak demand that sets the size of the grid link. A model that looks only at within-job behavior, and treats the facility as a fixed set of busy nodes, smooths out these swings and misses the true peak-to-average ratio. This paper develops a hierarchical semi-Markov Data-Center (HSM-DC) load model that couples two layers across two timescales. A job-scheduling layer creates jobs through a non-homogeneous compound-Poisson process shaped by daily, weekly, and seasonal patterns, gives each job a heavy-tailed node count and length, and places jobs on a fixed pool of nodes on a first-come basis. A within-job layer moves each busy node through a five-state semi-Markov chain for the BSP steps, with state-based Ornstein-Uhlenbeck noise. Facility power comes from this changing node count and the per-node power, set to match measured node data and the facility's straight-line power-versus-load curve. Configured to the reference facility at the same scale, the model matches mean power, its spread, and the peak-to-average ratio across load levels, with fit scores of 0.9997, 0.92, and 0.82. It also matches the share of queued jobs to within one point at high load. Facility-wide swings and peak demand come from how jobs arrive and get scheduled, so grid planning must model that process, not just scale up a single node's power curve.

Full text 中文海报
算电协同 论文图示
Research Article热管理与液冷

Zero-change foundry compatible silicon photonics MEMS optical switch

Arkadev Roy, Daniel Klawson, Jianheng Luo, Yiyang Zhi, Sirui Tang, Ming Wu

Published 2026-08-04 · arXiv · Credibility S

Large-scale photonic switches are emerging as essential devices for energy-efficient optical interconnect in data centers and AI/ML clusters as a key enabler for high-bandwidth and low-latency connectivity. Combining micro-electro-mechanical (MEMS) based mechanical reconfigurability with silicon photonic integrated circuits can enable a large-scale, low-loss, programmable platform required for large-scale optical ci…

Abstract, interpretation and reference

Abstract

Large-scale photonic switches are emerging as essential devices for energy-efficient optical interconnect in data centers and AI/ML clusters as a key enabler for high-bandwidth and low-latency connectivity. Combining micro-electro-mechanical (MEMS) based mechanical reconfigurability with silicon photonic integrated circuits can enable a large-scale, low-loss, programmable platform required for large-scale optical circuit switches. We demonstrate a broadband silicon photonics MEMS switch with more than 30 dB extinction ratio operating in C-band using a zero-change foundry-compatible process and Back-end-of-Line (BEOL) post-processing. The optical switch element exhibits an insertion loss of less than 1.5 dB with a low static power consumption of approx 20 nW at maximum actuation voltage. Our results illustrate that MEMS-based silicon photonics modulators and phase shifters can be used alongside standard silicon photonics components seamlessly in scenarios where performance in terms of footprint, extinction ratio, broad bandwidth, and low-loss operation is of paramount importance.

Full text 中文海报
热管理与液冷 论文图示
Research Article热管理与液冷

The Cost and Network Limits of Space-Based AI Compute

Kees van Berkel

Published 2026-07-15 · arXiv · Credibility S

This paper evaluates whether large-scale AI data centers deployed in low-Earth orbit (LEO) could become a cost-effective alternative to terrestrial facilities. The analysis compares orbital and ground-based systems across launch cost, power generation, cooling, radiation exposure, and atmospheric reentry, as well as compute-network performance. A key distinction is the shift from terrestrial Clos networks to space-b…

Abstract, interpretation and reference

Abstract

This paper evaluates whether large-scale AI data centers deployed in low-Earth orbit (LEO) could become a cost-effective alternative to terrestrial facilities. The analysis compares orbital and ground-based systems across launch cost, power generation, cooling, radiation exposure, and atmospheric reentry, as well as compute-network performance. A key distinction is the shift from terrestrial Clos networks to space-based mesh networks using laser inter-satellite links. Using bisection bandwidth, bisection intensity, and roofline-style models, we show that while LEO-based inference may be feasible, training frontier-scale LLMs in orbit is unlikely to be competitive with terrestrial data centers.

Full text 中文海报
热管理与液冷 论文图示
Research Article算电协同

The Environmental Cost of Digital Sovereignty: Water, Energy, and Emissions Impacts of Sovereign AI Infrastructure in the Global South

Muntaser Syed, Marius C. Silaghi, Sheikh Abujar, Sharun Akter Khushbu, Amal El Ahmad

Published 2026-07-15 · arXiv · Credibility S

Sovereign AI has become a strategic priority across the Global South, with over \$200 billion in state-led commitments announced between 2024 and 2026. Yet the physical infrastructure that compute sovereignty demands, above all data centers, imposes water, energy, and carbon costs that fall hardest on countries least equipped to absorb them. This paper presents a comparative environmental stress analysis across four…

Abstract, interpretation and reference

Abstract

Sovereign AI has become a strategic priority across the Global South, with over \$200 billion in state-led commitments announced between 2024 and 2026. Yet the physical infrastructure that compute sovereignty demands, above all data centers, imposes water, energy, and carbon costs that fall hardest on countries least equipped to absorb them. This paper presents a comparative environmental stress analysis across four cases: the United Arab Emirates, Bangladesh, India, and Africa (with a focus on Kenya). Using publicly available water stress data, grid carbon intensity factors, and GPU power specifications, we model the water consumption, energy demand, and carbon emissions of hypothetical sovereign AI deployments under multiple cooling technology scenarios. We find that a 1,024-GPU cluster using evaporative cooling in the UAE would consume over 30 million liters of water annually in a country classified as ``extremely high'' water stress. In Bangladesh, sovereign AI policy documents call for centralized GPU procurement but do not address where to site data centers in a country where more than a fifth of the land floods in an average year and the power grid struggles to deliver reliable supply. We identify a sovereignty-sustainability trilemma in which no country can simultaneously maximize AI sovereignty, minimize environmental impact, and maintain affordable resource access for citizens. We propose design principles for environmentally responsible sovereign AI, including mandatory water usage effectiveness reporting, climate-vulnerability siting assessments, and a preference for frugal small language models over frontier pre-training in resource-constrained settings.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

A Stackelberg-Bayesian Capacity-Market Game of Carbon Regulation and Second-Life Battery Investment under AI Data-Center Load Growth

Rouzbeh Haghighi, Ali Hassan, Sina Mohammadi, Marcus Chen I Wada, Wencong Su

Published 2026-08-04 · arXiv · Credibility S

Artificial intelligence (AI) data centers are driving rapid electricity load growth across all U.S. ISO/RTO regions, raising both system costs and carbon exposure. This study develops a three-level Stackelberg--Bayesian game in which a regulator (leader) sets carbon penalties and subsidies, a single ISO capacity market clears against an energy balance modeled as a classical generation-expansion problem, and technolo…

Abstract, interpretation and reference

Abstract

Artificial intelligence (AI) data centers are driving rapid electricity load growth across all U.S. ISO/RTO regions, raising both system costs and carbon exposure. This study develops a three-level Stackelberg--Bayesian game in which a regulator (leader) sets carbon penalties and subsidies, a single ISO capacity market clears against an energy balance modeled as a classical generation-expansion problem, and technology-specific investors (followers) decide capacity and operation under incomplete information, yielding a Bayesian Nash equilibrium. The AI impact is captured parsimoniously as an additional load-growth factor on a greenfield-incremental expansion, isolating how much new capacity the growth pulls in and which technology fills it. Within this framework, we consider second-life battery (SLB) storage competing against new/first-life storage for capacity-market revenue. We quantify how a carbon tax, a renewable subsidy, and an SLB subsidy reshape the equilibrium investment mix, carbon emissions, and profit. Different scenarios are compared at the end based on cost-effectiveness and reduced carbon emissions.

Full text 中文海报
算电协同 论文图示
Research Article热管理与液冷

Assessing Risks of Hydro-Generator Shaft Fatigue from Data Center Load Oscillations

Kaustav Chatterjee, Meghana Ramesh, Shuchismita Biswas, Brett A. Ross, Antos C. Varghese, Sameer Nekkalapu, Slaven Kincic

Published 2026-07-15 · arXiv · Credibility S

Large AI data center loads can introduce persistent sub-synchronous active-power oscillations that may impact nearby generators by exciting torsional modes and increasing shaft stress. This paper presents a model-based framework for evaluating hydro-generator shaft fatigue risk under oscillatory loading. An electromagnetic transient simulation model is developed using a two-mass turbine-generator shaft representatio…

Abstract, interpretation and reference

Abstract

Large AI data center loads can introduce persistent sub-synchronous active-power oscillations that may impact nearby generators by exciting torsional modes and increasing shaft stress. This paper presents a model-based framework for evaluating hydro-generator shaft fatigue risk under oscillatory loading. An electromagnetic transient simulation model is developed using a two-mass turbine-generator shaft representation with parameters from real-world generation units and a configurable AI data center load. The risk assessment is performed in two stages. First, a network transfer function quantifies the propagation of load oscillations from the data center point of interconnection to the hydro-generator terminal. A plant transfer function then characterizes the resulting shaft torque amplification. A frequency-scan approach identifies resonance regions and evaluates torque amplification at individual forcing frequencies. Parametric studies show that amplification is strongly affected by generator-to-turbine inertia ratio and torsional damping. Lower inertia ratios shift torsional modes to lower frequencies and increase amplification, indicating that some Kaplan-type units may be more susceptible than comparable Francis or Pelton units. Reduced damping further increases resonant response and fatigue exposure. A simplified fatigue assessment based on S--N curves and the Goodman diagram relates simulated torque response to mechanical integrity. The resulting Goodman safety factor provides a practical metric for evaluating the impact of persistent AI data center oscillations on hydro-generator service life and supports interconnection studies, oscillation limits, and plant-level monitoring strategies.

Full text 中文海报
热管理与液冷 论文图示
Research Article算电协同

A Phased Development Framework Enabling Islanded Operation of Sustainable AI Data Centers With Onsite Grid-Following and Grid-Forming Energy Architectures

Soham Ghosh, Nabil Mohammed, Mohammad Ashraf Hossain Sadi

Published 2026-07-19 · arXiv · Credibility S

As hyperscale and colocation AI data centers continue to expand, the electric grid is increasingly required to support large, concentrated loads, with individual facilities ranging from 500 MW to 2 GW. Current projections estimate that approximately 50 GW of AI data center capacity will require grid connectivity in the United States by 2030. While prior research has extensively examined the environmental and operati…

Abstract, interpretation and reference

Abstract

As hyperscale and colocation AI data centers continue to expand, the electric grid is increasingly required to support large, concentrated loads, with individual facilities ranging from 500 MW to 2 GW. Current projections estimate that approximately 50 GW of AI data center capacity will require grid connectivity in the United States by 2030. While prior research has extensively examined the environmental and operational impacts of AI data centers, as well as their potential role as grid-interactive assets, limited attention has been given to the challenges associated with their scalable deployment through engineering, procurement, and construction (EPC) processes. This manuscript addresses this gap by proposing a phased development framework for AI data center expansion. The approach is designed to enable developers to meet aggressive time-to-market objectives while navigating multi-year constraints associated with interconnection approvals and lead times associated with the procurement of component equipment. A modular construction architecture is presented, along with a detailed analysis of integrated energy systems and the role of hybrid on-site generation in supporting incremental capacity growth. Electromagnetic transient simulations (EMT) are used to evaluate system performance, demonstrating that a combination of on-site natural gas generation and grid-forming energy storage can reliably support data center operations during early and intermediate deployment phases. The study further examines the transition to full grid interconnection, including the capability of the data center to operate in islanded mode during grid disturbances. Finally, the manuscript compares grid-forming control strategies for system reconnection and restoration under varying conditions.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

Inter-Area Oscillation Damping in Data-Center-Integrated Power Systems

Ahmed Alfatlawi, Masoud H. Nazari

Published 2026-07-30 · arXiv · Credibility S

This paper develops explicit dynamic models of a hyperscale data center, including its heating, ventilation, and air conditioning (HVAC) and uninterruptible power supply (UPS) subsystems, and integrates them into a small-signal stability framework to investigate the impact of data center demand response on power system inter-area oscillations. Through eigenvalue analysis and time-domain simulations, the results demo…

Abstract, interpretation and reference

Abstract

This paper develops explicit dynamic models of a hyperscale data center, including its heating, ventilation, and air conditioning (HVAC) and uninterruptible power supply (UPS) subsystems, and integrates them into a small-signal stability framework to investigate the impact of data center demand response on power system inter-area oscillations. Through eigenvalue analysis and time-domain simulations, the results demonstrate that UPS-based demand response can enhance inter-area oscillation damping. In contrast, the HVAC subsystem is shown to be inherently incapable of providing effective oscillation damping due to its limited thermal response bandwidth. A gradient-based optimization algorithm is used to tune the UPS controller gain to maximize the damping ratio of the critical inter-area mode. The effectiveness of the proposed approach is validated using the IEEE 39-bus test system.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

A Configurable Thermal-Dynamic Model for AI Data Center Cooling Load Simulation

Cletus Ngwerume, Lang Tong, Chee-Wooi Ten, Yi Hu

Published 2026-07-31 · arXiv · Credibility S

Cooling demand constitutes a significant and flexible component of AI data center electricity consumption, but time-synchronized measurements are scarce and constant coefficient-of-performance models cannot represent thermal dynamics. This letter proposes a configurable thermal dynamic simulation model for hybrid air- and liquid-cooled data centers. Unlike existing models centered on temperature prediction or equipm…

Abstract, interpretation and reference

Abstract

Cooling demand constitutes a significant and flexible component of AI data center electricity consumption, but time-synchronized measurements are scarce and constant coefficient-of-performance models cannot represent thermal dynamics. This letter proposes a configurable thermal dynamic simulation model for hybrid air- and liquid-cooled data centers. Unlike existing models centered on temperature prediction or equipment-level cooling analysis, the proposed model is designed to generate dynamic cooling electricity profiles for long-duration power system studies. The model is validated using operational telemetry from the Marconi100 supercomputer. Compared with the baseline, the proposed model reduces the mean absolute error from 95.80 to 20.88~kW and the root-mean-square error from 109.79 to 27.27~kW. Evaluation over approximately 520 daily profiles further shows improved reproduction of daily peak demand and intraday variability. The proposed model provides a computationally tractable means of generating physically interpretable cooling load profiles for power system studies.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

Planning Waste-to-Energy-Coupled AI Data Centers Through Grade-Matched Cooling and Corridor Screening

Qi He, Chunyu Qu, Wenjie Zuo

Published 2026-07-28 · arXiv · Credibility S

AI data-center growth is increasingly constrained by limited deliverable electricity, interconnection capacity, and cooling demand. This study develops a boundary-consistent screening framework for waste-to-energy (WtE)-coupled AI data-center cooling. It treats cooling as an energy service that can be supplied through grade matching rather than only through electricity-driven mechanical chilling. The framework trans…

Abstract, interpretation and reference

Abstract

AI data-center growth is increasingly constrained by limited deliverable electricity, interconnection capacity, and cooling demand. This study develops a boundary-consistent screening framework for waste-to-energy (WtE)-coupled AI data-center cooling. It treats cooling as an energy service that can be supplied through grade matching rather than only through electricity-driven mechanical chilling. The framework translates plant-side exportable heat into corridor-level planning metrics by accounting for thermal attenuation, absorption conversion, and parasitic electricity for delivery and auxiliaries. In a reference case, a regulated WtE plant processing 1500 t/day of municipal solid waste at 10 MJ/kg provides about 78.1 MWth of exportable heat. At a 20 km corridor, this yields about 53.0 MW of delivered cooling and 8.0 MWe of net avoided cooling electricity after parasitic loads. The coupled system is governed by operating regimes rather than a single efficiency score. Under baseline assumptions, full thermal coverage extends to about 20.9 km, the quality-adjusted criterion remains positive to about 22.9 km, and net electricity relief remains positive to about 44.7 km. For a 1 GW IT campus at 70 percent utilization and a 5 km corridor, net grid relief ranges from about 116.9 to 264.4 MW across scenarios. The required WtE footprint ranges from about 3 to 148 representative plants, or 0.6 to 40 full-load-equivalent plants at a 25 percent displacement target. The framework identifies when WtE-coupled cooling is corridor-feasible, when hybrid operation is required, and when infrastructure scale becomes the binding constraint. It is intended for screening and comparison, not project-specific hydraulic or plant-cycle design.

Full text 中文海报
算电协同 论文图示
Research Article算电协同

Emission-Forecasting-Based Spatial-Temporal Carbon Response: A Multi-Agent Attention-Enhanced Deep Learning Framework

Feiyu Cai, Jing Qiu, Yi Yang, Chenxi Zhang, Xinlei Wang, Baichuan Liu, Junhua Zhao

Published 2026-07-29 · arXiv · Credibility S

As a major contributor to carbon emissions, the decarbonization of power systems has garnered significant societal attention. Nodal carbon intensity (NCI), a critical factor in carbon-oriented demand response, has traditionally been determined through ex-post calculations. However, this ex-post approach introduces latency in low-carbon dispatch. To address this, this paper presents a proactive ex-ante spatial-tempor…

Abstract, interpretation and reference

Abstract

As a major contributor to carbon emissions, the decarbonization of power systems has garnered significant societal attention. Nodal carbon intensity (NCI), a critical factor in carbon-oriented demand response, has traditionally been determined through ex-post calculations. However, this ex-post approach introduces latency in low-carbon dispatch. To address this, this paper presents a proactive ex-ante spatial-temporal carbon response framework. At its core, we develop a novel deep learning-based hierarchical design, enhanced by a dual-stage attention mechanism and a large language model (LLM)-based multi-agent cooperation system, to accurately forecast day-ahead NCI. This design effectively mitigates the impact of renewable energy uncertainty and enhances predictive resilience. On the demand side, the framework proposes a spatial-temporal carbon scheduling model that integrates geographically dispatchable loads (GDLs), including mobile energy storage systems (MESSs) and distributed data centers (DDCs). Leveraging high-accuracy day-ahead NCI predictions, the framework can effectively reduce system emissions by quickly responding to carbon intensity fluctuations. The proposed framework is tested on the modified IEEE 33-bus system. According to the simulation results, the impacts of proposed framework on dispatching latency and emission outcomes are analyzed. The results demonstrate that under a one-hour reduction in carbon scheduling latency, the proposed model and methodology can achieve over 30% emission reduction. This research breaks through the limitations of passive carbon accounting, advancing toward proactive carbon management. It offers an intelligent solution that accelerates the transition to cleaner power systems while directly supporting sustainable production goals.

Full text 中文海报
算电协同 论文图示
Research Article热管理与液冷

From Individual to Shared Ownership: A Coalitional Game Approach to Sustainable Co-investment

Published 2026-07-29 · arXiv · Credibility S

This paper proposes a cooperative game-theoretic framework for sustainable co-investment in shared infrastructure under regulatory incentives. Multiple heterogeneous operators co-invest in a common infrastructure whose production capability evolves over time and is subject to operational variability. A regulator supports the deployment through incentive mechanisms designed to align individual economic investment obj…

Abstract, interpretation and reference

Abstract

This paper proposes a cooperative game-theoretic framework for sustainable co-investment in shared infrastructure under regulatory incentives. Multiple heterogeneous operators co-invest in a common infrastructure whose production capability evolves over time and is subject to operational variability. A regulator supports the deployment through incentive mechanisms designed to align individual economic investment objectives with the coalitional one. We formulate the co-investment problem as a transferable-utility (TU) coalitional game in which the value generated by cooperation depends on heterogeneous operational profiles, dynamic resource availability, investment costs, and regulatory incentive level. We show that the proposed coalitional game can be reformulated as a linear production game (LPG), whose dual prices yield a constructive and stable allocation of the cooperative surplus. Finally, we illustrate the proposed framework through a case study on co-investment among data center operators in shared renewable energy infrastructure, supported by government subsidies promoting renewable energy consumption.

Full text 中文海报
热管理与液冷 论文图示
Research Article算电协同

A Predict-then-Schedule framework for Power Distribution Networks with AI Data Centers

Siqi Yan, Jiebao Zhang, Xi Yao, Juan Huang, Ye Shi

Published 2026-07-20 · arXiv · Credibility S

The surge of GPU-intensive workloads in artificial intelligence (AI) data centers drives massive energy demands, leading to soaring costs and significant stress on local power distribution networks. Coordinating delay-tolerant workload scheduling with power grid conditions via precise workload prediction can mitigate these issues. However, a critical gap remains in conventional approaches, i.e., minimizing predictio…

Abstract, interpretation and reference

Abstract

The surge of GPU-intensive workloads in artificial intelligence (AI) data centers drives massive energy demands, leading to soaring costs and significant stress on local power distribution networks. Coordinating delay-tolerant workload scheduling with power grid conditions via precise workload prediction can mitigate these issues. However, a critical gap remains in conventional approaches, i.e., minimizing prediction error does not necessarily lead to minimized downstream operational loss. Hence, this paper proposes an end-to-end Predict-Then-Schedule (PTS) framework that integrates upstream workload prediction with downstream scheduling optimization. By leveraging differentiable convex optimization, the PTS framework maps input features directly to optimal scheduling and enables gradient-based training. Furthermore, to respect the data center's capacity, a workload over-shifted loss combining electricity cost with a penalty for load-shedding is introduced to evaluate scheduling quality. Experiments demonstrate that the proposed framework significantly reduces operational cost and enhances system security compared to the conventional two-stage baseline.

Full text 中文海报
算电协同 论文图示