Research Article热管理与液冷
Arkadev Roy、Daniel Klawson、Jianheng Luo、Yiyang Zhi、Sirui Tang、Ming Wu
Published 2026-08-04 · arXiv · Credibility S
Large-scale photonic switches are emerging as essential devices for energy-efficient optical interconnect in data centers and AI/ML clusters as a key enabler for high-bandwidth and low-latency connectivity. Combining micro-electro-mechanical (MEMS) based mechanical reconfigurability with silicon photonic integrated circuits can enable a large-scale, low-loss, programmable platform required for large-scale optical ci…
Abstract, interpretation and reference
Abstract
Large-scale photonic switches are emerging as essential devices for energy-efficient optical interconnect in data centers and AI/ML clusters as a key enabler for high-bandwidth and low-latency connectivity. Combining micro-electro-mechanical (MEMS) based mechanical reconfigurability with silicon photonic integrated circuits can enable a large-scale, low-loss, programmable platform required for large-scale optical circuit switches. We demonstrate a broadband silicon photonics MEMS switch with more than 30 dB extinction ratio operating in C-band using a zero-change foundry-compatible process and Back-end-of-Line (BEOL) post-processing. The optical switch element exhibits an insertion loss of less than 1.5 dB with a low static power consumption of approx 20 nW at maximum actuation voltage. Our results illustrate that MEMS-based silicon photonics modulators and phase shifters can be used alongside standard silicon photonics components seamlessly in scenarios where performance in terms of footprint, extinction ratio, broad bandwidth, and low-loss operation is of paramount importance.
中文解读
背景:AI 数据中心负载、功率密度和能源约束同步上升,液冷、热管理和数据中心能效正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用文献摘要中的模型、实验或案例分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向冷却效率、能源利用或运维策略的改进方向。意义:对日报读者而言,它可用于判断液冷方案、热管理路线和高密度部署节奏。仍需结合全文实验条件、样本范围和成本假设核验。
参考文献
Arkadev Roy, Daniel Klawson, Jianheng Luo, 等. Zero-change foundry compatible silicon photonics MEMS optical switch[J/OL]. (2026-08-04)[2026-08-31]. http://arxiv.org/abs/2608.03146v1.
Research Article算电协同
Cletus Ngwerume、Lang Tong、Chee-Wooi Ten、Yi Hu
Published 2026-07-31 · arXiv · Credibility S
Cooling demand constitutes a significant and flexible component of AI data center electricity consumption, but time-synchronized measurements are scarce and constant coefficient-of-performance models cannot represent thermal dynamics. This letter proposes a configurable thermal dynamic simulation model for hybrid air- and liquid-cooled data centers. Unlike existing models centered on temperature prediction or equipm…
Abstract, interpretation and reference
Abstract
Cooling demand constitutes a significant and flexible component of AI data center electricity consumption, but time-synchronized measurements are scarce and constant coefficient-of-performance models cannot represent thermal dynamics. This letter proposes a configurable thermal dynamic simulation model for hybrid air- and liquid-cooled data centers. Unlike existing models centered on temperature prediction or equipment-level cooling analysis, the proposed model is designed to generate dynamic cooling electricity profiles for long-duration power system studies. The model is validated using operational telemetry from the Marconi100 supercomputer. Compared with the baseline, the proposed model reduces the mean absolute error from 95.80 to 20.88~kW and the root-mean-square error from 109.79 to 27.27~kW. Evaluation over approximately 520 daily profiles further shows improved reproduction of daily peak demand and intraday variability. The proposed model provides a computationally tractable means of generating physically interpretable cooling load profiles for power system studies.
中文解读
背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用实验验证、原型测试或测量对比,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向冷却效率、能源利用或运维策略的改进方向。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。
参考文献
Cletus Ngwerume, Lang Tong, Chee-Wooi Ten, 等. A Configurable Thermal-Dynamic Model for AI Data Center Cooling Load Simulation[J/OL]. (2026-07-31)[2026-08-31]. http://arxiv.org/abs/2607.28962v1.
Research Article余热回收
Wenyu Liu、Enea Figini、Mario Paolone
Published 2026-08-17 · arXiv · Credibility S
This paper proposes a two-layer model predictive control (MPC) framework for the real-time operation of data centers integrated with on-site photovoltaic generation, battery energy storage, waste heat recovery, and district heating. The upper layer employs scenario-based stochastic optimization to jointly optimize intraday market participation, workload scheduling, and energy management under uncertainty. The lower …
Abstract, interpretation and reference
Abstract
This paper proposes a two-layer model predictive control (MPC) framework for the real-time operation of data centers integrated with on-site photovoltaic generation, battery energy storage, waste heat recovery, and district heating. The upper layer employs scenario-based stochastic optimization to jointly optimize intraday market participation, workload scheduling, and energy management under uncertainty. The lower layer adopts an adaptive tube-based MPC strategy that compensates short-term disturbances while tracking the dispatch references given by the upper layer. The framework further integrates multi-horizon forecasting to support real-time decision making. Microservice-based simulation studies under representative clear-sky and overcast operating conditions demonstrate that the proposed framework accurately tracks dispatch plans despite fast photovoltaic and workload fluctuations. Compared with single-layer control strategies, the adaptive lower-layer controller substantially reduces real-time dispatch deviations and the associated imbalance costs. In addition, the proposed framework naturally adapts to seasonal operating conditions and responds to carbon-aware operating signals, offering a practical approach for economically efficient, sustainable, and grid-supportive operation of future data centers.
中文解读
背景:AI 数据中心负载、功率密度和能源约束同步上升,余热回收、热泵耦合和二次能源利用正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用建模优化、调度分析或算法评估,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断数据中心余热能否从成本项转化为能源资产。仍需结合全文实验条件、样本范围和成本假设核验。
参考文献
Wenyu Liu, Enea Figini, Mario Paolone. Real-Time Control of Sustainable Data Centers: A Two-Layer Model Predictive Control Framework with Workload Flexibility and Heat Recovery[J/OL]. (2026-08-17)[2026-08-31]. http://arxiv.org/abs/2608.16432v1.
Research Article芯片与算力
Hanzhao Wang、Jingxuan Wu、Yumeng Li、Yu Pan、Guanting Chen
Published 2026-08-19 · arXiv · Credibility S
The growing demand for AI-driven workloads, particularly from Large Language Models (LLMs), has raised concerns about the significant energy and resource consumption in data centers. This work introduces a novel LLM-based predictive scheduling system designed to enhance operational efficiency while reducing the environmental impact of data centers. Our system utilizes an LLM to predict key metrics such as execution …
Abstract, interpretation and reference
Abstract
The growing demand for AI-driven workloads, particularly from Large Language Models (LLMs), has raised concerns about the significant energy and resource consumption in data centers. This work introduces a novel LLM-based predictive scheduling system designed to enhance operational efficiency while reducing the environmental impact of data centers. Our system utilizes an LLM to predict key metrics such as execution time and energy consumption from source code, and it has the potential to extend to other sustainability-focused metrics like water usage for cooling and carbon emissions, provided the data center can track such data. The predictive model is followed by a real-time scheduling algorithm that allocates GPU resources, aiming to improve sustainability by optimizing both energy consumption and queuing delays. With fast inference times, the ability to generalize across diverse task types, and minimal data requirements for training, our approach offers a practical solution for data center scheduling. This framework demonstrates strong potential for advancing sustainability objectives in AI-driven infrastructure. Through our collaboration with a data center, we achieved a 32% reduction in energy consumption and a 30% decrease in waiting time.
中文解读
背景:AI 数据中心负载、功率密度和能源约束同步上升,芯片、服务器和高密度算力部署正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用建模优化、调度分析或算法评估,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向能效评价口径、运营指标和优化目标的系统化梳理。意义:对日报读者而言,它可用于判断芯片路线和服务器密度变化如何传导到机房设计。仍需结合全文实验条件、样本范围和成本假设核验。
参考文献
Hanzhao Wang, Jingxuan Wu, Yumeng Li, 等. LLM-Powered Predictive Decision-Making for Sustainable Data Center Operations[J/OL]. (2026-08-19)[2026-08-31]. http://arxiv.org/abs/2608.18503v1.
Research ArticleAI 运维优化
Kevin D. Gauld、Daniel J. Varon、Nicholas Balasus、Daniel H. Cusworth
Published 2026-08-23 · arXiv · Credibility S
AI data center power demand is spurring rapid deployment of on- and near-site natural gas turbines. Nitrogen oxide (NO$_x$) pollution from this equipment is a growing concern but has not previously been quantified with atmospheric observations. Here we demonstrate space-based detection and quantification of NO$_x$ emissions from the SpaceXAI Colossus 2 power plant in Southaven, Mississippi. Using observations from t…
Abstract, interpretation and reference
Abstract
AI data center power demand is spurring rapid deployment of on- and near-site natural gas turbines. Nitrogen oxide (NO$_x$) pollution from this equipment is a growing concern but has not previously been quantified with atmospheric observations. Here we demonstrate space-based detection and quantification of NO$_x$ emissions from the SpaceXAI Colossus 2 power plant in Southaven, Mississippi. Using observations from the geostationary TEMPO satellite instrument, we detect a strong increase in local mean NO$_2$ column concentrations after the plant began operations in late 2025. We then use TEMPO to estimate two-week-average NO$_x$ source rates from August 2025 to mid-August 2026, calibrating against continuous emission monitoring system (CEMS) data from US power plants. TEMPO first detected NO$_x$ emissions in December 2025 at 460$\pm$180 kg h$^{-1}$. We find that emissions increased through August 2026, averaging 730$\pm$185 kg h$^{-1}$ after February 2026, roughly 16 times higher than expected from the facility's March 2026 permit for 41 turbines operating under best available control technology (BACT) requirements ($\sim$47 kg h$^{-1}$). Emissions at the expected level would be undetectable by our TEMPO analysis.
中文解读
背景:AI 数据中心负载、功率密度和能源约束同步上升,AI 运维、负载预测和设施调优正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用仿真建模和情景分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断AI 工具是否能降低运维复杂度并提升可用性。仍需结合全文实验条件、样本范围和成本假设核验。
参考文献
Kevin D. Gauld, Daniel J. Varon, Nicholas Balasus, 等. Quantifying AI data center nitrogen oxide (NO$_x$) emissions from space[J/OL]. (2026-08-23)[2026-08-31]. http://arxiv.org/abs/2608.22153v1.
Research Article算电协同
Rouzbeh Haghighi、Ali Hassan、Sina Mohammadi、Marcus Chen I Wada、Wencong Su
Published 2026-08-05 · arXiv · Credibility S
Artificial intelligence (AI) data centers are driving rapid electricity load growth across all U.S. ISO/RTO regions, raising both system costs and carbon exposure. This study develops a three-level Stackelberg--Bayesian game in which a regulator (leader) sets carbon penalties and subsidies, a single ISO capacity market clears against an energy balance modeled as a classical generation-expansion problem, and technolo…
Abstract, interpretation and reference
Abstract
Artificial intelligence (AI) data centers are driving rapid electricity load growth across all U.S. ISO/RTO regions, raising both system costs and carbon exposure. This study develops a three-level Stackelberg--Bayesian game in which a regulator (leader) sets carbon penalties and subsidies, a single ISO capacity market clears against an energy balance modeled as a classical generation-expansion problem, and technology-specific investors (followers) decide capacity and operation under incomplete information, yielding a Bayesian Nash equilibrium. The AI impact is captured parsimoniously as an additional load-growth factor on a greenfield-incremental expansion, isolating how much new capacity the growth pulls in and which technology fills it. Within this framework, we consider second-life battery (SLB) storage competing against new/first-life storage for capacity-market revenue. We quantify how a carbon tax, a renewable subsidy, and an SLB subsidy reshape the equilibrium investment mix, carbon emissions, and profit. Different scenarios are compared at the end based on cost-effectiveness and reduced carbon emissions.
中文解读
背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用框架构建和频域/系统级分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向冷却效率、能源利用或运维策略的改进方向。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。
参考文献
Rouzbeh Haghighi, Ali Hassan, Sina Mohammadi, 等. A Stackelberg-Bayesian Capacity-Market Game of Carbon Regulation and Second-Life Battery Investment under AI Data-Center Load Growth[J/OL]. (2026-08-05)[2026-08-31]. http://arxiv.org/abs/2608.03989v1.
Research Article芯片与算力
Hao Chen、Zengqi Chen、Wu Zhou、Kaihang Lu、Mingyuan Zhang、Yuxiang Yin、Yiou Cui、Chaoran Huang
Published 2026-08-12 · arXiv · Credibility S
Increasing artificial intelligence (AI) workloads drive co-packaged optics (CPO), which integrates optical engines with electronic components. Optical interconnects can extend transmission distances and reduce latency, allowing distributed clusters in AI factories to operate as a unified computational unit. However, escalating data throughput necessitates greater parallelization of light within ultracompact form fac…
Abstract, interpretation and reference
Abstract
Increasing artificial intelligence (AI) workloads drive co-packaged optics (CPO), which integrates optical engines with electronic components. Optical interconnects can extend transmission distances and reduce latency, allowing distributed clusters in AI factories to operate as a unified computational unit. However, escalating data throughput necessitates greater parallelization of light within ultracompact form factors while maintaining stringent energy efficiency and latency constraints. Here, we present a multidimensional silicon photonic engine that achieves a communication capacity exceeding 1.8 terabit/s/lambda/s. By monolithically integrating transceivers, spatial and polarization (de)multiplexers, and optical signal processors on a single chip, we eliminate bulky discrete (de)multiplexers and power-hungry digital signal processing (DSP). In experiments, the photonic engine can be self-configured to identify two, four, or six concurrent spatial and polarization channels per fiber while mitigating dynamic channel crosstalk. Compared with the state-of-art DSP, our approach achieves >5,000-fold reductions in both power consumption and processing latency at a MIMO processing order of six. Furthermore, we demonstrate full-duplex, modulation-format-transparent inter-chip communication over 300-meter fiber. These results represent a paradigm shift for optical engines in future high-performance computing and AI-driven data centers.
中文解读
背景:AI 数据中心负载、功率密度和能源约束同步上升,芯片、服务器和高密度算力部署正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用框架构建和频域/系统级分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向能效评价口径、运营指标和优化目标的系统化梳理。意义:对日报读者而言,它可用于判断芯片路线和服务器密度变化如何传导到机房设计。仍需结合全文实验条件、样本范围和成本假设核验。
参考文献
Hao Chen, Zengqi Chen, Wu Zhou, 等. Towards Terabit/$λ$/s Multidimensional Silicon Photonic Engine[J/OL]. (2026-08-12)[2026-08-31]. http://arxiv.org/abs/2608.11639v1.
Research Article算电协同
Ze Yu、Hongwei Zhen、Chao Shen、Mingyang Sun
Published 2026-08-11 · arXiv · Credibility S
The rapid growth of large language model (LLM) services is expanding AI data centers (AIDCs), increasing electricity demand and associated carbon emissions. Renewable energy integration can mitigate these impacts but also strengthens the coupling between AIDC loads and inverter-interfaced generation, creating cross-domain cyber-physical vulnerabilities. Specifically, adversarial AI requests alter AIDC power demand, …
Abstract, interpretation and reference
Abstract
The rapid growth of large language model (LLM) services is expanding AI data centers (AIDCs), increasing electricity demand and associated carbon emissions. Renewable energy integration can mitigate these impacts but also strengthens the coupling between AIDC loads and inverter-interfaced generation, creating cross-domain cyber-physical vulnerabilities. Specifically, adversarial AI requests alter AIDC power demand, whereas inverter control tampering modifies source-side dynamics, and their combined impact on system stability varies with generation forecast and demand response uncertainties. To this end, we propose an uncertainty-aware AIDC microgrid vulnerability assessment framework under computing-power coordinated attacks. First, the framework maps adversarial AI requests to AIDC power variations and represents uncertainties in attack-induced demand responses and photovoltaic (PV) forecasts through confidence-weighted realizations. Then, impedance based stability analysis combines these realizations with bounded inverter parameter tampering to construct attack reachable domains and identify critical attack time windows. Furthermore, a separate criterion identifies fixed coordinated attack vectors that retain destabilizing capability throughout each selected window. Case studies demonstrate that, unlike either attack component applied alone, coordinated attacks within identified critical windows induce sustained inverter frequency oscillations with peak absolute deviations exceeding 20% of nominal frequency, whereas the evaluated out-of-window response remains bounded. The proposed method further identifies critical attack windows and the associated coordinated attack vectors.
中文解读
背景:AI 数据中心负载、功率密度和能源约束同步上升,算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题:论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法:摘要显示作者采用框架构建和频域/系统级分析,把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果:研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义:对日报读者而言,它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。
参考文献
Ze Yu, Hongwei Zhen, Chao Shen, 等. AIDC Microgrid Vulnerability Assessment Under Computing-Power Coordinated Attacks[J/OL]. (2026-08-11)[2026-08-31]. http://arxiv.org/abs/2608.10645v2.