Huihuo Zheng

Huihuo Zheng

Computer Scientist  ·  AI/ML Group

Argonne National Laboratory — Lemont, IL

Agentic AI Large-scale Distributed Training Data Management for AI HPC & Parallel I/O
2,256
Citations
20
h-index
25
i10-index
15+
Years

I am a Computer Scientist in the AI/ML Group at Argonne National Laboratory (ALCF). My research interests include agentic AI systems for autonomous scientific discovery, large-scale distributed training, and data management for AI. I apply high-performance computing and deep learning to domain sciences including physics, chemistry, and materials science. I co-lead the MLPerf Storage Benchmarking group, developing benchmark suites for evaluating the performance of storage systems for AI applications.

Ph.D. in Physics, University of Illinois at Urbana-Champaign (2016) • M.Phil., Hong Kong University of Science and Technology (2010) • B.Sc., Tsinghua University (2008)

News & Updates

Research Areas

Agentic AI

Agentic AI for Scientific Discovery

Designing multi-agent systems and LLM-driven workflows that autonomously plan, execute, and reason over scientific workloads at DOE leadership computing facilities.

HPC & Parallel I/O

HPC & Parallel I/O

Advanced HDF5 features (Cache VOL, node-local storage, topology-aware collective I/O), storage benchmarking (MLPerf Storage, h5bench, DLIO), and exascale I/O optimization.

Scientific AI

Scientific AI & Deep Learning

Large-scale distributed training (AuroraGPT), LLMs for science, AI-accelerated electron tomography, gravitational wave detection, and drug/materials discovery.

Computational Physics

Computational Physics

Quantum Monte Carlo, density functional theory, many-body perturbation theory, and dielectric-dependent hybrid functionals for heterogeneous materials.

Selected Publications

Sorted by year (most recent first). For full list see Publications or Google Scholar.

LLM for Batteries 2025
Large language models for batteries 📚 20
W. Zuo, H. Zheng, T. He, V. Vishwanath, M. K. Y. Chan, R. L. Stevens, K. Amine, et al.
Joule 9(8), 2025
LLMs can reason over electrochemical literature to autonomously propose, evaluate, and prioritize battery material hypotheses, suggesting a new role for AI in accelerating energy storage research.
HiPerRAG 2025
HiPerRAG: High-Performance Retrieval Augmented Generation for Scientific Insights 📚 19
O. Gokdemir, C. Siebenschuh, A. Brace, A. Wells, H. Zheng, et al.
PASC Conference, 2025
HiPerRAG achieves high-throughput scientific document retrieval by co-designing the retrieval pipeline with HPC storage, enabling exascale-grade RAG for large scientific corpora.
AskHPC 2025
AskHPC: A ChatBot for High Performance Computing User Support 📚 1
A. Bondapalli, H. Zheng, O. T. Ajayi, M. Keceli, et al.
SC’25 Workshops, 2025
An LLM-powered chatbot built on RAG over ALCF documentation answers HPC user queries about system usage, job scheduling, and software debugging with high accuracy.
tomviz tomography 2022
Real-time 3D analysis during electron tomography using tomviz 📚 67
J. Schwartz, R. Harris, J. Pietryga, H. Zheng, P. Kumar, et al.
Nature Communications 13, 4458, 2022
Real-time 3D tomographic reconstruction during live electron microscopy lets scientists make mid-experiment decisions at near-atomic resolution, eliminating the traditional hours-long post-acquisition wait.
Gravitational wave AI 2022
Inference-optimized AI and high performance computing for gravitational wave detection at scale 📚 39
P. Chaturvedi, A. Khan, M. Tian, E. A. Huerta, H. Zheng
Frontiers in Artificial Intelligence 5:828672, 2022
Inference-optimized deep learning on HPC systems achieves real-time gravitational wave detection throughput, enabling continuous monitoring of LIGO data at production scale.
HDF5 Cache VOL 2022
HDF5 Cache VOL: Efficient and Scalable Parallel I/O through Caching Data on Node-local Storage 📚 21
H. Zheng, V. Vishwanath, Q. Koziol, H. Tang, J. Ravi, J. Mainzer, S. Byna
IEEE CCGrid, 2022
A transparent HDF5 caching plugin stages I/O on node-local NVMe SSDs, delivering up to 10× speedup for checkpoint-intensive HPC applications with no application code changes required.
DLIO benchmark 2021
DLIO: A Data-Centric Benchmark for Scientific Deep Learning Applications 📚 54
H. Devarajan, H. Zheng, A. Kougkas, X-H. Sun, V. Vishwanath
IEEE/ACM CCGrid, 2021
DLIO faithfully reproduces the I/O patterns of scientific deep learning workloads, giving storage vendors and HPC centers a reproducible, configurable tool for measuring AI storage performance.
Galaxy catalogs DES 2019
Deep learning at scale for the construction of galaxy catalogs in the Dark Energy Survey 📚 66
A. Khan, E. A. Huerta, S. Wang, R. Gruendl, E. Jennings, H. Zheng
Physics Letters B 795, 248–258, 2019
Deep learning galaxy classification scales to 2,048 GPUs on Blue Waters, enabling petabyte-scale Dark Energy Survey catalog construction without expensive human labeling.
Dielectric hybrid DFT 2019
Dielectric dependent hybrid functionals for heterogeneous materials 📚 80
H. Zheng, M. Govoni, G. Galli
Physical Review Materials 3, 073803, 2019
A position-dependent hybrid functional adapts to the local dielectric environment in heterogeneous systems, achieving GW-level band-gap accuracy at DFT cost.
View All Publications →

Projects & Contributions

🤖 Agentic AI for Scientific Discovery

Multi-agent LLM framework for autonomous scientific workflows on DOE leadership computing facilities. Integrates ClearML, Globus, and facility APIs for end-to-end automation.

LLMMulti-AgentALCFMCP

📊 DLIO Benchmark

Deep Learning I/O benchmark for characterizing and optimizing storage performance for AI training workloads. Used in the MLPerf Storage benchmark suite.

BenchmarkHPC I/OMLPerf

💾 HDF5 Cache VOL

Virtual Object Layer plugin for HDF5 enabling transparent caching on node-local NVMe storage, delivering up to 10× I/O acceleration for parallel scientific applications.

HDF5Parallel I/OExascale

📊 h5bench

Unified benchmark suite for evaluating HDF5 I/O performance patterns on pre-exascale and exascale platforms. Covers diverse access patterns and VOL plugins.

HDF5BenchmarkStorage

🌎 AuroraGPT

Training large language models on the Aurora exascale supercomputer for science. Targeting domain-specific LLMs for materials, biology, climate, and energy research.

LLMAuroraExascaleAI for Science

💰 MLPerf Storage

Co-leading the MLPerf Storage working group to develop community benchmarks for evaluating storage system performance under real AI training workloads.

BenchmarkStorageMLPerf

🔬 Tomo_TV / tomviz

Tomographic reconstruction algorithms with real-time 3D analysis during electron tomography experiments. Enables dynamic compressed sensing at atomic resolution.

TomographyReal-timeMaterials

💬 AskHPC

LLM-powered chatbot for HPC user support at leadership computing facilities. Integrates facility documentation and knowledge bases via RAG. Published at SC’25.

LLMRAGUser Support

🧐 AskALCF

AI-powered question-answering service for ALCF users, providing instant answers about systems, software, job scheduling, and facility policies using curated HPC documentation.

LLMRAGALCFUser Support
View All Projects →

Selected Talks

All Talks →

Contact

🏠

Office

Room 3117, Building 240
Argonne Leadership Computing Facility
Argonne National Laboratory
Lemont, IL 60439

📞

Phone

630-252-8643

🎓

Scholar

Google Scholar Profile

Prospective collaborators & students: I am always interested in collaborations at the intersection of HPC, scientific AI, and autonomous workflows. Please send an email with a brief description of your interests and your CV. See the About page for more details.