Projects & Contributions

Open-source tools, active research projects, and community benchmarking efforts.

Agentic AI & Autonomous Science

🤖 Agentic AI for Autonomous Scientific Discovery

Status: Active • ALCF / DOE

A multi-agent LLM framework enabling autonomous end-to-end scientific workflows on DOE leadership computing facilities (Polaris, Aurora, Frontier, Perlmutter). The system uses specialized agents for job submission, software building, data staging, model serving, and result analysis — all orchestrated through a natural-language interface backed by MCP (Model Context Protocol) tools.

The framework integrates with ClearML for experiment tracking, Globus for data movement, and facility-specific APIs (ALCF IRI, NERSC IRI, OLCF S3M). Applications include training LLMs, running simulations, and driving iterative optimization loops without human-in-the-loop intervention.

💬 AskHPC

Status: Active • Published at SC'25

An LLM-powered chatbot for HPC user support at leadership computing facilities. Uses Retrieval-Augmented Generation (RAG) over facility documentation, software manuals, and accumulated user-support tickets to provide accurate, contextualized answers to user queries about system usage, job scheduling, software builds, and debugging.

🧐 AskALCF

Status: Active • ALCF

AI-powered question-answering service for ALCF users, providing instant and accurate answers about systems, software, job scheduling, and facility policies. Built on a curated knowledge base of ALCF documentation and user-support experience, AskALCF reduces the burden on support staff and helps users get unblocked faster.

HPC & Parallel I/O

📊 DLIO Benchmark

Status: Active • Used in MLPerf Storage

Deep Learning I/O (DLIO) is a benchmark suite for characterizing and optimizing storage and I/O performance for AI training workloads. It models realistic data ingestion pipelines for image classification, object detection, NLP, and scientific ML workflows.

DLIO is a core component of the MLPerf Storage benchmark and is used by storage vendors, HPC centers, and research groups worldwide to evaluate system readiness for AI workloads.

💾 HDF5 Cache VOL

Status: Active • ExaIO / ExaHDF5

A Virtual Object Layer (VOL) plugin for HDF5 that enables transparent, asynchronous caching of dataset writes and reads on node-local NVMe storage. The Cache VOL reduces contention on shared parallel file systems and can deliver up to 10× I/O speedup for checkpoint-heavy workloads.

📊 h5bench

Status: Active • HPC-IO

A unified benchmark suite for evaluating HDF5 I/O performance across diverse access patterns — write, read, and metadata operations — on pre-exascale and exascale platforms. Supports multiple VOL plugins including the Cache VOL and async VOL.

Scientific AI & Machine Learning

🌎 AuroraGPT

Status: Active • ALCF / DOE

Training large language models on the Aurora exascale supercomputer (Intel Xe GPUs) for scientific applications. AuroraGPT targets domain-specific LLMs in materials science, climate, biology, chemistry, and energy, leveraging Aurora's 10+ ExaFLOPS peak performance.

💰 MLPerf Storage Benchmark

Status: Active • MLCommons

Co-leading the MLPerf Storage working group to develop community-standardized benchmarks for evaluating the performance of storage systems under real AI training workloads. The benchmark covers data loading bottlenecks, metadata performance, and scalability across different storage architectures.

Scientific Imaging & Tomography

🔬 Tomo_TV & tomviz

Status: Active • University of Michigan / ALCF

Real-time 3D analysis algorithms for electron tomography experiments. Combines dynamic compressed sensing with GPU-accelerated reconstruction to enable live structural analysis during experiments. The tomviz platform provides an interactive environment for materials scientists, published in Nature Communications (2022, 67 citations).