Open-source tools, active research projects, and community benchmarking efforts.
A multi-agent LLM framework enabling autonomous end-to-end scientific workflows on DOE leadership computing facilities (Polaris, Aurora, Frontier, Perlmutter). The system uses specialized agents for job submission, software building, data staging, model serving, and result analysis — all orchestrated through a natural-language interface backed by MCP (Model Context Protocol) tools.
The framework integrates with ClearML for experiment tracking, Globus for data movement, and facility-specific APIs (ALCF IRI, NERSC IRI, OLCF S3M). Applications include training LLMs, running simulations, and driving iterative optimization loops without human-in-the-loop intervention.
An LLM-powered chatbot for HPC user support at leadership computing facilities. Uses Retrieval-Augmented Generation (RAG) over facility documentation, software manuals, and accumulated user-support tickets to provide accurate, contextualized answers to user queries about system usage, job scheduling, software builds, and debugging.
AI-powered question-answering service for ALCF users, providing instant and accurate answers about systems, software, job scheduling, and facility policies. Built on a curated knowledge base of ALCF documentation and user-support experience, AskALCF reduces the burden on support staff and helps users get unblocked faster.
Deep Learning I/O (DLIO) is a benchmark suite for characterizing and optimizing storage and I/O performance for AI training workloads. It models realistic data ingestion pipelines for image classification, object detection, NLP, and scientific ML workflows.
DLIO is a core component of the MLPerf Storage benchmark and is used by storage vendors, HPC centers, and research groups worldwide to evaluate system readiness for AI workloads.
A Virtual Object Layer (VOL) plugin for HDF5 that enables transparent, asynchronous caching of dataset writes and reads on node-local NVMe storage. The Cache VOL reduces contention on shared parallel file systems and can deliver up to 10× I/O speedup for checkpoint-heavy workloads.
A unified benchmark suite for evaluating HDF5 I/O performance across diverse access patterns — write, read, and metadata operations — on pre-exascale and exascale platforms. Supports multiple VOL plugins including the Cache VOL and async VOL.
Training large language models on the Aurora exascale supercomputer (Intel Xe GPUs) for scientific applications. AuroraGPT targets domain-specific LLMs in materials science, climate, biology, chemistry, and energy, leveraging Aurora's 10+ ExaFLOPS peak performance.
Co-leading the MLPerf Storage working group to develop community-standardized benchmarks for evaluating the performance of storage systems under real AI training workloads. The benchmark covers data loading bottlenecks, metadata performance, and scalability across different storage architectures.
Real-time 3D analysis algorithms for electron tomography experiments. Combines dynamic compressed sensing with GPU-accelerated reconstruction to enable live structural analysis during experiments. The tomviz platform provides an interactive environment for materials scientists, published in Nature Communications (2022, 67 citations).