Skip to main content
Seminar | Computing, Environment and Life Sciences

DeepVELA: Agentic AI for High Performance Computing Software, Data and Workloads in the Genesis Mission

TPC Seminar

Abstract: I will describe The University of Virginia’s (UVA) part of DeepVELA, a Genesis Phase I project led by Matt Sinclair (University of Wisconsin–Madison) with Oak Ridge National Laboratory, Fermilab and UVA. DeepVELA uses agentic AI for both the software engineering and the data and process management that scientific computing on high performance computing (HPC) systems needs, applied to Genesis and MLCommons codes.

Its workloads range from small manufacturing-control time series on a few nodes to the largest U.S. Department of Energy simulations. Energy use is a first-class performance measure alongside time to solution. Specialized agents will be trained with reinforcement learning.

Underneath is a knowledge graph built on fifteen linked ontologies: time-series, simulation, code-task and measured science datasets; AI models, time-series models, machine learning surrogates and models of measured data; benchmarks of every kind; parallel programming environments and HPC systems; linear algebra and kernels; AI motifs and task types; metadata standards; and workflows and complete workloads. Each defines classes (entity types such as dataset, model, benchmark, software and system, and taxonomies of methods), their attributes with controlled vocabularies, and typed relations with their inverses. Each is reviewed in turn by several frontier language learning models (LLMs) and tested by filling in real entries. About half are built; the rest are partly built or planned.

Every value in the graph is a claim that cites its source and carries a confidence, and the graph is updated as new evidence arrives. Values are found by cellular retrieval‑augmented generation: one targeted retrieval for each attribute of each entity, that is, each cell of the table. Each retrieval either finds a value with its evidence or records that no source states one. We compare local open-weight and external frontier LLMs, both for discovering entities and for filling in their attributes. Because entries from different disciplines are described by the same ontologies, the graph can act as a recommender. For a new problem, it finds benchmarks, datasets, models and codes in other fields that use similar methods, so a benchmark’s value carries across disciplines.

Bio: Geoffrey Fox received a Ph.D. in Theoretical Physics from Cambridge University, where he was Senior Wrangler. He is now a professor at the Biocomplexity Institute and Computer Science department at the University of Virginia. He received the High-Performance Parallel and Distributed Computing (HPDC) Achievement Award and the Association for Computing Machinery-Institute of Electrical and Electronics Engineers Computer Society Ken Kennedy Award for Foundational contributions to parallel computing in 2019. He is currently active in the industry consortium MLCommons/MLPerf.