The US Department of Energy (DOE) Argonne National Laboratory has launched what it describes as the first large-scale artificial intelligence (AI) inference service for open science, giving researchers cloud-like access to large language models and domain-specific foundation models running directly on high-performance computing (HPC) systems.
The new service is designed to help scientists analyze massive datasets, test hypotheses, and accelerate research workflows without the financial and operational burden of building and maintaining dedicated AI infrastructure.
Bridging AI model development and scientific research
The Argonne Leadership Computing Facility (ALCF) Inference Service aims to bridge the gap between AI model development and real-world scientific applications. In scientific computing, inference refers to the process of using trained AI models to identify patterns, analyze data, and generate predictions from new information.
“Our inference service helps close the gap between developing AI models and putting them to work in scientific research,” says Michael Papka, PhD, director of the ALCF. “By offering AI inference as a shared resource, we enable researchers to apply AI at scale to their data, simulations, and experiments, without the burden of building and maintaining their own infrastructure.”
Rather than spending weeks configuring local hardware and software environments, researchers can use the service to rapidly interpret experimental results and refine active hypotheses. The platform currently operates on ALCF computing systems, including Sophia and Metis, with support planned for additional NVIDIA-based systems such as Tara and Minerva.
Built for distributed HPC environments
The ALCF Inference Service is built on a framework designed to provide secure, scalable AI inference across distributed HPC environments. The platform supports commercial and open-weight model families, including Google Gemma, Meta Llama, and OpenAI’s GPT-OSS models, alongside specialized in-house systems developed at Argonne, such as AuroraGPT.
The service also incorporates Globus Compute and Globus Auth technologies to support federated access and distributed inference workflows across institutions.
Authorized researchers can access the platform using credentials from their home institutions—including universities, industry partners, and national laboratories such as Brookhaven National Laboratory, Lawrence Berkeley National Laboratory, and Oak Ridge National Laboratory.
Supporting fusion energy and chemistry research
The infrastructure is already supporting research initiatives, including the DOE Genesis Mission. Argonne says the platform can support fusion energy applications in which AI models analyze experimental data streams and help predict plasma disruptions. In chemistry and materials science, tools such as ChemGraph are using the system to automate molecular simulation workflows.
Reducing the cost of AI research workflows
For research organizations evaluating the rising costs of AI-driven computational science, the service addresses a major operational challenge: token consumption and infrastructure costs. Advanced agentic research workflows often require repeated communication between AI models and simulation software through a process known as tool calling.
According to Murat Keçeli, PhD, a computational scientist at Argonne who helped develop ChemGraph, these workflows process information in units known as tokens. Because agentic applications can consume extremely high token volumes, running them through commercial cloud AI platforms can quickly become prohibitively expensive. Using a shared, institution-backed inference service helps reduce those external costs.
By leveraging centralized federal computing infrastructure, laboratories and research institutions can avoid major capital expenditures associated with purchasing high-end graphics processing units (GPUs), maintaining complex software stacks, and paying recurring subscription fees to commercial AI vendors. The shift from providing raw compute power to delivering integrated AI-enabled scientific services could allow research organizations to direct more funding toward experiments, materials development, and staffing.
This article was created with the assistance of Generative AI and has undergone editorial review before publishing.








