MCS Menu
Large language models (LLMs) are rapidly becoming capable of acting as “AI scientists,” helping manage many stages of the research process — from experiment design and data analysis to interpreting results and drafting publications. As these systems take on larger roles in scientific discovery, ensuring their reliability and security becomes increasingly important.
One challenge is that LLMs can be vulnerable to adversarial attacks — carefully crafted inputs designed to manipulate model behavior. In scientific settings, such attacks could lead to inaccurate findings, unsafe recommendations or reduced trust in AI-generated results.
Researchers have developed adversarial robustness benchmarks to measure how well AI systems withstand these attacks. Most existing benchmarks were designed for general-purpose AI applications, however, and often overlook the unique challenges of scientific research. As a result, scientists lack effective tools for evaluating vulnerabilities in AI systems used for research.
To address this gap, researchers at the U.S. Department of Energy’s Argonne National Laboratory have developed a two-part framework: a multi-agent mechanism for generating scientific security and a multilayered defense architecture designed to protect scientific AI systems from malicious threats. Their work was presented at the Trillion Parameter Consortium 2026 workshop in June 2026 by Saket Chaturvedi, a postdoctoral researcher at Argonne.
Multiple agents, specialized roles
Most current approaches for generating adversarial benchmarks rely on a single-agent system. While efficient, such a system can create competing responsibilities, with one model acting simultaneously as domain expert, attacker and evaluator.
The solution proposed by the Argonne researchers involves a collaborative system of specialized agents, each with a distinct role:
-
Orchestration agent — coordinates the process and works with scientific experts to identify realistic vulnerabilities.
-
Domain agent — contributes scientifically grounded concepts and threat scenarios.
-
Adversary agent — creates targeted attack prompt using domain-specific knowledge.
-
Refiner agent — reviews prompts for scientific plausibility and works with the adversary agent to improve them.
-
Quality control agent — removes redundancies and validates results across multiple LLMs.
“Using multiple specialized agents is important,” Chaturvedi said. “It reduces the limitations of single-agent systems by dividing benchmark creation into focused tasks that can be handled more effectively.”
A key feature of the framework is iterative refinement. If a generated prompt does not meet quality standards, it is returned to the appropriate agent with feedback for improvement. This process repeats until the prompt is accepted for inclusion in the benchmark.
A multilayered approach to defense
Generating scientific security benchmarks is only part of the solution. In addition to evaluating LLM threats, researchers want to defend against threats. To this end, the Argonne team designed a multilayered defense architecture in which (1) a red teaming layer continuously tests the system with automated adversarial attacks, (2) an internal safety layer incorporates features such as safety-aligned LLMs to protect communication between agents, and (3) an external safety layer provides boundary controls and additional protections against outside threats.
“Together, these layers help address the different pathways an attacker might try to exploit,” said Joshua Bergerson, Infrastructure Security & Risk Analytics Team Leader at Argonne and coauthor of the study.
Building the next generation of secure scientific AI
The team is currently developing a prototype implementation of both the multi-agent benchmark generator and the three-layer defense system.
“The framework provides a foundation for secure scientific AI,” said Tanwi Mallick, a computer scientist at Argonne and coauthor of the study.
The work also opens the door to important future research questions. For example, can adversarial agents trained in one scientific domain be adapted to others? How can AI systems balance robust security measures with the need for rapid responses in an emergency? And what is the most effective way for human experts to provide feedback within automated workflows?
“By addressing these questions, we can move closer to realizing the potential of LLMs as safe, reliable and transformative partners in scientific discovery,” Mallick said.
Additional details about the framework are available in the preprint “Toward Reliable, Safe, and Secure LLMs for Scientific Applications” by S. Chaturvedi, J. Bergerson, and T. Mallick, https://arxiv.org/pdf/2603.18235
Argonne National Laboratory seeks solutions to pressing national problems in science and technology by conducting leading-edge basic and applied research in virtually every scientific discipline. Argonne is managed by UChicago Argonne, LLC for the U.S. Department of Energy’s Office of Science.
The U.S. Department of Energy’s Office of Science is the single largest supporter of basic research in the physical sciences in the United States and is working to address some of the most pressing challenges of our time. For more information, visit https://energy.gov/science.