About
📋 Biography
Sandeep Madireddy is a Computer Scientist in the Mathematics and Computer Science Division at Argonne National Laboratory. His research spans probabilistic machine learning, bio-inspired and energy-efficient learning, high-performance computing, and generative AI with an emphasis on safety and robustness. His work integrates algorithmic advances with applied research across fusion energy sciences, cosmology and high-energy physics, weather and climate, and material science.
He previously served as a postdoc and assistant computer scientist in Mathematics and Computer Science Division, advised by Prasanna Balaprakash and Stefan Wild. Before joining Argonne, he earned his Ph.D. in mechanical and materials engineering (probabilistic machine learning) from the University of Cincinnati as part of the UC Simulation Center (a Procter & Gamble collaboration). He previously earned his master's at Utah State University and his bachelor's from BITS Pilani, India.
🔬 Projects & Grants
Sandeep serves as a Co-investigator (and AI lead) for DOE and NSF-funded projects:
- ModCon/Genesis Mission, Lead for Multimodal Scientific AI Reasoning Models BASE Capability Thrust
- AuroraGPT, Co-lead for Evaluation and AI-Safety Thrust
- RAPIDS3: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence, AI thrust co-lead
- CETOP: A Center for Edge of Tokamak OPtimization, Co-investigator and AI Lead
📝 Professional Service
Sandeep also provides professional services to machine learning, high-performance computing, and domain-science conferences and journals.
- PC member (HPC): Supercomputing 2023 and 2026, CCGRID 2024, INFOCOMP 2019-2023; reviewer for IPDPS and IEEE Cluster
- Reviewer (AI): ICML (2021-2024), ICLR (2021-2024), NeurIPS (2021-2025), AISTATS (2023, 2025)
- Journals: Nature, Neural Networks, NeuroComputing, JMLMC, SIAM SISC, JPDC, TPDS, Parallel Computing, TCC, JoS, MNRAS, CISE, IEEE TPSC, and JoLT
Research Areas
Probabilistic Machine Learning
Research at the intersection of Bayesian Inference and Information-Theoretic learning.
Scientific Machine Learning and AI for Science
Theoretical and Applied Machine learning for Large-scale Science including Fusion Energy, Astrophysics, HPC, and Climate Science
Multi-modal Foundation Models
Theoretical and Applied research on developing and evaluating (for skill and safety) Multimodal Language Models as scientific research assistants and Spatio-Temporal Foundation models for science.
Energy Efficient AI
Energy-efficient AI with Bio-inspired architectures such as Neuromorphics for large scale AI models and life-long learning in this paradigm.
News & Announcements
Presented an invited talk on multimodal reasoning models for science at the 2026 Monterey Data Conference. View the LinkedIn post.
Presented an invited talk on flow-based generative models for scientific applications at the Artificial Intelligence for Robust Engineering and Science (AIRES) 7 Workshop. View the LinkedIn post.
Senior Personnel (ANL), Lab 25-3560, (Lead PI: Rick Stevens, Argonne National Laboratory), FY26-FY28
Co-PI (ANL), Foundational Models for Energy Applications, LDRD Prime Focus Area, (Lead-PI: Pinaki Pal, Argonne National Laboratory), FY26-FY28
Co-PI (ANL), LAB 25-3510.000002, (Lead PI: Robert Ross, Argonne National Laboratory), FY26-FY31
Publications
2026
Composing Flow-Matching Energies with Known Physics: Generation, OOD Detection, and Inversion on PDE Fields
PreprintYixuan Sun, Anirban Samaddar, Sandeep Madireddy
arXiv preprint arXiv:2608.18004, 2026
This work shows that potential-parameterized flow-matching models expose explicit time-dependent scalar energies while retaining standard regression training and flow-ODE sampling. The energies enable physics-corrected generation, out-of-distribution detection, and posterior sampling for PDE inverse problems.
SURE: data-efficient safety guardrailing via internal representations and uncertainty-weighted pseudo-labels
JournalHanwen Li, Jinhao Duan, Chenxi Yuan, …
npj Artificial Intelligence, 2026
SURE trains a lightweight external classifier on LLM hidden-state representations, using uncertainty-weighted pseudo-labels to detect unsafe inputs without modifying the model's inference path. It achieves a strong safety-performance balance with limited labeled data and generalizes to held-out hazard categories.
GLANCED-IO: Taming I/O Optimization for Deep Learning at Scale
ArticleRAO Sinurat, W Nixon, P Carns, …
Scientific deep learning at scale faces significant challenges in optimizing I/O configurations due to costly experiments and complex parameter spaces, leading to suboptimal training efficiency. GLANCED-IO, a cross-layer I/O optimization framework, addresses this by combining high-fidelity approximation with efficient configuration exploration, significantly improving performance beyond independent application or system tuning.
KORAL: Knowledge Graph Guided LLM Reasoning for SSD Operational Analysis
PreprintM Akewar, S Madireddy, D Luo, …
arXiv preprint arXiv:2602.10246, 2026
KORAL is a knowledge-driven reasoning framework that combines Large Language Models with a structured Knowledge Graph to analyze SSD performance and reliability using fragmented telemetry and existing literature. This approach enables explainable, evidence-based insights without requiring large datasets or extensive expert input, effectively addressing challenges posed by workload shifts, architectural changes, and environmental factors.
Multi-task Modeling for Engineering Applications with Sparse Data
PreprintY Comlek, RM Krishnan, SK Ravi, …
arXiv preprint arXiv:2601.05910, 2026
This paper presents a Multi-Task Gaussian Processes (MTGP) framework designed for engineering systems with multi-source, multi-fidelity data, addressing issues of data scarcity and varying task correlations. By leveraging inter-task relationships, the framework improves prediction accuracy and reduces computational costs, as demonstrated through benchmarks including the Forrester function, 3D void modeling, and friction-stir welding scenarios.
Uncovering Physical Drivers of Dark Matter Halo Structures with Auxiliary-Variable-Guided Generative Models
PreprintA Ganguli, A Samaddar, F Kéruzoré, …
arXiv preprint arXiv:2602.23518, 2026
This research introduces an auxiliary-variable-guided framework to disentangle representations of thermal Sunyaev-Zel’dovich maps by using halo mass and concentration as physical factors in the latent space. Their Disentangled Latent-CFM model preserves generative flexibility while linking latent dimensions to interpretable astrophysical properties, enabling recovery of established mass-concentration relations and identification of unusual halo formations.
2025
Aeris: Argonne earth systems model for reliable and skillful predictions
ConferenceV Hatanpää, E Ku, J Stock, …
Proceedings of the International Conference for High Performance Computing …, 2025
This research introduces AERIS, a large-scale pixel-level Swin diffusion transformer, and SWiPe, a technique for efficient parallelism in window-based transformers, to improve weather forecasting with diffusion models. AERIS achieves state-of-the-art performance and scalability on the ERA5 dataset, outperforming existing methods and maintaining stability for seasonal-scale predictions up to 90 days.
Ailuminate: Introducing v1. 0 of the ai risk and reliability benchmark from mlcommons
PreprintS Ghosh, H Frase, A Williams, …
arXiv preprint arXiv:2503.05731, 2025
The paper presents AILuminate v1.0, the first comprehensive industry-standard benchmark designed to evaluate AI systems' resistance to prompts eliciting dangerous, illegal, or undesirable behaviors across 12 hazard categories. This benchmark features a novel evaluation framework, extensive prompt datasets, a five-tier grading scale, and an entropy-based response assessment, supported by infrastructure for ongoing development and cross-disciplinary input.
★AstroMLab 1: Who wins astronomy jeopardy!?
ArticleYS Ting, TD Nguyen, T Ghosal, …
Astronomy and Computing 51, 100893, 2025
This study evaluates large language models on a novel astronomy-specific benchmark of 4,425 multiple-choice questions, revealing varied performance across subfields and highlighting response calibration for research use. Claude-3.5-Sonnet leads with 85.0% accuracy, while open-weight models like LLaMA-3-70b and Qwen-2-72b show rapid improvement, though non-English-focused models face challenges in certain astrophysical topics.
AstroMLab 1
ArticleYS Ting, TD Nguyen, T Ghosal, …
This study evaluates large language models using a novel astronomy-specific benchmark of 4,425 multiple-choice questions across diverse astrophysical topics, revealing performance differences and response calibration critical for research use. Claude-3.5-Sonnet leads with 85.0% accuracy, while open-weight models like LLaMA-3-70b have significantly improved, though challenges remain in less represented areas such as exoplanets and instrumentation due to limited training data.
AuroraGPT Data Collection Interface
ArticleR Underwood, A Maurya, Z Li, …
Argonne National Laboratory (ANL), Argonne, IL (United States), 2025
The software package enables efficient data collection to evaluate the performance of large language models (LLMs) on scientific topics. It is designed to support comprehensive assessment, ensuring accurate measurement of LLM capabilities in understanding and processing scientific information.
Automated MCQA Benchmarking at Scale: Evaluating Reasoning Traces as Retrieval Sources for Domain Adaptation of Small Language Models
ConferenceO Gokdemir, N Getty, R Underwood, …
Proceedings of the SC'25 Workshops of the International Conference for High …, 2025
The authors present a scalable, modular framework that automates the creation of multiple-choice question-answering benchmarks from large scientific corpora, demonstrated by generating over 16,000 questions from 22,000 radiation and cancer biology papers. Evaluating small language models on these questions reveals that retrieval-augmented generation using reasoning traces significantly improves performance, enabling several smaller models to outperform GPT-4 on a specialized 2023 exam.
Can LLMs Model the Environmental Impact on SSD?
ConferenceM Akewar, G Quan, S Madireddy, …
Proceedings of the 17th ACM Workshop on Hot Topics in Storage and File …, 2025
Environmental stressors significantly affect SSD performance and reliability, but studying these effects is difficult due to limited data, complex interrelated factors, and the specialized equipment required. Existing storage management techniques neglect environmental impacts, and accurately modeling these effects remains challenging because of varied NAND flash responses and the cascading influence of historical exposure.
Chance-constrained Flow Matching for High-Fidelity Constraint-aware Generation
PreprintJ Liang, Y Sun, A Samaddar, …
arXiv preprint arXiv:2509.25157, 2025
This paper introduces Chance-constrained Flow Matching (CCFM), a training-free method that incorporates stochastic optimization during sampling to enforce hard constraints while preserving high-fidelity generation. CCFM guarantees feasibility comparable to repeated projection but operates on noisy intermediates, reducing distributional distortion and complexity.
Data-Efficient Dimensionality Reduction and Surrogate Modeling of High-Dimensional Stress Fields
JournalA Samaddar, SK Ravi, N Ramachandra, …
Journal of Mechanical Design 147 (3), 031701, 2025
This study presents a convolutional autoencoder framework enhanced by an information bottleneck loss to model tensor datatypes in simulations under limited data conditions, improving robustness by filtering nuisance information while retaining relevant features. The approach is combined with hyperparameter optimization and is numerically compared with dimensionality reduction-based surrogate modeling methods to demonstrate improved predictive performance for engineering applications.
Eaira: Establishing a methodology for evaluating ai models as scientific research assistants
PreprintF Cappello, S Madireddy, R Underwood, …
arXiv preprint arXiv:2502.20309, 2025
This paper presents EAIRA, a comprehensive methodology developed at Argonne National Laboratory for evaluating Large Language Models as scientific research assistants through multiple choice questions, open responses, lab-style experiments, and field-style experiments. These evaluations provide a holistic assessment of LLMs' factual recall, reasoning, problem-solving, and real-world research interactions across diverse scientific domains.
Efficient Flow Matching using Latent Variables
PreprintA Samaddar, Y Sun, V Nilsson, …
arXiv preprint arXiv:2505.04486, 2025
Flow matching models typically overlook the clustering structure in target data, leading to inefficient learning, especially for high-dimensional datasets on low-dimensional manifolds. Latent-CFM addresses this by conditioning on pretrained deep latent features, achieving better generation quality with less training and computation across synthetic, image, and physical spatial field datasets.
Ensembles of Neural Surrogates for Parametric Sensitivity in Ocean Modeling
PreprintY Sun, R Egele, SHK Narayanan, …
arXiv preprint arXiv:2508.16489, 2025
This research improves ocean simulation accuracy by using deep learning surrogates combined with large-scale hyperparameter search and ensemble learning to enhance forward predictions and sensitivity estimations. The ensemble approach also quantifies epistemic uncertainty, increasing the reliability of neural surrogates for parameter tuning and decision making.
Evaluating the Safety and Skill Reasoning of Large Reasoning Models Under Compute Constraints
PreprintA Balaji, L Chen, R Thakur, …
arXiv preprint arXiv:2509.18382, 2025
This work investigates reducing the computational cost of reasoning language models by using reasoning length constraints and model quantization, focusing on their effect on safety performance. It proposes fine-tuning with length controlled policy optimization and quantization to balance computational efficiency with model safety within user-defined compute limits.
Evaluation of Test-Time Compute Constraints on Safety and Skill Large Reasoning Models
ConferenceA Balaji, L Chen, R Thakur, …
Proceedings of the SC'25 Workshops of the International Conference for High …, 2025
This research examines two compute constraint strategies—reasoning length constraint and model quantization—to manage the trade-off between computational efficiency and safety in reasoning language models. It proposes fine-tuning models with a length-controlled policy optimization method and applying quantization to maximize chain-of-thought generation within user-defined compute limits.
Talks
Multimodal Reasoning Models For Science
Monterey Data Conference
Flow-based Generative Models for Science
Artificial Intelligence for Robust Engineering and Science (AIRES) 7 Workshop
Multimodal Scientific AI
18th Joint Laboratory on Extreme Scale Computing Workshop
Language Model Evaluation and Safety for Scientific Tasks
Argonne Training Program on Extreme-Scale Computing
Half Day Tutorial: Evaluation of AI Model Scientific Skills
Trillion Parameter Consortium's 2025 all-hands conference and exhibition
Projects
Team
Post Doctoral Researchers
No current members listed.
Summer Interns
No current members listed.








