Latest Adversarial Robustness Research Papers
The newest Adversarial Robustness papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Adversarial Robustness so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Adversarial Robustness papers in your inbox — free →Recent papers
- BlueSTAR: Tiered Agentic Architecture for Autonomous Cyber DefenseSimona Boboila, Xavier Cadet, Edward Koh, Daniel Balasubramanian et al. · arXiv · Sep 10, 2026
Cyber attacks are increasingly automated, narrowing the time available for human analysts to detect, reason about, and respond to intrusions. Large language models (LLMs) offer a promising foundation for autonomous cyber defense because the…
- SpecGuard: Inference-Time Backdoor Detection For FreeRui Wen, Ahmed Salem, Andrew Paverd, Mark Russinovich et al. · arXiv · Sep 10, 2026
Large language models are often fine-tuned, shared, or downloaded from third parties, so a deployed model may carry a hidden backdoor that behaves normally on benign inputs but switches to attacker-controlled behavior when a secret trigger …
- Predicting Privacy Leakage from Weight Spectral DensityRichard J. Preen, Jim Smith · arXiv · Sep 10, 2026
Membership inference attacks (MIAs) are widely used to audit the privacy disclosure risk of machine learning models, however current state-of-the-art attacks require training computationally expensive shadow models, making large-scale priva…
- Signing the Transaction but Not the Decision: Whisper Attacks and a Binding Defense for AP2Yedidel Louck, Amit Dvir, Ariel Stulman · arXiv · Sep 10, 2026
Software agents are beginning to shop and pay on a person's behalf. Agent payment protocols such as AP2 produce cryptographically valid signatures for completed purchases, yet do not constrain the decisions that lead to them. Consequently, …
- Certifying Adversarial Robustness of Quantum Classifiers under Known-Readout Query AccessJi Guan, Mingyu Huang · arXiv · Sep 10, 2026
A quantum classifier assigns labels by evolving an input quantum state and measuring the output, so repeated executions reveal only a distribution over labels. We study certified adversarial robustness for such classifiers under known-reado…
- On Identifying Sound Conditions for Frontrunning ResistanceSebastian Holler, Anna Piscitelli, Jannik Albrecht, Stephan Dübler et al. · arXiv · Sep 10, 2026
Blockchains enable decentralized applications through smart contracts---interactive programs executed through consensus. However, the inherently asynchronous nature of blockchain transaction ordering introduces a class of vulnerabilities kn…
- Chypothermia: Clock Freezing for Static Side-channel AttacksFatemeh Khojasteh Dana, Mehmet Ali Cetin, Xinrui Wang, Andrew Butler et al. · arXiv · Sep 10, 2026
Static side-channel attacks, which exploit halted-clock conditions to extract sensitive information, pose an increasing threat to chip security. To counter these attacks, various defenses have been proposed that monitor for abnormal clock b…
- Deep-Fake CAPTCHA: Mitigating Next-Generation Social Engineering AttacksGuy Frankovits, Lior Yasur, Fred M. Grabovski, Yisroel Mirsky · arXiv · Sep 10, 2026
This paper presents DF-CAPTCHA, an active defense against real-time deepfake impersonation in voice and video calls. Instead of passively searching for artifacts, DF-CAPTCHA prompts the caller to perform simple challenge-response tasks that…
- Few-Shot Learning for Network Intrusion Detection: Methods, Datasets, and PerformanceArne Roszeitis, Victor Jüttner, Erik Buchmann · arXiv · Sep 10, 2026
Anomaly-based network intrusion detection systems (NIDS) are an important first line of defense. However, training NIDS for new attack types is challenging, because labeled attack data are rarely available. Few-shot learning (FSL) addresses…
- Domain-Incremental Learning for Multi-Channel Replay Speech DetectionMichael Neri · arXiv · Sep 10, 2026
Replay attacks are the most accessible threat to voice-controlled systems, and the acoustic cues that expose them are strongly modulated by the environment in which the attack is mounted. A detector deployed in the field therefore has to ab…
- ToxicRAG: Compromising Retrieval-Augmented Generation Systems via Single-Shot Knowledge Poisoning AttacksHaozhe Lu, Jiaqi Li, Xinyuan Zhu, Xiang Li · arXiv · Sep 10, 2026
Retrieval-Augmented Generation (RAG) can ground large language model (LLM) outputs in external evidence, but it also exposes the system to knowledge poisoning. Representative attacks use multiple injected documents or templates that directl…
- The Missing Boundary: How Autonomous Agents Lose ControlZonghao Ying, Xiangfan Wu, Huiyu Wu, Xing Zheng et al. · arXiv · Sep 10, 2026
Autonomous agents increasingly perform long-horizon tasks involving tool use, persistent state, and consequential actions, raising a fundamental question: \emph{under what conditions does an agent cross the boundary of authorized execution …
- DeFiFusion: Combining Transaction Events with Smart Contracts to Detect Price Manipulation AttacksRui Cao, Shaojing Fan, Liming Fang, Yuchan Liu et al. · arXiv · Sep 10, 2026
Decentralized Finance (DeFi) has emerged as a rapidly growing blockchain-based financial service, where market transaction dynamics and underlying smart contract logic are intricately intertwined. This autonomous interplay, while eliminatin…
- Empirical Evaluation of Data Poisoning Attacks in Supervised LearningToshif Khan, Muhammad Abusaqer · arXiv · Sep 10, 2026
Data poisoning corrupts training data to degrade a model or to plant attacker-controlled behavior. This study evaluates two representative training-time attacks, label flipping and backdoor poisoning, on MNIST and Fashion-MNIST with three b…
- Empirical Evaluation of Membership Inference Attacks on NLP Text Classifiers: A Baseline Study on SST-2William Novak, Muhammad Abusaqer · arXiv · Sep 10, 2026
Membership inference attacks (MIAs) try to determine whether a specific record was used to train a model, a privacy risk that matters in natural language processing (NLP), where training data can contain sensitive user text. This paper pres…
- Robustness Evaluation of Image Perception Using Adversarial Alignment-Based Domain Adaptation and Cross-Adverse-Weather Domain GeneralizationJinha Kim, Seok-Cheol Kee · Transactions of Korean Soci... · Sep 10, 2026
- DriftNet: A Dual-Head Trajectory Transformer for Detecting and Localizing Prompt Injection in LLM AgentsAsif Pinjari, Mithun Paul Saint-Germain · arXiv · Sep 9, 2026
When an indirect prompt injection succeeds against an LLM agent, the compromise is visible in the agent's own behavior: a benign prefix of tool calls, a poisoned observation, and a suffix of actions that serve the attacker. An operator need…
- Lower Bounds for Private Graph Optimization Problems using Reconstruction AttacksJacob Imola, Rasmus Pagh, Lukas Retschmeier · arXiv · Sep 9, 2026
This paper studies fundamental graph optimization problems under differential privacy (DP) and shows new, reconstruction-based lower bounds. We consider a graph $G = (V, E, \vec{w})$ where the vertex set $V$ and edges $E$ are public and the…
- Beyond Static Guarantees: Measuring the Static-Pass Dynamic-Fail Gap in Security-Sensitive and LLM-Generated Python CodeJessica Pourleyli, Maitreyee Das Urmi, Glaucia Melo · arXiv · Sep 9, 2026
Advances in large language models (LLMs) fuel the quest for scalable methods to assess the security of generated and security-sensitive software. Static analysis is widely adopted as a scalable, reproducible, and inexpensive security gate, …
- Temporal and Multimodal Deep Learning for Cyberattack Detection in LEO Satellite SystemsKyle Stein, Guillermo Francia, Eman El-Sheikh, Hossain Shahriar · arXiv · Sep 9, 2026
The growing reliance on Low-Earth Orbit (LEO) satellite communication systems has increased the need for intelligent methods capable of detecting cyberattacks across complex and dynamic space environments. Unlike conventional network intrus…
- Wicked Problem, Parsimonious Solution: Securing Electric Vehicle Charging Station SoftwareEmma Sheppard, Zachary Wadhams, Dalton Arford, Clemente Izurieta et al. · arXiv · Sep 9, 2026
Electric vehicle charging infrastructure presents a suite of novel cyber-physical threats. Among this infrastructure, charging stations are the most vulnerable elements. The software in the charging station supply equipment is particularly …
- Learning Intrusion Response Strategies for OT SystemsDuc Huy Le, Rolf Stadler · arXiv · Sep 9, 2026
Cyberattacks against Operational Technology (OT) systems, which monitor and control industrial processes, pose an increasing threat to essential societal services. For this reason, developing automated intrusion response strategies is highl…
- Are Unreachable Nodes Truly Safe? Fully Eclipsing Monero's P2P Network!Ruisheng Shi, Jiaqi Zeng, Lina Lan, Shihan Zhang et al. · arXiv · Sep 9, 2026
Eclipse attacks isolate a blockchain node by monopolizing its network connections. Existing attacks on Monero (NDSS'25), Bitcoin (USENIX'15/21, S&P'20) and Ethereum (WWW'26) implicitly assume that the adversary can establish inbound con…
- Friend-Safe Adversarial Attack for selective evasion in personalized federated learningHyun Kwon, Dae-Jin Kim · Discover Computing · Sep 9, 2026
Federated Learning (FL) enables collaborative model training across distributed clients while preserving data privacy; however, the personalization of client models in non-independent and identically distributed (non-IID) settings creates a…
- Two-stage adversarial collaborative diagnosis framework via transferable perturbation mechanism for HST bogies with long-tailed dataYuanhong Chang, Yujian Xie, Jiayi Li, Wenping Zhang et al. · Measurement Science and Tec... · Sep 9, 2026
Abstract Deep learning models for high-speed train (HST) bogie transmission components fault diagnosis frequently suffer from severe performance degradation under unseen operating conditions due to domain shift and long-tailed data distribu…
- Simultaneous Faults and Cyber-Attacks Diagnosis in Wind Turbines; an LMI Approach Using Memory-Based Dynamic Residual FieldMehdi Shakeri, Mehrdad Babazadeh, Mahdi Khodabandeh · Iranian Journal of Science ... · Sep 8, 2026
- Security Evaluation of Classical Machine Learning Models Under Poisoning, Evasion, and Model Extraction AttacksHannelore Sebestyen, Elisa Valentina Moisi, Simina Coman, Daniela Elena Popescu · Applied Sciences · Sep 8, 2026
Classical machine learning (ML) models, including Logistic Regression (LR), Support Vector Machines (SVM), Random Forests (RF), and XGBoost, remain widely used in practical applications because of their efficiency, interpretability, and rel…
- Architectural robustness in lane detection systems: Geometric constraints and adversarial stabilityEric Yocam, Varghese Vaidyan, Mark Ngotonie, Denis Ruganuza et al. · Transportation Research Int... · Sep 8, 2026
- Guided Adversarial Robust Transfer Learning with Source MixingXin Xiong, Zijian Guo, Tianxi Cai · Journal of the American Sta... · Sep 8, 2026
- Offensive and Defensive Dynamics of Adversarial AI in CybersecurityRosemary Chisom Dimakunne, Paul Clement Uwamotobon Akpabio, Monsuru Olarewaju Moshood · Zenodo (CERN European Organ... · Sep 8, 2026
This study provides a dual perspective analysis of adversarial AI in cybersecurity, examining both attack (offensive) and defense (protective) strategies. The motivation stems from the growing adoption of machine learning (ML) based intrusi…