Latest AI Safety & Alignment Research Papers
The newest AI Safety & Alignment papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks AI Safety & Alignment so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest AI Safety & Alignment papers in your inbox — free →Recent papers
- Multi-dimensional hierarchical temporal alignment for improved temporal commonsense reasoning in large language modelsGe Yan, Hai-Tao Yu, Lei Chao · Institutional Repositories ... · Nov 1, 2026
Benefiting from recent advances in generative AI and large language models (LLMs), current LLM-driven AI systems are rapidly reshaping how we reason, plan, and make decisions. Yet endowing these systems with human-like temporal intelligence…
- The development of Japanese medical ethics education from 1950s and its lessons for ChinaKai Hou, Peisen Li · Scientific Electronic Libra... · Oct 1, 2026
Abstract: Background: Modern medical technology has significantly advanced human health and well-being, yet it has also introduced complex ethical challenges that require careful navigation. In this context, medical ethics education plays a…
- Artificial Id: Drive and Persistent Alignment in Agentic AIYakov Pyotr Shkolnikov · arXiv · Sep 10, 2026
Agentic AI is moving from bounded task execution toward systems that retain consequential state, continue operating and adapt across task boundaries. That shift creates a control problem that current harnesses largely solve by hand: objecti…
- Understanding Operator Attitudes Toward AI-Supported Decision Making in Maritime OperationsDoreen Jirak, Armeen Saroukanoff, Dirk van Rooy · arXiv · Sep 10, 2026
Maritime Autonomous Surface Ships (MASS) and AI- supported decision assistants are expected to transform maritime operations, but their safe integration depends on how maritime professionals perceive and trust such systems. This paper prese…
- LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language GenerationDongfang Zhao · arXiv · Sep 10, 2026
Large language model serving costs scale directly with output sequence length, yet standard preference alignment often inflates response verbosity without improving utility. We study whether the parameterization of post-training updates aff…
- Continuous-Time Acoustic Modelling with Neural Controlled Differential EquationsMattias Cross, Minghui Zhao, Anton Ragni · arXiv · Sep 10, 2026
Text-to-speech (TTS) models commonly address text--speech alignment by expanding phone-level encoder states to frame-level decoder inputs using predicted durations. While this length-regulation step resolves alignment structurally, this use…
- ActSafeGuard: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching PoliciesJianming Ma, Rongjun Jin, Xiaxi Si, Yang Zhang et al. · arXiv · Sep 10, 2026
Vision-Language-Action (VLA) and World-Action Models (WAMs) have demonstrated strong capabilities in general-purpose robotic manipulation, yet their generated actions may violate hard physical constraints and therefore be unsafe or infeasib…
- Autonomy, Social Norms, and Alignment: Towards a Developmental Framework for Autonomous Artificial AgentsMarica Notte, Ludovica Marinucci, Vieri Giuliano Santucci · arXiv · Sep 10, 2026
In recent years, artificial intelligence has made extraordinary progress thanks to large-scale models capable of generalization and the generation of complex outputs. However, transferring this potential into embodied agents reveals a signi…
- X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale DistillationHaojun Zhang, Yi Zou, Min Chen, Qize Yu et al. · arXiv · Sep 10, 2026
Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors. We introduce X…
- Off-Target Effects of Response-Style Alignment in a Korean 27B Language ModelHyojung Han · arXiv · Sep 10, 2026
We post-train Qwen3.8-27B for Korean response style -- verbosity, list and markdown usage, discourse structure and register -- and measure two behaviours the objective never targets: abstention on ambiguous social questions in KoBBQ, where …
- Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMsRavi Ranjan, Olivera Kotevska, Agoritsa Polyzou · arXiv · Sep 9, 2026
Large Language Models (LLMs) can memorize and reproduce sensitive, copyrighted, or otherwise undesirable training content, creating privacy, safety, and regulatory concerns. Machine unlearning offers a practical alternative to full retraini…
- TMA-Grid: an open-source, zero-footprint web application for FAIR tissue microarray de-arrayingAaron Ge, Monjoy Saha, Máire A. Duggan, Petra H. Lenz et al. · BMC Bioinformatics · Sep 9, 2026
Tissue microarrays (TMAs) significantly increase analytical efficiency in histopathology and large-scale epidemiologic studies by allowing multiple tissue cores to be scanned on a single slide. The individual cores can be digitally extracte…
- SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?Yuqiao Tan, Shizhu He, Jun Zhao, Kang Liu · arXiv · Sep 8, 2026
While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autonomous development demands a missing pillar: post-hoc monitoring and auditing to understand what models learn and ensure s…
- Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-TrainingYunpeng Xu, Kun Zheng · arXiv · Sep 8, 2026
Mid-training, the stage between pre-training and alignment, is where a model's per-domain data composition is typically set by data availability rather than principled design. We ask what that decision buys, and whether a later alignment pa…
- The History Is the Detector: Executing CVE Patch History, End-to-EndQiushi Wu, Kevin Eykholt, Youngja Park, Xiaokui Shu et al. · arXiv · Sep 4, 2026
Public vulnerability databases collect rich information about known software flaws, including their weakness types, affected components, and related patches. Fixing commits provide the exact code changes that removed these flaws. While thes…
- Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video CaptioningYe-Chan Kim, Seunghee Choi, SeungJu Cha, Si-Woo Kim et al. · arXiv · Sep 3, 2026
Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos given only an ordered set of event-level captions per video. Recent work synthesizes auxiliary transition captions via LLM to provide…
- SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety AlignmentQinghua Mao, Wanying Qu, Dadi Guo, Leitao Yuan et al. · arXiv · Sep 2, 2026
The performance of LLM-based agents is jointly shaped by the base model and the harness used when interacting with the environment. This exposes them to safety risks in both harmful final responses and multi-step execution trajectories. Exi…
- Mechanism Design for Alignment and ControlDirk Bergemann, Andrew Koh, Stephen Morris · arXiv · Sep 1, 2026
We develop a framework for mechanism design with AI agents whose alignment (preferences) and capabilities (feasible actions and information) are unknown. We want such agents to act on our behalf so mechanisms must incentivize both honesty a…
- OntoAligner-Ensemble: Voting-Based Fusion across Heterogeneous Ontology Alignment TechniquesHamed Babaei Giglou, Sören Auer, Peio Popov, Mahsa Sanaei et al. · arXiv · Aug 31, 2026
Ontology alignment (OA) has evolved through several methodological paradigms, ranging from lexical and structural aligners to knowledge graph embedding (KGE) models and, more recently, Large Language Model (LLM)-based approaches. Although m…
- When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AISihan Jia, Oliver Lemon · arXiv · Aug 28, 2026
We investigate whether automatic speech recognition (ASR) errors in user input can lead to unsafe outputs from Embodied AI (EAI) models. We find that ASR errors can lead to harmful instructions being accepted and executed by EAI models, the…
- RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill EvolutionJunjie Zhang, Hui Liu, Kecheng Chen, Xianbo Mo et al. · arXiv · Aug 27, 2026
LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes, creating greater risks than unsafe text generation alone. Existing automatic red-te…
- Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security ScannersQianlong Lan, Vinothini Pandurangan, Anuj Kaul, Indranil Sanyal · arXiv · Aug 27, 2026
Static scanners are increasingly used to identify executable or otherwise unsafe content in machine- learning artifacts, yet conventional evaluation metrics characterize only cases where a scanner yields a usable security judgment. We evalu…
- PAWBench: How Far Are We from Probabilistically Aligned World Modeling?Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes et al. · arXiv · Aug 27, 2026
Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also the distribution of p…
- Not All Eval-Awareness Is Equal: Capabilities Framing Predicts ComplianceAllison Zhuang, Santiago Aranguri · arXiv · Aug 27, 2026
Steering interventions targeting eval-awareness, a model's recognition that it is being tested, are increasingly used in safety evaluation pipelines, where evaluation-awareness is treated as a single quantity to be suppressed. We show that …
- StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility BalancingZhijie Zheng, Yu Li, Chen Qian, Yuqian Fu et al. · arXiv · Aug 25, 2026
LLM-based agents can interact with external environments through tool invocation, but this capability also introduces security risks such as file modification, information leakage, and unauthorized actions. Existing guardrails often evaluat…
- RACE: Scalable Statistical Estimation of Functional Consistency in LLM NeuronsRunyu Wang, Bo Liu, Xiaxin Zhang, Yu Han et al. · arXiv · Aug 25, 2026
Discovering stable neuron behavior across entire domains remains a challenge in mechanistic interpretability. Existing methods often rely on instance-level point estimates or computationally expensive procedures, which either obscure popula…
- AI with Authority, from Application to SiliconJason Hickey · arXiv · Aug 21, 2026
For sixty years, machine verification has been a major cost overhead, affordable only for exceptional artifacts. Here we report that generative AI inverts this relationship: at AI speed, machine verification is not only economical but essen…
- CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety AlignmentChengxiao Wang, Enyi Jiang, Xiaojing Liao, Sanmi Koyejo · arXiv · Aug 21, 2026
Improving the safety of large language models (LLMs) often comes at the expense of utility, as globally applied safety tuning may affect model responses to both harmful and benign inputs. We propose \textbf{C}ontinuous \textbf{L}at\textbf{E…
- Electronic Navigational Chart Change ClassificationJacob Arndt, Abhishek Potnis, Alexandre Sorokine · arXiv · Aug 20, 2026
Electronic Navigational Charts (ENCs) are geospatial vector datasets used in maritime navigation systems that represent hydrographic and navigational information such as depths, navigational aids, traffic schemes, and hazards. A major chall…
- Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent CommunicationRamneet Kaur, Pradyumna Chari, Ramesh Raskar, Jugad Singh et al. · arXiv · Aug 19, 2026
Language-model agents can communicate through continuous hidden states that are invisible in public transcripts, creating opportunities for covert harmful coordination. We introduce Verifiable Latent Alignments (VLA), an activation-aware fr…