Latest Agentic AI & LLM Agents Research Papers
The newest Agentic AI & LLM Agents papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Agentic AI & LLM Agents so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Agentic AI & LLM Agents papers in your inbox — free →Recent papers
- On Generating Text: Editorial Guidance on Generative and Agentic AI for Authors and ReviewersHameed Chughtai · Lancaster EPrints (Lancaste... · Sep 30, 2026
Let me be clear: generating text is not the same as writing research.It is beside the point whether a text is generated by or with an AI-powered tool.Why? Research is conducted by researchers and written by the authors who are those researc…
- Consolidation Without Weights: What the Complementary Learning Systems Analogy Licenses in LLM Agent Memory, and Why the Systems That Borrow Its Name Do Not Inherit Its GuaranteePranay M. Mahendrakar · Zenodo (CERN European Organ... · Sep 11, 2026
(c) 2026 Pranay Mahendrakar. Licensed under CC BY 4.0. Memory systems for language-model agents almost all contain a step called consolidation, and almost all of them cite, or gesture at, the complementary learning systems account of hippoc…
- Consolidation Without Weights: What the Complementary Learning Systems Analogy Licenses in LLM Agent Memory, and Why the Systems That Borrow Its Name Do Not Inherit Its GuaranteePranay M. Mahendrakar · Zenodo (CERN European Organ... · Sep 11, 2026
(c) 2026 Pranay Mahendrakar. Licensed under CC BY 4.0. Memory systems for language-model agents almost all contain a step called consolidation, and almost all of them cite, or gesture at, the complementary learning systems account of hippoc…
- Artificial Id: Drive and Persistent Alignment in Agentic AIYakov Pyotr Shkolnikov · arXiv · Sep 10, 2026
Agentic AI is moving from bounded task execution toward systems that retain consequential state, continue operating and adapt across task boundaries. That shift creates a control problem that current harnesses largely solve by hand: objecti…
- Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM AgentsRuiqing Yue, Yu Cui, Zhuoyu Sun, Sicheng Pan et al. · arXiv · Sep 10, 2026
Self-evolving runtime harnesses can substantially improve the capabilities of large language model (LLM) agents and provide a promising paradigm for optimizing agent execution. Existing harness evolution methods typically rely on iterative …
- Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation FailsZhou Yu, Bin Bi, Shiva Kumar Pentyala, Shubham Mehrotra et al. · arXiv · Sep 8, 2026
Agent harnesses (the system prompt, tool set, execution hooks, and context-management scaffolding around a model) are a critical determinant of agentic task success. Automated harness evolution can enable smaller models to perform well on d…
- MeClear: Cooperative Game-Theoretic Attribution and Risk-Aware Memory Clearance for Long-Horizon LLM AgentsBoyu Yang, Jiazheng Sun, Zilong Lu, Zhi Qiu et al. · arXiv · Sep 8, 2026
Long horizon Large Language Model (LLM) agents rely on external memory systems to preserve user preferences and task knowledge across extended interactions. Conventional retrieval mechanisms optimize semantic compatibility rather than downs…
- Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis RecipeDain Kim, Eungi Cho, Kyumin Kim, Shinyeong Noh et al. · arXiv · Sep 4, 2026
Data-sovereignty regulations increasingly require public institutions to deploy open-source, on-premise LLM agents that chain multiple tool-calls across live government APIs. However, open-source models consistently underperform in this mul…
- A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research SwarmsDavide Paglieri, Logan Cross, Tim Genewein, Joel Z. Leibo et al. · arXiv · Sep 3, 2026
Multi-agent AI science ecosystems rely on agents possessing tools that allow them to communicate, coordinate, and build on each other's work. Yet this shared infrastructure can also introduce vulnerabilities by creating a substrate for the …
- SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations CenterUday Vallabhaneni, Cassie L. Cagwin, David J. Wild · arXiv · Sep 3, 2026
Large language model (LLM) agents are increasingly proposed as autonomous SOC analysts, but two limitations make them unreliable at enterprise scale: a finite context window cannot hold a multi-thousand-host authentication graph, and free-f…
- Evolving transition-state search with agentic large language modelsJan A Meissner, Philipp Kuboth, Jan Meisner · npj Computational Materials · Sep 3, 2026
Abstract Accurate identification of transition states (TSs) is fundamental to computational chemistry. Modern reaction-discovery efforts increasingly rely on curating and completing large reaction datasets, where even a small fraction of TS…
- Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM SystemsYihang Chen, Yuxiang Chen, Yuxuan Huang, Meng Fang et al. · arXiv · Sep 2, 2026
Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through textual reflection. Despite strong empirical results, these systems lack a unified account of coordination, memory impro…
- Beyond static responses: multi-agent LLM systems as a new paradigm for social science researchJennifer Haase, Sebastian Pokutta · Humanities and Social Scien... · Sep 2, 2026
Abstract As large language models (LLMs) transition from static tools to fully agentic systems, their potential for transforming social science research is well recognized. This paper introduces a structured framework for understanding the …
- EvoSCM: Scientific Belief Revision Through Causal Model Evolution and ExperimentationQing Zhao, Haowei Li, Weijian Deng, Pengxu Wei et al. · arXiv · Sep 1, 2026
Scientific agents must learn not only how to reason, but also what to believe. However, existing LLM agents typically express scientific hypotheses in free-form text, leaving their beliefs implicit and difficult to test or revise. We introd…
- When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce EvaluationPeiying Zhu, Sidi Chang · arXiv · Sep 1, 2026
Interactive simulations increasingly evaluate policies in markets populated by language-model agents. Their outputs can look economic---prices, profits, consumer surplus, and welfare---without instantiating the behavior named in the claim. …
- Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured DataMilad Rezaei Hajidehi, Qitong Wang, Stratos Idreos · arXiv · Aug 31, 2026
Valuable data remains embedded in unstructured sources: web pages, reports, contracts, filings, earnings calls, and PDFs. The big bet in enterprise AI is deploying LLM agents that reason over this data to answer complex questions for every …
- Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy OptimizationJingxiao Yang, Wangjie Gan, Yingxuan Zhuang, Wenqi Zhang et al. · arXiv · Aug 31, 2026
Outcome-based reinforcement learning provides verified feedback for language-model agents, but assigns trajectory-level advantage uniformly to all decisions, yielding coarse credit over long-horizon interactions. On-policy self-distillation…
- Scaling Large Reasoning Models beyond Human Supervision: A Path toward SuperintelligenceZhiqin Yang, Jingwen Fu, Yuhan Liu, Hengyu Liu et al. · arXiv · Aug 31, 2026
Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this …
- MNIST-PRO: MNIST is Back as a Partially Observable World for AI AgentsVernon Toh, Navonil Majumder, Zhengyuan Liu, Nancy F. Chen et al. · arXiv · Aug 31, 2026
AI agents in partially observable environments need to coordinate active sensing with working memory to maintain an evolving perceptual state. However, existing benchmarks struggle to isolate this perceptual-state construction and interpret…
- ASSAY: No Action Without a PredictionArjmandi Mohsen · Zenodo (CERN European Organ... · Aug 30, 2026
ASSAY is a general agent harness built so that an LLM agent reasons its way through an unfamiliar world, learns that world from interaction at test time, and carries what it learns into later runs. A world is attached through one small adap…
- On the Maintenance and Co-evolution of Agent Plugins: An Empirical Study of Claude Code Plugin MarketplacesAhmed Hereiz, Yingzhe Lyu, Hao Li, Bram Adams et al. · arXiv · Aug 28, 2026
AI coding agents, software tools that automate development tasks through reasoning and tool use, are increasingly extended through plugin marketplaces, yet the structure, maintenance, and co-evolution dynamics of these emerging repositories…
- Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical ReasoningMinghui Xu, Zi Wang · arXiv · Aug 28, 2026
Current large language models (LLMs) increasingly benefit from external tool integration, especially for tasks requiring reliable computation and verification. Motivated by this, we study calculator tool calling for improving mathematical r…
- RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill EvolutionJunjie Zhang, Hui Liu, Kecheng Chen, Xianbo Mo et al. · arXiv · Aug 27, 2026
LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes, creating greater risks than unsafe text generation alone. Existing automatic red-te…
- Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution AuditYisen Xi · arXiv · Aug 27, 2026
Large language model (LLM) agents in governed organizations must let the persona (instructions, tone, self-presentation) evolve freely, while keeping execution (stateful, audited work) traceable. A single trust domain does not satisfy both …
- An Epistemic-Content Taxonomy of Human Intervention in Agentic CollaborationJongsun Suh · Zenodo (CERN European Organ... · Aug 27, 2026
When a human intervenes in an agent's work, their input supplies epistemic content; yet existing taxonomies index such moves only by timing, authority, granularity, defect class, or interaction state, not by what the input supplies. Two mon…
- AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMsSheng Liang, Yongyue Zhang, Nathanael Brian, Hang Lv et al. · arXiv · Aug 26, 2026
Agentic LLM pipelines face escalating inference costs as context accumulates across retrieval, tool use, and multi-turn interactions. To control latency, deployments routinely compress inputs, but this degrades task accuracy. Speculative de…
- AUTOMATION OF SOFTWARE DEVELOPMENT LIFECYCLE MANAGEMENT PROCESSES USING LLM AGENTS: INDUSTRIAL DEPLOYMENT EXPERIENCEБ. Поляков, Pashko A. · Zenodo (CERN European Organ... · Aug 26, 2026
The paper presents an industrial deployment experience of an LLM-agent-based automation loop for software development lifecycle processes — agents integrated into individual stages of the traditional lifecycle. The loop comprises planning w…
- The Impact of Claude Agents and Autonomous AI Agents on Higher Education: Opportunities, Challenges, and Implications for Educational LeadershipBogdan Costache · International Journal of Ed... · Aug 26, 2026
Artificial intelligence (AI) is rapidly evolving from conversational large language models (LLMs) toward Agentic AI, represented by autonomous AI agents capable of reasoning, planning, memory management, tool use, and autonomous workflow ex…
- SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RLKai Ruan, Jinghao Lin, Qianshan Wei, Ziqi Zhou et al. · arXiv · Aug 25, 2026
Group-relative reinforcement learning waits for sibling rollouts of the same prompt, which is costly for long and variable tool-use trajectories. Single-stream Policy Optimization (SPO) removes this dependency with a persistent prompt-level…
- Meta$^n$: Recursive Self-Improvement through Emergent DepthZae Myung Kim, Young-Jun Lee, Seungyeon Jwa, Dongyeop Kang · arXiv · Aug 25, 2026
Self-improving LLM agents refine answers, not the process that produces those answers. Systems that add a meta-level hold that level fixed, and those that edit themselves must leave part of their own editing machinery untouched to stay stab…