Latest Reinforcement Learning Research Papers
The newest Reinforcement Learning papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Reinforcement Learning so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Reinforcement Learning papers in your inbox — free →Recent papers
- Digital-Twin-Driven Predictive Maintenance and Fault-Tolerant Control for Electrified Agricultural Machinery: A Multiphysics and Deep Reinforcement Learning Framework for PMSM In-Wheel DrivesHongyu Xia · OpenAlex · Dec 31, 2026
Electrified agricultural machinery increasingly relies on permanent-magnet synchronous motor drives that must remain dependable under variable traction, terrain-induced vibration, thermal cycling, contamination, and intermittent duty. Exist…
- Digital-Twin-Driven Predictive Maintenance and Fault-Tolerant Control for Electrified Agricultural Machinery: A Multiphysics and Deep Reinforcement Learning Framework for PMSM In-Wheel DrivesHongyu Xia · Knowledge Commons (Lakehead... · Dec 31, 2026
Electrified agricultural machinery increasingly relies on permanent-magnet synchronous motor drives that must remain dependable under variable traction, terrain-induced vibration, thermal cycling, contamination, and intermittent duty. Exist…
- Near-Optimal Reinforcement Learning with Multi-Step Transition LookaheadCorentin Pla, Hugo Richard, Marc Abeille, Vianney Perchet · arXiv · Sep 10, 2026
We study reinforcement learning (RL) with transition look-ahead, where the agent may observe which states would be visited upon playing any sequence of $\ell$ actions before deciding its course of action. Although look-ahead can substantial…
- Reflex-Informed Neuromuscular Reinforcement Learning for Muscle-Driven LocomotionJian Zhou, Xingyu Zhang, Rui Ma, Yu Cao et al. · arXiv · Sep 10, 2026
Muscle-driven locomotion provides a physically grounded approach to generating realistic human movement. However, achieving both physiological plausibility and adaptability to changes in musculoskeletal capacity and external disturbances re…
- Surrogate-Assisted Deep Reinforcement Learning for Bi-Level Multi-Objective Capacity Sizing of Railway Collaborative Power Supply SystemYaozhen Chen, Mingli Wu, Jingtao Lu, Zheng Liu et al. · Electric Power Systems Rese... · Sep 10, 2026
- Effects of declarative language and an indirect reward system on participation in adult-selected activities in an autistic child with features of pathological demand avoidanceElizabeth A. Donovan, Ya‐yu Lo, Rebecca A. Payton · Research in Autism · Sep 10, 2026
Pathological Demand Avoidance (PDA) is a behavioral profile that is often associated with autism spectrum disorder and has received increased attention in the field of child psychiatry, child neurology, and pediatrics. Despite its prevalenc…
- Multi-Agent Reinforcement Learning for Autonomous UAV Exploration in Wildfire ResponseCaden Chandra, Jerry Ng · arXiv · Sep 9, 2026
This study develops a deep reinforcement learning framework for training Unmanned Aerial Vehicle (UAV) agents to navigate and monitor simulated wildfire environments. Results show that agents learn increasingly stable and effective behavior…
- Searching for New Physics with Reinforcement LearningJacky Kumar, Marianne Bouchard, David London · arXiv · Sep 9, 2026
Finding new physics (NP) is the most important problem in particle physics today. Studying ``anomalies'', i.e., measurements of low-energy observables whose values disagree with the predictions of the Standard Model (SM), is a powerful sear…
- TRACE: Training Reasoning Agents for Causal Exploration with Synthesized RewardsRui Sun, Zhan Shi, Bing He · arXiv · Sep 9, 2026
Reinforcement learning with verifiable rewards (RLVR) has advanced language-model reasoning in domains such as mathematics and code, where objective answers are inexpensive to check. Diagnostic reasoning over complex data lacks this advanta…
- Deep reinforcement learning for dynamic origin-destination matrix estimation in microscopic traffic simulations considering credit assignmentDonggyu Min, Seongjin Choi, Dong‐Kyu Kim · Transportation Research Par... · Sep 9, 2026
- PPORLD-EDNetLDCT: A Proximal Policy Optimization-based reinforcement learning framework for adaptive low-dose CT denoisingDebopom Sutradhar, Ripon Kumar Debnath, Mohaimenul Azam Khan Raiaan, Yan Zhang et al. · Biomedical Signal Processin... · Sep 9, 2026
Low-dose computed tomography (LDCT) is critical for minimizing radiation exposure, but it often leads to increased noise and reduced image quality. Traditional denoising methods, such as iterative optimization or supervised learning, often …
- Entropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code GenerationJiacheng Xu, Feng Chen, Xiuneng Xu, Bo An · arXiv · Sep 8, 2026
Existing methods for test-time reinforcement learning (TTRL) derive rewards from answer-level self-voting on unlabeled test-time tasks with canonical answers, but this breaks down for code generation because programs cannot be compared by s…
- ThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVRTommy Sha, Skylar Zhai, Siqi Zhao · arXiv · Sep 8, 2026
In reinforcement learning with verifiable rewards (RLVR) trained with group relative policy optimization (GRPO), the KL-free reward-advantage term studied here depends on within-group reward variation. If all rollouts in a group are correct…
- PlayTrain: An Efficient Reinforcement Learning Framework for LLM-Generated Adaptable JavaScript GamesRyan Truong, Lance Ying, Samuel J. Gershman, Kazuki Irie · arXiv · Sep 8, 2026
While many video-game environments (VGEs) have played crucial roles in advancing reinforcement learning (RL), developing novel VGEs or modifying existing ones to support new features, has been a laborious process requiring extensive hand-co…
- Reinforcement Learning-Assisted Quantum Simulation of Many-Body Excited States and Real-Time DynamicsJiaji Zhang, Lipeng Chen, Carlos L. Benavides-Riveros · Journal of Chemical Theory ... · Sep 8, 2026
Abstract The computation of electronic excited states and real-time quantum dynamics of many-Fermion systems is among the most promising applications of near-term quantum computing. In this work, we generalize the reinforcement learning con…
- Artificial Intelligence in Quantitative Trading: Application Pipeline, Prospects, and Risk GovernanceZichong Long · Applied and Computational E... · Sep 8, 2026
Artificial intelligence (AI) has become a significant force in the field of quantitative trading because it extends traditional rule-based systems to areas such as adaptive prediction, dynamic configuration and automated execution. At the s…
- Reinforcement Learning Methods and Optimization Strategies in Preference Alignment Techniques for Large Language ModelsHongyu Jiang · Applied and Computational E... · Sep 8, 2026
The growing pre-training scale has improved large language models' performance in language generation and knowledge representation. However, their training objectives remain limited to fitting data distributions, thus making it difficult to…
- A Survey on Reinforcement Learning Optimization Methods for Multi-Agent Collaboration of Large Language ModelsQianling Zhang · Applied and Computational E... · Sep 8, 2026
The integration of reinforcement learning (RL) into the optimization of multi-agent collaboration for Large Language Models (LLMs) is an important combination of two advanced areas, Multi-Agent Systems (MAS) and LLMs. This paper thoroughly …
- Controllable molecular generation with fine-tuned flow-matching modelK.-H. Wang, Jon Paul Janet, Alessandro Tibo · Communications Chemistry · Sep 5, 2026
Abstract Three-dimensional molecular generative models have emerged that produce de novo molecules both unconditionally and conditionally, e.g., within protein pockets. However, steering those models in a specific region of the chemical spa…
- Trans-SAC: gamification design for sustained user engagement in residential demand responseTianxiao Peng · Journal of Engineering and ... · Sep 5, 2026
Abstract Residential Demand Response (DR) is a critical mechanism for maintaining the supply-demand balance in modern smart grids. While gamification has recently emerged as a promising strategy to incentivize residential participation, exi…
- Decision-Relative Observation Quotients for Sequential ControlKaliel Williamson · Zenodo (CERN European Organ... · Sep 5, 2026
This paper studies when a reinforcement-learning system should acquire additional information before acting. It introduces causal observability: a finite criterion for whether an observation regime preserves the intervention-relevant distin…
- Regime-Aware Reinforcement Learning: A Mixture-of-Experts Framework for Dynamic Asset AllocationYirui Luo, John M. Mulvey · The Journal of Financial Da... · Sep 5, 2026
Recent advances in reinforcement learning (RL) have spurred growing interest in its application to multi-period financial planning. Existing literature broadly follows three paradigms: hybrid RL, scalable RL, and end-to-end RL. This paper d…
- EOL-DQN: An intelligent cutting decision of potato seed tubers based on improved deep reinforcement learningXuyang Wang, Yun Liu, Tianshuo Su, Yulu Sun et al. · Computers and Electronics i... · Sep 5, 2026
- Online Change-point Detection for Cooperative Multi-Agent Reinforcement LearningFatemeh Saberi Khomami, Julita Vassileva · arXiv · Sep 4, 2026
Cooperative multi-agent reinforcement learning (MARL) systems rely on past experience for learning coordinated behaviour, but this experience may become unreliable if the environment or task objective changes during training. In such cases,…
- Adaptive rocket trajectory optimization via reinforcement learning and FPGA accelerationWen-Gang Yao, Yang-Ming Guo, Limin Mao, Lei-Lei Zhang et al. · Scientific Reports · Sep 4, 2026
- Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought ReasoningKevin Du, Alexander Hoyle, Laura Ruis, Acyr Locatelli · arXiv · Sep 3, 2026
Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its answer. A growing body of work treats them as such, using LLM judges to diagnose errors, evaluate faithfulness, and provide step-…
- Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVRBoyan Li, Bingsen Chen, Chenghao Yang, Ping Nie et al. · arXiv · Sep 3, 2026
Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have emerged as two dominant methods for post-training reasoning LLMs. Prior work uses OPD's dense token-level supervision to complement the sparse RL re…
- DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent TrainingShubham Gandhi, Saurabh Goyal, Kiran Kate, Yara Rizk · arXiv · Sep 3, 2026
Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most long-horizon agent domains have none. We work in the outcome-blind setting, where ground-truth success signals are not available. Mul…
- Subspace Inference Enables Efficient Active Reward Learning from PreferencesYutai Zhou, Erdem Bıyık · arXiv · Sep 3, 2026
Reinforcement learning from human feedback (RLHF) has emerged as a powerful yet sample-inefficient approach for learning reward models from human preferences, making active learning a critical component in synthesizing informative preferenc…
- MOAC-RT: Personalized multi-objective radiotherapy using actor-critic deep reinforcement learningAva Khosravi, Mehdy Roayaei · Soft Computing · Sep 3, 2026