Latest Robot Manipulation & VLA Research Papers
The newest Robot Manipulation & VLA papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Robot Manipulation & VLA so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Robot Manipulation & VLA papers in your inbox — free →Recent papers
- UniMPA: A Unified Memory-Prediction-Action Model via Action-Grounded Transition ModelingWei Li, Rui Shao, Jie He, Lingsen Zhang et al. · arXiv · Sep 10, 2026
Recent advances in Vision-Language-Action (VLA) models have improved robotic manipulation, yet observation-to-action learning remains limited by a fundamental transition realizability gap, manifested in three tightly coupled problems: (i) T…
- SEED-UMI: Sharing the Exoskeleton between human and robot for onE-to-one Dexterous demonstrationTengbo Yu, Jiahao Wu, Daohan Li, Bingxu Chen et al. · arXiv · Sep 10, 2026
Imitation learning for dexterous hands is bottlenecked by the difficulty of collecting contact-rich demonstrations that transfer faithfully to the robot. Prior wearable-exoskeleton systems record only on the human side and retarget via open…
- ActSafeGuard: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching PoliciesJianming Ma, Rongjun Jin, Xiaxi Si, Yang Zhang et al. · arXiv · Sep 10, 2026
Vision-Language-Action (VLA) and World-Action Models (WAMs) have demonstrated strong capabilities in general-purpose robotic manipulation, yet their generated actions may violate hard physical constraints and therefore be unsafe or infeasib…
- Quasi-static analysis of passive stability in a novel underactuated multi-finger handLéonie Plancoulaine, Sylvain Guégan, Franck Plestan, Damien Chablat · arXiv · Sep 10, 2026
Underactuated robotic hands achieve adaptive and robust grasping with a reduced number of actuators, but predicting the stable equilibrium pose of the grasped object remains a significant challenge. This paper introduces a quasi-static anal…
- 2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon ManipulationYutong Hu, Fengjiao Chen, Xuezhi Cao, Renaud Detry · arXiv · Sep 10, 2026
Long-horizon robot manipulation requires memory, but not necessarily inside the action policy. To address such tasks, current agentic systems often combine VLAs with planners and geometric tools, sometimes using additional depth or calibrat…
- Beyond Noise Steering: Dual-Latent Space Reinforcement Learning for Generative Robot PolicyPengfei Zhang, Teng Sun, Xianchao Xiu · arXiv · Sep 10, 2026
Pretrained generative robot policies learn expressive action priors from demonstrations. However, existing reinforcement learning methods only steer the noisy space but fail to modulate intermediate action representations during the generat…
- ObstaDiff: Generalizable Diffusion Policy Learning via Obstacle-aware RepresentationsJiawen Wang, Kevin Yao, Khalid Jawed · arXiv · Sep 10, 2026
Imitation learning has achieved impressive results in robotic manipulation, yet most existing approaches assume clean backgrounds and lack explicit mechanisms for obstacle-aware motion generation. Extending such policies to cluttered, real-…
- IMLE-VLA: Fast Single-Step Action Generation for Vision-Language-Action PoliciesKian Hosseinkhani, Qinhe Peng, George Shramko, Mehran Aghabozorgi et al. · arXiv · Sep 10, 2026
Vision-language-action (VLA) policies leverage pretrained vision-language backbones to achieve strong cross-task generalization. A leading design couples this backbone with a dedicated continuous action head trained via diffusion or flow ma…
- DUET-DINO: Simultaneous Cross-View World Modeling for Latent Planning in Robot ManipulationNisarga Nilavadi, Ralf Römer, Moritz Reuss, Michael Krawez et al. · arXiv · Sep 9, 2026
Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However, their predictions for fine-grained spatial and rotational actions are unreliable for full…
- Frequency-Conditioned Flow Matching for Vision-Language-Action ModelsHaochen Niu, Shengye Dong, Hao Liu, Peiwen Lin et al. · arXiv · Sep 9, 2026
Robot actions are temporally correlated trajectories whose frequency components encode motion at different scales with highly non-uniform energy distributions. Yet Flow Matching--based vision-language-action (VLA) models typically generate …
- FolDeX: A Physical-World Benchmark for Long-Horizon Robotic Manipulation of Deformable ObjectsChenhuan Liu, Yi Xu, Feng Wu, Hanyang Wang et al. · arXiv · Sep 9, 2026
Embodied AI, including vision-language-action and world-action models, must operate reliably in the physical world. Yet methods that perform well in simulation can degrade substantially on real robots, especially in long-horizon deformable-…
- Assembling Two Parts in One HandLiuao Pei, Tianyue Wu, Hui Zhang, Ping Luo et al. · arXiv · Sep 9, 2026
A hallmark of human dexterity is the cooperative use of fingers, where different fingers take on distinct yet coordinated roles to accomplish fine manipu- lation, such as capping a pen with the hand that holds it. We study this finger-level…
- Toward human-compatible compliant manipulation with an anthropomorphic dexterous handQi Luo, Sheng Jin, Yongzhi Luo, Jiayu Cao et al. · Bioinspiration & Biomimetics · Sep 9, 2026
Achieving human-like compliant manipulation remains a fundamental challenge in robot hands due to the difficulty of realizing biomechanical compatibility, compliant interaction, and dexterous operation. To address this, we propose a human-c…
- TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action ModelAnqi Li, Yuxin Chen, Zhaobo Li, Zhuo Cao et al. · arXiv · Sep 8, 2026
We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional methods that model navigation as a 2D path planning problem, humanoid traversal in cluttered environments requires continuous geomet…
- DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-ImaginationYankai Fu, Ning Chen, Junkai Zhao, Heng Zhang et al. · arXiv · Sep 8, 2026
Dexterous manipulation involves contact-rich and fine-grained interactions with the physical world, posing significant challenges for existing vision-language-action (VLA) models due to severe visual occlusions and complex contact dynamics.…
- Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action ManipulationVivek Chavan, Yahuan Shi, Oliver Heimann, Kevin Haninger et al. · arXiv · Sep 4, 2026
Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon procedures requiring persistent task state, dependency-aware reasoning, conditional decisions, and reliable grounding. We investig…
- RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan et al. · arXiv · Sep 4, 2026
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight in…
- Temporal Tactile Encoding and Compliance for Intent-Aware Robot-to-Human Bimanual HandoverPasquale Marra, Stefano Berti, Gabriele Mario Caddeo, Lorenzo Natale · arXiv · Sep 4, 2026
Reliable robot-to-human handover requires the robot to infer when the person is ready to receive the object, and release it safely, comfortably, and at the right time. This is challenging because visual observations alone may not disambigua…
- LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation ModelsLin Liu, Zhicheng Bao, Lu Zhang, Ziying Song et al. · arXiv · Sep 4, 2026
Vision-Language-Action (VLA) or World Action (WAM) models have recently demonstrated remarkable performance in robotic manipulation. On LIBERO, SOTA method have achieved nearly 100\% success rates, seemingly suggesting that the models are r…
- A unified framework for estimating fingertip forces and muscle activations in human grasping with anatomical modeling and soft-finger contactRyuki Nohara, Naomichi Ogihara, Mitsunori Tada · Scientific Reports · Sep 4, 2026
Abstract We developed a method to estimate fingertip forces, torques, and muscle activations during object grasping with a digital hand model. Grasping an object requires several equilibrium conditions: force and moment balance on the objec…
- Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp SynthesisSixu Yan, Shikang Wang, Binhua Huang, Xuanlai Tang et al. · arXiv · Sep 3, 2026
This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generalizable grasp synthesis across different robotic hands. Unlike existing VLG methods that tightly couple foundation models with end-to-…
- Continuous Actions from Discrete Minds: Latent-Aligned Planning for End-to-End Autonomous DrivingRuoyu Yao, Yusen Xie, Qingzhao Liu, Pei Liu et al. · arXiv · Sep 3, 2026
Bridging the gap between the discrete reasoning of Vision-Language Models and the continuous, physics-constrained nature of autonomous driving remains a significant challenge. In this work, we introduce LaPla, a unified Vision-Language-Acti…
- Toward Unified Robot Learning: Bridging Representation, Vision-Language-Action, and World ModelsShaunak A. Mehta, Ananya Hazarika, Haochen Zhang, Fan Yang et al. · arXiv · Sep 3, 2026
For robots to operate reliably in real-world environments, they need to perceive their surroundings, act, and reason about the consequences of those actions. Rapid progress in the domains of representation learning, VLA models, and world mo…
- Revisiting Topological Graphs for Macro Action based Closed-loop Reinforcement Learning of Vision Language Navigation in Continuous EnvironmentShuhao Ye, Sitong Mao, Yuxiang Cui, Yufei Wei et al. · arXiv · Sep 3, 2026
Vision-Language Navigation in Continuous Environments (VLN-CE) requires an agent to follow natural language instructions through unseen environments. Existing imitation learning (IL) pipelines struggle in this closed-loop setting: behavior …
- FWBC-VLA: Force-Aware Whole-Body Compensation for Contact-Rich Loco-ManipulationYutian Zhang, Siyuan Ma, Liwen Yang, Yang Li et al. · arXiv · Sep 3, 2026
Contact-rich loco-manipulation requires a bridge between semantic action generation and physical interaction control. Existing Vision-language-action (VLA) models generate task-level actions from visual and linguistic observations, but cann…
- HINT: Human-Intent Inception for Long-Horizon Robot ManipulationMingyu Mei, Haojie Xu, Shihao Jin, Zibo Dai et al. · arXiv · Sep 2, 2026
Humans can perform complex manipulations given a simple intent through an overall instruction, while continuously adapting to evolving visual observations. However, current vision-language action (VLA) models and other action policies strug…
- Latent Cluster Analysis for Vision-Language-Action ModelsTheodor Wulff, Sergio Lanza, Tamara Bila, Angelo Cangelosi et al. · arXiv · Sep 2, 2026
Vision-Language-Action (VLA) Models are increasingly used in robotics for their ability to ground language and perception into action, yet the internal representations driving their behaviour remain poorly understood. We propose LAVLA, a fr…
- ZETA: A Controlled Study of Zero-Shot Cross-Embodiment VLA Transfer for Tabletop ManipulationMi Yan, Wenhao Zhang, Zhiqi Zhang, Yu Peng et al. · arXiv · Sep 2, 2026
Zero-shot generalization to unseen embodiments is important for generalizable vision-language-action (VLA) models as robot hardware evolves and task-specific data collection remains costly. However, a systematic understanding of this proble…
- Spatially Aware World Action Model via Geometric Latent DiffusionJavier Alejandro Lopetegui Gonzalez, Paul Pacaud, Cordelia Schmid · arXiv · Sep 2, 2026
World Action Models (WAMs) leverage the capabilities of large-scale pretrained video diffusion models to jointly predict future observations and actions, inheriting rich visual and physical priors from internet-scale video. This has made th…
- A Physics-Consistent Benchmark for Contact-Rich Human-Robot Interaction in Assistive CareChengxiao He, Shanghai Yuan, Liuqun Fan, Shenzhen Zhu · arXiv · Sep 2, 2026
Conventional task-level evaluation asks whether a robot policy completes a specified action, but can miss failures that emerge only during physical human contact. This limitation is critical in contact-rich assistive tasks, where meaningful…