Latest Embodied AI Research Papers
The newest Embodied AI papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Embodied AI so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Embodied AI papers in your inbox — free →Recent papers
- The development of Japanese medical ethics education from 1950s and its lessons for ChinaKai Hou, Peisen Li · Scientific Electronic Libra... · Oct 1, 2026
Abstract: Background: Modern medical technology has significantly advanced human health and well-being, yet it has also introduced complex ethical challenges that require careful navigation. In this context, medical ethics education plays a…
- UniMPA: A Unified Memory-Prediction-Action Model via Action-Grounded Transition ModelingWei Li, Rui Shao, Jie He, Lingsen Zhang et al. · arXiv · Sep 10, 2026
Recent advances in Vision-Language-Action (VLA) models have improved robotic manipulation, yet observation-to-action learning remains limited by a fundamental transition realizability gap, manifested in three tightly coupled problems: (i) T…
- Rapid Learning of Dexterous In-Hand Pen Writing through Real-Time Jacobian EstimationKai Stewart, Yasunori Toshimitsu, Robert K. Katzschmann · arXiv · Sep 10, 2026
Dexterous in-hand manipulation of a grasped object with an anthropomorphic hand is an unsolved frontier for robot dexterity. The contact-richness and highly dynamic nature of object-hand interactions tend to require extensive modeling or da…
- ORCH: Organizational Principles Enable Collective Intelligence in Embodied AIZhengran Ji, Jonathan Hyun, Boyuan Chen · arXiv · Sep 10, 2026
Collective intelligence depends not only on the capabilities of individual members, but also on how those members are organized. Yet artificial multi-agent systems are typically assembled using fixed organizational structures, even when the…
- ActSafeGuard: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching PoliciesJianming Ma, Rongjun Jin, Xiaxi Si, Yang Zhang et al. · arXiv · Sep 10, 2026
Vision-Language-Action (VLA) and World-Action Models (WAMs) have demonstrated strong capabilities in general-purpose robotic manipulation, yet their generated actions may violate hard physical constraints and therefore be unsafe or infeasib…
- Contact-Aware Incremental Model Predictive Control for an Underactuated Aerial ManipulatorDarwin Liu, Tamas Keviczky, Sihao Sun · arXiv · Sep 10, 2026
We present a robust contact-aware control framework for aerial writing on an underactuated platform. The framework combines nonlinear model predictive control (NMPC) for accurate end-effector position and normal-force tracking at small refe…
- Memory as Plans: World-Action Modeling with Memory-Grounded PlanningSizhe Zhao, Haozhe Xie, Weiyu Zhao, Chenchu Zhang et al. · arXiv · Sep 10, 2026
Mainstream robotic policies often adopt a Markovian formulation, but many complex real-world manipulation tasks are inherently non-Markovian, requiring long-horizon memory beyond the current observation. Existing memory mechanisms often rel…
- Morphology-Aware Human Motion Retargeting for Wheeled-Humanoid Loco-ManipulationChenbo Xia, Chao Ye · arXiv · Sep 10, 2026
Human-to-humanoid retargeting has largely been studied on legged platforms, while comparatively few wheeled-humanoid systems support coupled locomotion and manipulation from general human motion. Building on GMR's configurable general-motio…
- 2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon ManipulationYutong Hu, Fengjiao Chen, Xuezhi Cao, Renaud Detry · arXiv · Sep 10, 2026
Long-horizon robot manipulation requires memory, but not necessarily inside the action policy. To address such tasks, current agentic systems often combine VLAs with planners and geometric tools, sometimes using additional depth or calibrat…
- Harness Robotic OS: A Unified Embodied-Agent Runtime for Closed-Loop Quadruped InspectionYaoyuan Yan, Zhiyou Heng, Haoxiang Jie, Gang Liu et al. · arXiv · Sep 10, 2026
Autonomous property inspection requires more than robust robot navigation: a deployable system must connect heterogeneous sensing, reusable autonomy capabilities, multimodal scene understanding, human interaction, and enterprise response wi…
- LTLDiff: Finite Linear Temporal Logic-Guided Data Generation and Diffusion Policies for Multi-agent Robotic ManipulationChuhan Meng, Haiyan Yin · arXiv · Sep 10, 2026
Multi-agent robotic manipulation tasks require coordination among agents to satisfy task-level temporal, logical, and safety constraints. Recently, diffusion policies have been used to perform the task. However, they still suffer from desyn…
- ObstaDiff: Generalizable Diffusion Policy Learning via Obstacle-aware RepresentationsJiawen Wang, Kevin Yao, Khalid Jawed · arXiv · Sep 10, 2026
Imitation learning has achieved impressive results in robotic manipulation, yet most existing approaches assume clean backgrounds and lack explicit mechanisms for obstacle-aware motion generation. Extending such policies to cluttered, real-…
- Coupling Intelligence to an ExistenceJulien Frappa · Zenodo (CERN European Organ... · Sep 10, 2026
A thought experiment and experimental protocol asking whether irreversibility, lineage, vulnerability and physical coupling alter strategic behaviour in artificial agents. The proposal synthesises existing work in mortal computation, embodi…
- The Prospect of Embodied Intelligence in DentistrySiwei Wang, Huancai Lin, Liangyue Pang · International Dental Journal · Sep 10, 2026
Conventional artificial intelligence in dentistry is limited by a ‘diagnostic-executive disconnect’, functioning primarily as static ‘bystander intelligence’. This review aims to delineate the technical architecture of Embodied Artificial I…
- Planning along Differentiable Charts of Constraint Manifolds with General-Purpose IK SolversThomas Cohn, Seiji Shaw, Harel Biggie, Travis Manderson et al. · arXiv · Sep 9, 2026
Planning trajectories for robot manipulators under kinematic equality constraints restricts feasible motions to a measure-zero submanifold of the configuration space, requiring special algorithmic treatment. A promising strategy is parametr…
- Show-Harness: Just a VLM Agent Can Play RobotsYanzhe Chen, Zechen Bai, Zhijun Cao, Wenzheng Zeng et al. · arXiv · Sep 9, 2026
Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to "play" robots t…
- DUET-DINO: Simultaneous Cross-View World Modeling for Latent Planning in Robot ManipulationNisarga Nilavadi, Ralf Römer, Moritz Reuss, Michael Krawez et al. · arXiv · Sep 9, 2026
Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However, their predictions for fine-grained spatial and rotational actions are unreliable for full…
- Multi-Agent Reinforcement Learning for Autonomous UAV Exploration in Wildfire ResponseCaden Chandra, Jerry Ng · arXiv · Sep 9, 2026
This study develops a deep reinforcement learning framework for training Unmanned Aerial Vehicle (UAV) agents to navigate and monitor simulated wildfire environments. Results show that agents learn increasingly stable and effective behavior…
- Deformable Object Manipulation under Partial Observability via Real-Time Full-Shape EstimationKosar Behnia, Ville Kyrki, Gokhan Alcan · arXiv · Sep 9, 2026
Manipulating deformable objects (DOs) is challenging due to their high-dimensional state space, underactuated dynamics, and partial observability. In this paper, we propose cRVAE, a lightweight conditional recurrent variational autoencoder …
- FolDeX: A Physical-World Benchmark for Long-Horizon Robotic Manipulation of Deformable ObjectsChenhuan Liu, Yi Xu, Feng Wu, Hanyang Wang et al. · arXiv · Sep 9, 2026
Embodied AI, including vision-language-action and world-action models, must operate reliably in the physical world. Yet methods that perform well in simulation can degrade substantially on real robots, especially in long-horizon deformable-…
- Multi-Robot Scanner for Automated Full-Body Dermoscopic ImagingValerio Franchi, Rafael Garcia, Nuno Gracias, Ricard Campos et al. · arXiv · Sep 9, 2026
This paper outlines the specifications and design approach used to construct a full body imaging scanner capable of capturing skin lesions at a dermatoscopic level using cameras mounted on the end-effectors of four UR10 manipulators. The sy…
- TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action ModelAnqi Li, Yuxin Chen, Zhaobo Li, Zhuo Cao et al. · arXiv · Sep 8, 2026
We study the problem of navigating cluttered indoor environments with a humanoid robot. Unlike conventional methods that model navigation as a 2D path planning problem, humanoid traversal in cluttered environments requires continuous geomet…
- Proxy Policy SteeringChuanruo Ning, Tianrui Wang, Wei-Chiu Ma, Kuan Fang · arXiv · Sep 8, 2026
Generalist robot policies carry broad manipulation priors from large-scale data, but specializing them to a new task remains the deployment bottleneck. This requires eliciting task-specific behavior from limited demonstrations without degra…
- DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-ImaginationYankai Fu, Ning Chen, Junkai Zhao, Heng Zhang et al. · arXiv · Sep 8, 2026
Dexterous manipulation involves contact-rich and fine-grained interactions with the physical world, posing significant challenges for existing vision-language-action (VLA) models due to severe visual occlusions and complex contact dynamics.…
- Online, Reachability-Aware, Sampling-Based Motion PlanningBrendan Gould, Zhiyuan Zhang, Panagiotis Tsiotras, Samuel Coogan · arXiv · Sep 8, 2026
Sampling-Based Model-Predictive Control (MPC) algorithms are a flexible class of controllers used for navigation on a wide range of robotic systems. Historically, such approaches have lacked hard safety guarantees, a shortcoming which we re…
- DYAD: A Multimodal Dataset of Co-Located Human AssistanceAkhil Ajikumar, Mahya Qorbani, Sakib Reza, Sean Andrist et al. · arXiv · Sep 8, 2026
An embodied assistant working beside a person must track task state, recognize help seeking, choose how to intervene, and produce an appropriate response. Existing procedural datasets richly describe individual execution, while interactive …
- Visible-Reachable Workspace for Perception-Aware Humanoid DesignBoxi Xia, Zijiang Yang, Ryan Shin, Bokuan Li et al. · arXiv · Sep 8, 2026
Workspace analysis measures where a robot can place its end effector. For visually guided manipulation, reachability alone is insufficient: a kinematically reachable target may not be visible in the specific pose required to reach it. The r…
- Graph-Based Safe Reinforcement Learning for Multi-Agent Systems with Time-Varying TopologyXiao Sizhe, Dong Lijing, Bai Rui, Tan Xin · arXiv · Sep 8, 2026
This paper presents a graph-based safe multi-agent reinforcement learning (MARL) framework for cooperative navigation with time-varying topology. To address the critical challenge of ensuring safety in environments with sensing constraints,…
- FOCI Policy: Focus on Object-Centric Interactions for Relational Manipulation PoliciesZe Fu, Pinhao Song, Yutong Hu, Renaud Detry · arXiv · Sep 8, 2026
Object-centric manipulation policies improve generalization by modeling object motion instead of directly predicting robot actions. However, existing methods are often limited by representations which are either too simplistic to capture in…
- DCLP++: Learning to Navigate with Footprint Clearance and Relative MotionShanze Wang, Wei Zhang · arXiv · Sep 8, 2026
We present DCLP++, a local navigation frameworkthat uses footprint clearance as the geometric basis for studying relative motion features in dynamic environments. Each valid LiDAR return is mapped to its shortest Euclidean distance from the…