Latest Spatial AI & SLAM Research Papers
The newest Spatial AI & SLAM papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Spatial AI & SLAM so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Spatial AI & SLAM papers in your inbox — free →Recent papers
- Visual-SLAM for the detection of hidden tomatoes in greenhouses by Hierarchical Localization and GLOMAPfor robotized harvestingFernando Cañadas-Aránega, José C. Moreno, José L. Blanco-Claraco, Francisco Rodríguez · arXiv · Sep 10, 2026
Advanced crop monitoring inside greenhouses is becoming one of the primary objectives of research centers. High-performance sensors, such as LiDAR or stereo cameras, have traditionally been employed for this purpose, though these often have…
- Harness Robotic OS: A Unified Embodied-Agent Runtime for Closed-Loop Quadruped InspectionYaoyuan Yan, Zhiyou Heng, Haoxiang Jie, Gang Liu et al. · arXiv · Sep 10, 2026
Autonomous property inspection requires more than robust robot navigation: a deployable system must connect heterogeneous sensing, reusable autonomy capabilities, multimodal scene understanding, human interaction, and enterprise response wi…
- Spheriverse: 3D Scene Understanding from Spherical Observations in the WildFei Teng, Sheng Wu, Mengfei Duan, Guoqiang Zhao et al. · arXiv · Sep 8, 2026
Spherical observations provide global visual context for 3D scene understanding. However, visual information is encoded in an angular domain, whereas the physical world is represented in Cartesian coordinates. This cross-space representatio…
- FIRE-LIVWO: Robust LiDAR-Inertial-Visual-Wheel Odometry via Failure-Immune mmWave Radar EnhancementKun Hu, Menggang Li, Kaidi Wu, Zhiwen Jin et al. · arXiv · Sep 4, 2026
Achieving robust SLAM in large-scale underground coal mines with complex structures and severe degeneracies remains highly challenging. Dense smoke and dust cause substantial loss of visual information and degrade LiDAR point-cloud features…
- GraFT: A Training-Free Framework for Spatial Reasoning in Multimodal Large Language Models via 3D Scene GraphsJunqing Du, Fernando Ropero, Erkin Turkoz, Yanfeng Zhang et al. · arXiv · Sep 3, 2026
3D spatial reasoning underpins understanding and acting in the physical world, yet it remains unreliable in current multimodal large language models (MLLMs). These models falter at precise geometric measurement, at transforming between egoc…
- A hybrid pipeline for dynamic ontology-based semantic mappingKonstantinos Dimitropoulos, Ioannis Hatzilygeroudis · arXiv · Sep 3, 2026
Semantic mapping plays a crucial role in the ability of a robot to interact with objects, operate and navigate a complex environment. The most common pipeline for semantic mapping consists of geometric mapping and localization (SLAM), perce…
- MS-MEM: Multi-Skill Manipulation-Enhanced Mapping via Uncertainty- and Disturbance-Aware Action SelectionYitian Shi, Jesper Mücke, Nils Dengler, Sicong Pan et al. · arXiv · Sep 2, 2026
Accurate scene understanding in confined, cluttered spaces such as shelves is essential for service robots, as many everyday tasks require them to locate and retrieve objects reliably. Yet, it remains challenging due to severe occlusions, r…
- SG-AMP: Scene-Graph-Guided Active Perception and Semantics-Aware Motion Planning for Pepper PlantsRohit Menon, Shiva Rudra Lolla, Niklas Mueller-Goldingen, Gokul Chenchani et al. · arXiv · Sep 1, 2026
We present SG-AMP, integrating robust depth completion with input-conditioned uncertainty, persistent panoptic mapping, plant scene-graph reasoning, and semantics-aware active view-motion planning. Beyond inspecting uncertain observed regio…
- Failure or Drift? Evaluating Monocular SLAM under Synthetic and Real-World CorruptionsAbhay Skaria Thomas, Shashank Agnihotri, Margret Keuper · arXiv · Aug 31, 2026
Visual SLAM is commonly evaluated on clean trajectories, although deployment failures are often caused by adverse weather, illumination, blur, and sensor artifacts. Controlled corruptions are attractive because they isolate such factors, bu…
- STEGNav: Spatio-Temporal Event Graph Reasoning for Multimodal Lifelong Object NavigationYang Chen, Zhenyu Huang, Wenbo Fu, Danyang Peng et al. · arXiv · Aug 28, 2026
Multimodal lifelong navigation requires an agent to autonomously explore unseen environments while sequentially completing navigation tasks specified by object categories, language descriptions, or reference images. Existing methods primari…
- CGS-SLAM: Collaborative Gaussian Splatting based SLAM for Multi-Agent ReconstructionJean-Daniel de Ambrogi, Aladine Chetouani, Vincent Nguyen, Aurélien Chateigner · arXiv · Aug 27, 2026
Recent advances in SLAM have leveraged 3DGS for photorealistic reconstruction and novel view synthesis. However, most methods rely on RGB-D input, which is unavailable on consumer-grade smartphones, and few integrate 3DGS within a collabora…
- LT-Mem: Volatility-Aware Spatio-Temporal Memory for Lifelong Scene UnderstandingYumin Lee, Hyoseok Ju, Giseop Kim · arXiv · Aug 19, 2026
Long-term robot operation in evolving environments requires object-level understanding that persists across repeated revisits. Existing systems either overwrite history to maintain an up-to-date map or store semantic snapshots without consi…
- Evaluation of Monocular SLAM Systems on High-Altitude Nadir UAV FootageGašper Spagnolo, Matej Dobrevski, Danijel Skočaj · arXiv · Aug 19, 2026
Aerial nadir video combines weak geometric constraints with severe perceptual aliasing, making it a difficult regime for monocular SLAM. We benchmark five monocular SLAM systems on local UAV flights, synthetic city-scale imagery, and long-r…
- Jetson-ORB-SLAM3: Accuracy-Preserving GPU Implementation for Edge Computing DevicesRajat Roy, Aditya Arun Kumar Yadav, Hardik Jain · arXiv · Aug 18, 2026
Visual-inertial SLAM on low-power edge platforms is constrained by the cost of dense feature extraction and loop closure. Prior GPU ports of ORB-SLAM trade accuracy for speed by approximating the ORB detector, altering the feature set and t…
- OVIP-SG: Open-Vocabulary Instance-Preserving Scene Graphs for Mapping and Retrieval of Small, Fine-Grained ObjectsTianjing Hao, Haiyu Lan, Angsong Li, Cheng Chen et al. · arXiv · Aug 18, 2026
Integrating open-vocabulary perception into object-level 3D scene graphs is a double-edged sword. While vision-language detectors recover long-tail categories and small, fine-grained objects overlooked by closed-set models, they also tend t…
- Scalix: Uncertainty-Aware Scale-Consistent Monocular SLAMSebastian Barbas Laina, Tianyi Zhang, Panagiotis Petropoulakis, Simon Schaefer et al. · arXiv · Aug 18, 2026
Cameras are ubiquitous sensors in robotics due to their compact form factor and the perceptual richness captured through visual information. Monocular SLAM enables robots to understand the environment with a minimum setup, however, it inher…
- ViHaTeleop: A Low-Cost, Lightweight Visual-Haptic Teleoperation System for Dexterous Manipulation LearningFucai Zhu, Yanhou Lai, Paul Maestre, Koichi Hashimoto · arXiv · Aug 17, 2026
Learning from demonstration is a promising approach for dexterous manipulation, but collecting high-quality contact-critical demonstrations remains difficult with low-cost teleoperation hardware. We present ViHaTeleop, a lightweight (0.7 kg…
- From Principles to Practice: Engineering Responsible AI for Geospatial IntelligenceMuhammad Hassan, Bilal Sardar, Shareeful Islam, Maryam Imani et al. · SN Computer Science · Aug 13, 2026
Abstract AI-enabled geospatial applications increasingly inform high-stakes decisions in crop type classification, flood risk assessment, and land use monitoring; yet current practice prioritises predictive accuracy whilst leaving responsib…
- Prior-SG: Task and Prior Driven Region Segmentation for Scene Graphs in Arbitrarily-Structured EnvironmentsGiorgio Tonetti, Laurent Kneip, Abel Gawel, Marco Hutter · arXiv · Aug 6, 2026
Hierarchical 3D scene graphs are a promising representation for high-level spatial reasoning in autonomous mobile platforms. However, existing extraction frameworks typically rely on purely local visual clustering or strict geometric heuris…
- Topometric Autonomous Vehicle Localization by Combining Visual Embeddings and Feed-Forward 3D ModelsEulogio Quemada-Torres, Alberto Jaenal, Francisco-Angel Moreno, Javier Gonzalez-Jimenez · arXiv · Aug 6, 2026
Effective Visual Localization (VL) requires a map of the environment that combines compactness for efficient scalability with robustness against visual appearance changes and metric precision. Through low-dimensional image embeddings, Visua…
- SLAMFormer-$\infty$: Infinite SLAM Transformer for Unbounded Frontend and Backend ProcessingZhijian Fang, Weicheng Zheng, Yijun Yuan, Weibang Wang et al. · arXiv · Aug 4, 2026
We introduce the Infinite SLAM Transformer (SLAMFormer-$\infty$), the first geometric transformer capable of supporting both long-range frontend and backend processing without an explicit distance bound. Instead of relying on a first-frame-…
- Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action ModelsJin Cui, Yanbin Hu, Xinyue Long, Linkai Li et al. · arXiv · Aug 3, 2026
Visual representations of VLA models remain unreliable for spatially precise robotic manipulation. We uncover that vision encoders in VLAs also exhibit attention artifacts previously documented in generic Vision Transformers, and further sh…
- Accuracy potential of visual localization exploiting high-end street-level imageryJonas Meyer, Stephan Nebiker, Pascal Theiler, Norbert Haala · arXiv · Jul 27, 2026
Accurate and reliable pose information with respect to a reference frame is increasingly demanded across applications such as autonomous navigation, surveying, robotics, and augmented and mixed reality. Visual localization can serve as a co…
- GLAM-SLAM: Real-time Gaussian Large-scale Mapping via Flow Densification and Spatial DecompositionPanagiotis Mermigkas, Argyris Manetas, Petros Maragos · arXiv · Jul 23, 2026
Existing Gaussian-splatting-based monocular Simultaneous Localization and Mapping (SLAM) systems are either tailored to short sequences, are not real-time, or suffer from prohibitive GPU memory requirements, limiting their applicability in …
- DINS-IO: Learned Inertial Odometry via Differentiable INS ConsistencyHao Qiao, Yan Wang, Jian Kuang, Xiaoji Niu · arXiv · Jul 22, 2026
The training of learned inertial odometry depends on dense, high-precision position ground truth from motion capture, visual-inertial odometry or SLAM, which is costly and hard to acquire at scale. We propose DINS-IO, which learns inertial …
- Cognitive Dual-Process Planning for Autonomous Driving with Structured Scene Knowledge and Verifiable Reasoning-Action ConsistencyZhongyao Yang, Haoyu Li, Yu Yan, Zhuangxuan Yu et al. · arXiv · Jul 21, 2026
High-level planning for autonomous driving is a knowledge-intensive engineering decision task that requires accurate scene understanding, timely inference, and internally consistent action selection. Vision-language models (VLMs) can make i…
- Driving Scene Understanding: How much temporal context and spatial resolution is necessary?Ramashish Gaurav, Bryan P. Tripp, Apurva Narayan · Canadian AI 2021 · Jan 1, 2021
Driving Scene Understanding is a broad field which addresses the problem of recognizing a variety of on-road situations; namely driver behaviour/intention recognition, driver-action causal reasoning, pedestrians’ and nearby vehicles’ intent…