Latest Object Detection Research Papers
The newest Object Detection papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Object Detection so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Object Detection papers in your inbox — free →Recent papers
- MC-DeTra: Motion-Consistent Joint Object Detection and Socially-Aware Trajectory Forecasting in Bird's-Eye-View ImagesVladislav Diuzhev, Dmitry Yudin · arXiv · Sep 10, 2026
Unified models for object detection and trajectory forecasting aim to merge perception and prediction for autonomous driving, refining actor trajectories directly over shared bird's-eye-view (BEV) images rasterized from LiDAR and high-defin…
- FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow EstimationVladislav Bargatin, Alexander Yakovenko, Khaled Abud, Dmitriy Vatolin · arXiv · Sep 10, 2026
Optical flow methods typically rely on task-specific inductive biases, such as correlation volumes, feature warping, and iterative refinement, among others, to reach high accuracy. While effective, such biases constrain the model to predefi…
- Your Model Already Knows Don't Teach It, Learn to Ask It: Soft Prompting for Few-Shot Adaptation of Vision-Language ModelsGautam Rajendrakumar Gare, Siyi Li, Hewei Wang, Cesar Daniel Hernandez et al. · arXiv · Sep 10, 2026
We address few-shot object detection with vision-language models (VLMs) in out-of-domain settings such as aerial, industrial, and medical imagery, using only ten annotated images for supervision. Existing adaptation methods are discrete pro…
- Improving Faint Object Detection for Space Situational Awareness with Variational AutoencodersAngela Cratere, Luca Ghilardi, Vishnu Reddy, Francesco Dell'Olio et al. · arXiv · Sep 10, 2026
We present a deep-learning pipeline for enhancing the detection of faint moving objects in optical space situational awareness (SSA) imagery through automated star removal and background reconstruction. Detecting low signal-to-noise ratio (…
- Segmentation Residual–Gated Multi-Scale Region Enhancement in Monocular 3D Object DetectionTao Peng, Jinsu An, Byeong-Woo Kim · Transactions of Korean Soci... · Sep 10, 2026
- A Power-Efficiency Analysis and Evaluation Framework for Embedded AI-Based Object Detection ModelsJinho Yoo, Seokkan Ki, Seok-Cheol Kee · Transactions of Korean Soci... · Sep 10, 2026
- UAV Based Automated Surveillance of Ganoderma boninense in Oil Palm Canopies Using YOLO26 ArchitectureMuhammad Rizky Pribadi, Hafiz Irsyad, Eka Puji Widiyanto, Muhammad Tri Setianto et al. · Journal of Embedded Systems... · Sep 6, 2026
Purpose – This study develops and evaluates a UAV-based automated surveillance approach using the YOLO26 architecture to detect visible symptoms associated with Ganoderma boninense infection in oil palm canopies. The study addresses the lim…
- Efficient Multi-Timescale Event Representations for Feed-Forward Object DetectionFredrik Lundell, Per-Erik Forssen, Mårten Wadenbäck, Astrid Lundmark · arXiv · Sep 4, 2026
Autonomous systems require robust low-latency perception under rapidly changing scene dynamics and challenging illumination. In event cameras object detection commonly relies on recurrent architectures to accumulate sparse temporal informat…
- SimFuse3D: Source-Guided Target Simulation and Confidence-Guided Multi-Stage Localization Reweighting for Cross-Platform 3D Object DetectionYongchun Lin, Xinliang Zhang, Yun Zou, Zhixuan Xiao et al. · arXiv · Sep 4, 2026
Changes in sensor height and viewpoint alter object-level point distributions, making cross-platform LiDAR unsupervised domain adaptation (UDA) difficult. Self-training uses labeled source scans and unlabeled target scans, yet a retained pr…
- Moving Object Detection from Moving Camera Using Focus of Expansion Likelihood and SegmentationMasahiro Ogawa, Qi An, Atsushi Yamashita · International Journal of Au... · Sep 4, 2026
Separating moving and static objects from a moving camera viewpoint is essential for 3D reconstruction, autonomous navigation, and scene understanding in robotics. Existing approaches often rely primarily on optical flow, which struggles to…
- Stereo 4D Radar for 3D Object Detection: Integrating Geometric Alignment and Absolute Velocity EstimationSeung-Hyun Song, Dong-Hee Paek, Woong-Chan Byun, Seung-Hyun Kong · arXiv · Sep 2, 2026
Four-dimensional (4D) Radar is a powerful sensing modality capable of detecting surrounding three-dimensional (3D) objects under diverse weather conditions and providing Doppler-based motion information. However, raw 4D Radar signals contai…
- Information Density Imbalance in Visual Object DetectionZiwei Zhao, Yanxi Lu, Yuwei Hu, Shiyang Su et al. · arXiv · Sep 2, 2026
In object detection, the number of instances is typically used to determine whether a dataset exhibits a long-tailed distribution, implicitly assuming that the model will perform poorly on categories with fewer instances. This assumption ha…
- Domain shift-robust object detection with GenAI image editingIsabel D. Stein, Thijs A. Eker, Sebastiaan P. Snel, Ella P. Fokkinga et al. · arXiv · Sep 2, 2026
Object detectors often degrade under domain shifts such as changes in lighting, weather, or occlusion. These shifts alter object appearance and expose a reliance on visual shortcuts learned from the training distribution that do not general…
- If It Moves, Radar Knows: A Physics-Aware Radar Transformer for Class-Agnostic Moving-Object DetectionYinghao Sun, Shuguang Li, Jinliang Shao, Tieshan Li · arXiv · Sep 2, 2026
Detectors trained on closed-set annotations can miss rare moving objects outside the training taxonomy. Automotive radar provides category-independent Doppler motion cues and is less affected by adverse illumination and weather, but sparse,…
- An improved RT-DETR algorithm for small-object detection in UAV aerial imagesQiyu Long, Zhixun Liang, Peng Chen, Peng Tang · Scientific Reports · Sep 2, 2026
Abstract To address the challenges of UAV aerial imagery, including the prevalence of small objects, complex background interference, and difficulty in feature extraction that lead to high missed detection rates and compromise detection acc…
- Real-Time Video Anomaly Detection Using YOLO Pose Estimation and CLIP-Based Semantic ScoringVanodhya G. Warnasooriya, Amir Hajian, Watchara Ruangsang, Supavadee Aramvith · arXiv · Aug 31, 2026
We propose a lightweight two-stage framework for real-time video anomaly detection. The first stage employs YOLO v11n-pose to detect persons and extract seventeen skeletal keypoints in a single forward pass. The second stage encodes each cr…
- A Composition-Aware Pretraining Framework for Geospatial Foundation ModelsAryan Kashyap Naveen, Abhishek Srinivas, Pranav Moothedath, Shrutilipi Bhattacharjee · arXiv · Aug 31, 2026
Geospatial foundation models have emerged as state-of-the-art methods for downstream Earth observation tasks. However, existing pretraining methodologies process imagery through a single-concept lens, failing to capture the highly compositi…
- RailGen: Improving Railway Intrusion Detection via Agent-Guided Small-Scale Foreign Object GenerationQuan Hao, Ziyang Tao, Chenxi Zhang, Yudong Wang et al. · arXiv · Aug 31, 2026
Small-object detection under long-tailed data distributions is a fundamental yet challenging problem in multimedia. Railway Foreign Object Detection (RFOD) epitomizes this challenge with easily confused small intrusions and scarce samples. …
- RailSyn: Diagnosis-Guided Image Generation for Traceable Data Completion in Railway Foreign Object DetectionQuan Hao, Chenxi Zhang, Ziyang Tao, Yuyuan Zhou et al. · arXiv · Aug 31, 2026
Railway foreign object detection (RFOD) is critical to safe railway operation, yet scarce real positive samples incompletely represent task-relevant variations in object scale, intrusion relation, railway scene, illumination, and adverse we…
- GeBDA: Building Damage Assessment as Text-Based Sequence PredictionOlivier Dietrich, Krishna Sapkota, Konrad Schindler, Genady Beryozkin · arXiv · Aug 28, 2026
Conventionally, Building Damage Assessment (BDA) is tackled either with dedicated network architectures or by fine-tuning geospatial image foundation models. In this work, we ask whether a general-purpose Vision-Language Model (VLM) can loc…
- WALDO: One-Shot Exemplar-Conditioned Object Detection in Cluttered ScenesKishor Datta Gupta, Ahmed Rafi Hasan, Md. Mahfuzur Rahman, Md. Sadman Haque et al. · arXiv · Aug 28, 2026
Locating a specific object instance in a cluttered scene using a single reference image and a short description, and reporting when that instance is absent, large vision-language models usually address this task. We ask whether the same cap…
- CF-YOLO: Context-Aware Feature Refinement for Camouflaged Industrial Micro-Defect DetectionXinda Yu, Kunxin Zheng, Chunan Yu, Qingbo Song et al. · arXiv · Aug 28, 2026
Automated detection of surface micro-defects on industrial components, such as copper tubes, is critically important for quality assurance but remains challenging due to the minute scale of anomalies and their visual camouflage against comp…
- TADP: Task-Aware Deformable Prediction for Single-Stage 3D Object DetectionSu Wang, Yaochen Li, Min Yang, Jiaohao Nie et al. · arXiv · Aug 27, 2026
Most single-stage 3D object detectors complete different tasks with the same extracted features. Nevertheless, it is impossible to project features into a common space that is adaptive for all the tasks. We present a novel task-aware deform…
- CODE: Cross-Modal Calibration and Dynamic Suppression for Open World Object DetectionHao Xu, Zhaoning Shi, Hehe Jin, Bo Ma · arXiv · Aug 27, 2026
Open World Object Detection (OWOD) built on multimodal foundation models often suffers from semantic ambiguity caused by unidirectional text-to-vision matching, while rigid outlier penalties may over-suppress unknown objects near known-clas…
- Geometric Feature-Based Adaptive Voxel Sampling Technique for LiDAR-Based 3D Object DetectionSeung-Tak Ra, Seung-Ho Lee · Journal of the Institute of... · Aug 27, 2026
- TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency DistillationXiaoda Yang, Yuxiang Liu, Kaiwen Zheng, Yuan Liu et al. · arXiv · Aug 25, 2026
Joint text-to-video-audio generation produces synchronized visual and acoustic content, but the long sampling trajectories and heterogeneous multimodal computation of large models make inference prohibitively expensive. We present TurboT2VA…
- Automated detection and classification of dental findings in radiographs using a computer vision (YOLO) model: a tool for enhancing dental educationBehnaz Shirgir, Gülsüm Aşiksoy, Usman Abidemi Sarumi, Fadi Al-Turjman · PeerJ Computer Science · Aug 21, 2026
There is considerable value to dental radiographs in the diagnosis and treatment planning of dentistry in modern practice, but the skill (and knowledge) required for a diagnostic interpretation is largely one of low skill but of high demand…
- Dual-Level Spatial–Frequency Collaborative Detector for Oriented Object Detection in Remote Sensing ImagesXuehuai Shi, Jingru Sun, Kun Yu, Zhihui Wei et al. · Remote Sensing · Aug 21, 2026
Oriented object detection (OOD) in remote sensing images (RSIs) suffers from insufficient feature representation caused by arbitrary rotation angles and small spatial resolutions. Existing spatial–frequency fusion paradigms merely implement…
- Open-Vocabulary 3D Object Detection with Co-Distillation Discovery and Dual Guidance Robust TrainingShangbo Yuan, Jie Xu, Xiaofeng Zhu, Na Zhao · arXiv · Aug 20, 2026
Recently, open-vocabulary 3D object detection (3D-OVD) has gained increasing attention for its ability to detect unseen objects in 3D scenes. Existing approaches typically adopt a two-stage pipeline that first discovers novel objects using …
- Where Grounding Accuracy Lives on the IoU Curve: Label-Free Inference-Time Boundary RefinementBo Ma · arXiv · Aug 20, 2026
Vision--language models can identify the correct referent while returning an imprecise bounding box. We study whether a frozen direct-answer model can use its own prediction to allocate one additional localized observation without accessing…