Latest Computer Vision Research Papers
The newest Computer Vision papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Computer Vision so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Computer Vision papers in your inbox — free →Recent papers
- SOAR: Smooth Online Activation Routing for Stable Neural Learning from Evolving StreamsSizhen Niu · OpenAlex · Dec 31, 2026
Online neural learning requires models that update after each incoming example, remain calibrated under distributional change, and avoid brittle gradient transmission. The original version of this work used a small static benchmark, a shall…
- Research on Attention Guidance and User Autonomy in AI-Powered Immersive EnvironmentsMulin Qiao · PhilPapers (PhilPapers Foun... · Dec 31, 2026
As artificial intelligence gets deeply embedded into immersive environments, attention-guiding technology delivers ever more precise and efficient outcomes. Still, this improved efficiency carries an overlooked design risk: excessive AI gui…
- X-ray crystal structures of the cannabinoid synthases CBCAS, CBDAS and THCASGIDEON JAMES GROGAN, Jack Domenech, Jared Cartwright, Andrew King et al. · White Rose Research Online ... · Dec 1, 2026
The enzymes Cannabichromenic Acid Synthase (CBCAS), Cannabidiolic Acid Synthase (CBDAS) and Tetrahydrocannabinolic Acid Synthase (THCAS) are together the major cannabinoid synthase enzymes responsible for the biosynthesis of their respectiv…
- Catalyst poisoning influences from various functional groups of energy carriers towards electrochemical oxidation reactions on non-noble high-entropy alloy anodes in acidic mediaTahawy Rafat, Muflihah Salma Aridha, Hara Kosuke, Ohto Tatsuhiko et al. · Institutional Repositories ... · Dec 1, 2026
Electrolytic synthesis of energy carriers using renewable energy and fuel cells that use energy carriers for regeneration are important technologies for achieving our carbon-neutral society. However, electrochemical reactions in electrolyte…
- Interpol review of gunshot residue, 2022 to 2024.S. Charles, Nadia Geusens, Clémence Tarenne, Florian Vanneste et al. · PubMed · Dec 1, 2026
- Developing Radiomic Neuroimaging Biomarkers Using Machine Learning for Predicting Post-Traumatic EpilepsyMark Chao · Digital Commons-TMC (Texas ... · Oct 7, 2026
- Binder-Free Mesoporous Vanadium Oxide Electrode: Anodic Electrodeposition, Characterization, and Supercapacitor ApplicationM Kazazi, Soheila Kazemi, Behzad Koozegar Kaleji · DOAJ (DOAJ: Directory of Op... · Oct 1, 2026
Mesoporous vanadium oxide (V2O5) was galvanostatically electrodeposited into nickel foam to obtain a binder-free electrode with a three-dimensional (3D) porous structure for supercapacitive performance. The anodic electrodeposition process …
- Polysaccharide-based compound from the fungus Pycnoporus sanguineus for the protection of tomato plants against Meloidogyne incognitaBruna L. de Oliveira, Thaísa M. Mioranza, José R. Stangarlin, Odair J. Kuhn et al. · Scientific Electronic Libra... · Oct 1, 2026
ABSTRACT New control strategies for Meloidogyne incognita are needed to meet the global demand for environmentally safe methods. The objective of this study was to use crude polysaccharide extract (CPE) from the basidiocarp of Pycnoporus sa…
- Natur – Tiere – Krieg. Beziehungen und Wechselwirkungen von der Antike bis zur GegenwartTU Dortmund University · Open MIND · Sep 16, 2026
Vom 16. bis zum 18. September 2026 findet an der TU Dortmund eine internationale Konferenz zum Thema “Natur – Tiere – Krieg. Beziehungen und Wechselwirkungen von der Antike bis zur Gegenwart” statt. Sie verfolgt das Ziel, die bislang häufig…
- SenseNova-U1.5: Towards Native Unified Visual IntelligenceHaiwen Diao, Jiahao Wang, Chenjing Ding, Hanming Deng et al. · arXiv · Sep 10, 2026
We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. We strengthen its visual interface through spatially coheren…
- MindTopo: Can Foundation Models Reason in Topological Space?Yunfei Ge, Anbang Liu, Qineng Wang, Johnalbert Garnica et al. · arXiv · Sep 10, 2026
Spatial reasoning depends not only on metric properties such as distance, angle, and shape, but also on topological relations that remain invariant under continuous deformation. Cognitive science identifies these relations as foundational t…
- Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video UnderstandingWeitong Cai, Hang Zhang, Yukai Huang, Yiqiao Xie et al. · arXiv · Sep 10, 2026
Long-video understanding on edge devices must reason over hours of content under tight compute and bandwidth budgets. Subsampling visual tokens loses temporal structure, while text-only video memories lose fine-grained visual attributes. We…
- 3D Point Splatting for mmWave Radar Novel View SynthesisAdnan Armouti, Yixuan Gao, Rajalakshmi Nandakumar · arXiv · Sep 10, 2026
Solving novel view synthesis (NVS) for millimeter-wave (mmWave) radar requires a renderer that is physically faithful, complex-valued, and multi-viewpoint-tractable. No prior method achieves these three properties simultaneously. Differenti…
- Guided Super-Resolution of Digital Elevation Models with Diffusion-Based Image GeneratorsArmand Mihai Nicolicioiu, Dominik Narnhofer, Nando Metzger, Daniel Panangian et al. · arXiv · Sep 10, 2026
High-resolution digital surface models (DSMs) play an important role in urban analysis, 3D building reconstruction, and infrastructure monitoring, yet their availability remains limited due to the high cost and complexity of data acquisitio…
- CoRA-NAS: Coarse Ranking and Anchor-Residual Refinement for Neural Architecture SearchYifan Yang, Zhaoyan Wang, Zheng Gao, Xiaoyu Li et al. · arXiv · Sep 10, 2026
Zero-cost proxies rank architectures cheaply, but their reliability varies across search spaces. We introduce CoRA-NAS (COarse Ranking + Anchor-residual), a two-stage framework combining a static ranking prior with low-cost learning-curve r…
- Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency ModelingMeimingwei Li, Stefan Andreas Baumann, Felix Krause, Björn Ommer · arXiv · Sep 10, 2026
Visual Autoregressive Models (VAR) generate images through next-scale prediction, producing all tokens within each scale in parallel. We show that this parallel decoding constitutes a mean-field-style approximation that discards spatial dep…
- Revisiting Avatar-As-Image: High-Fidelity Registration is All You NeedMargaret Kostyrko, Yuxuan Xue, Garvita Tiwari, Gerard Pons-Moll · arXiv · Sep 10, 2026
The representation of 3D clothed humans as standardized 2D UV texture and displacement maps over an underlying body model has long been studied. This compact representation is enticing as it enables pretrained image networks to process, gen…
- MC-DeTra: Motion-Consistent Joint Object Detection and Socially-Aware Trajectory Forecasting in Bird's-Eye-View ImagesVladislav Diuzhev, Dmitry Yudin · arXiv · Sep 10, 2026
Unified models for object detection and trajectory forecasting aim to merge perception and prediction for autonomous driving, refining actor trajectories directly over shared bird's-eye-view (BEV) images rasterized from LiDAR and high-defin…
- Language-Augmented Semantic Priors for B-Spline Surface FittingYunzhong Lou, Yusheng Luo, Jiahao Li, Yu Song et al. · arXiv · Sep 10, 2026
The use of B-splines and Non-Uniform Rational B-Splines surfaces constitutes the mathematical foundation of contemporary computer-aided design (CAD) systems. Despite long-term progress, geometric kernels in traditional CAD still rely heavil…
- Spectral Adapters for Segment Anything Model-based Segmentation of Colorectal Liver Metastases in Computed TomographyRamtin Mojtahedi, Mohammad Hamghalam, Jacob J. Peoples, Natalie Gangai et al. · arXiv · Sep 10, 2026
Accurate segmentation of colorectal liver metastases (CRLM) in contrast-enhanced computed tomography (CT) is important for response assessment, surgical planning, and follow-up. We propose two parameter-efficient spectral adapters for the S…
- Single-Stream Multi-Feature Fusion with Temporal Robustness for Gait Emotion RecognitionShirong Lyu, Silu Quan, Yixuan Ding, Chengpeng Wang · arXiv · Sep 10, 2026
3D skeleton-based gait emotion recognition faces high annotation costs, data scarcity, and poor generalization on heterogeneous data. This paper proposes SV-GCN, a single-stream multi-feature fusion framework with temporal invariance. We in…
- Multimodal Taxonomic Conditioning for Generative Plankton ImageryDaniela Ivanova, Ozgu Goksu, Nicolas Pugeault · arXiv · Sep 10, 2026
Automated plankton imaging produces severely long-tailed datasets, where the rare taxa of greatest ecological interest have too few images to train or evaluate classifiers reliably. We generate synthetic plankton imagery conditioned on taxo…
- Self-Supervised Cardiac Phase Detection via Single-Parameter Latent OrbitsJohn Bonnici, Matthew Baugh, Aleksandra Kulbaka, Sarah Cechnicka et al. · arXiv · Sep 10, 2026
Accurate identification of end-diastole (ED) and end-systole (ES) in echocardiography underpins the quantification of ventricular function, yet manual selection of these key frames is subjective and introduces clinically significant inter-o…
- Vidu S2: Real-Time Interactive, Editable, and Spatial Video GenerationJintao Zhang, Kai Jiang, Jintao Chen, Xu Wang et al. · arXiv · Sep 10, 2026
We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. Moreover, we explore the feasibility of real-time spatial video generation for both V…
- LangStreet: Persistent Language Fields for Anchor-Decoded Street GaussiansRunyi Yang, Deheng Zhang, Xiaoye Wang, Mengjiao Ma et al. · arXiv · Sep 10, 2026
Language Gaussian fields implicitly assume that the primitive carrying semantics remains identifiable across views. This assumption breaks in scalable anchor-decoded representations, where persistent anchors generate view-conditioned child …
- MMGait: Benchmarking and Unifying Gait Recognition across Heterogeneous ModalitiesSaihui Hou, Chenye Wang, Qingyuan Cai, Aoqi Li et al. · arXiv · Sep 10, 2026
Gait recognition is commonly studied using RGB videos or their derived silhouettes and poses. Yet human walking produces heterogeneous photometric, geometric, and motion cues that cannot be systematically examined with RGB-centered benchmar…
- OmniKVQuant: KV Cache Quantization for Omni-LLMsSuho Yoo, Hyunjong Ok, Jongmin Choi, Jihoo Jung et al. · arXiv · Sep 10, 2026
As Omni-modal large language models (Omni-LLMs) take in audio, video and text together, their KV cache memory cost grows. KV cache quantization is the de facto approach in text-only LLMs, but its application to Omni-LLMs remains unexplored.…
- Learn the Solid, Not the File: Canonical Inputs for Neural Networks on CAD Boundary RepresentationsHeinrich Jiang, Hager Yasser Mohamed, Alexander Hitt, Valeriia Lomakina et al. · arXiv · Sep 10, 2026
Boundary representation (B-rep) is the standard format used by modern CAD systems for parametric 3D models. It turns out, the exact same solid can be represented by different B-reps: for example, two engineers using different operations, a …
- A Comparative Evaluation of Pre-trained Convolutional Neural Networks for Melanoma DetectionWagner Moreno Schmitz, Marco Antonio de Castro Barbosa, Thiago Magalhães Amaral, Dalcimar Casanova et al. · arXiv · Sep 10, 2026
Early diagnosis of melanoma is critical for improving patient survival rates. However, accurately distinguishing melanoma from other skin lesions remains a significant clinical challenge due to the high visual similarity among lesion types …
- World in World: Explore the World with World ModelsChenxi Song, Yanming Yang, Chi Zhang · arXiv · Sep 10, 2026
Autoregressive video world models enable interactive, long-horizon exploration, but flexible control remains challenging. Exploring a source video from new viewpoints requires the generated rollout to remain synchronised with the recorded e…