Latest Multimodal Learning Research Papers
The newest Multimodal Learning papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Multimodal Learning so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Multimodal Learning papers in your inbox — free →Recent papers
- Conserved Multimodal Nonlinear Dynamics in Brain Activity and StructureCalvin Grant · Zenodo (CERN European Organ... · Dec 5, 2026
SUMMARY The brain has been generating the same rhythms — delta, theta, alpha, beta, gamma — in every subject ever recorded. The boundaries between them are reproducible to within a hertz across species, ages, anesthesia, wakefulness, epilep…
- Conserved Multimodal Nonlinear Dynamics in Brain Activity and StructureCalvin Grant · Zenodo (CERN European Organ... · Dec 5, 2026
SUMMARY The brain has been generating the same rhythms — delta, theta, alpha, beta, gamma — in every subject ever recorded. The boundaries between them are reproducible to within a hertz across species, ages, anesthesia, wakefulness, epilep…
- Towards Early and Accurate Disease Detection Through Multimodal Predictive Modeling: Fusion of Electronic Health Records, Medical Imaging, And Omics Data Using Interpretable Machine Learning.Muhammad Ahsan Hayat, Jahangir Baig, Shayan Ahmed, Ahmed Faraz Ayubi · Zenodo (CERN European Organ... · Nov 3, 2026
Early detection of disease is a cornerstone for improving patient outcomes, reducing costs, and enabling preventative interventions. Traditional predictive models often rely on a single type of data (e.g., imaging, clinical labs, or genomic…
- Towards Early and Accurate Disease Detection Through Multimodal Predictive Modeling: Fusion of Electronic Health Records, Medical Imaging, And Omics Data Using Interpretable Machine Learning.Muhammad Ahsan Hayat, Jahangir Baig, Shayan Ahmed, Ahmed Faraz Ayubi · Zenodo (CERN European Organ... · Nov 3, 2026
Early detection of disease is a cornerstone for improving patient outcomes, reducing costs, and enabling preventative interventions. Traditional predictive models often rely on a single type of data (e.g., imaging, clinical labs, or genomic…
- Impact of a Multimodal Diagnostic Stewardship Quality Improvement Initiative on Inpatient Urine Culture Utilization and Diagnostic Performance Across an 11-Hospital Health SystemJessica Garciga, Richard Levine, Timothy Gauthier, Denise Payne et al. · Scholarly Commons - Baptist... · Oct 23, 2026
- Cross-modal Image Recommendation for News Articles by Multimodal Foundation Models-based Retrieval-RerankingDamianos Galanopoulos, Andreas Goulas, Vasileios Mezaris · Open MIND · Oct 1, 2026
Retrieving relevant images for a given news article is challenging and can be considered a special version of the cross-modal retrieval problem. This notebook paper presents our solution for the MediaEval NewsImages 2025 benchmarking task. …
- AI-Driven Biomarker Discovery & Progression Modeling for Precision Diagnosis of GlaucomaCheng Huang · SMU Scholar (Southern Metho... · Oct 1, 2026
This dissertation presents a comprehensive study on the integration of artificial intelligence (AI) for glaucoma diagnosis and retinal image analysis. Leveraging multimodal imaging data including fundus photography, Optical Coherence Tomogr…
- SenseNova-U1.5: Towards Native Unified Visual IntelligenceHaiwen Diao, Jiahao Wang, Chenjing Ding, Hanming Deng et al. · arXiv · Sep 10, 2026
We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. We strengthen its visual interface through spatially coheren…
- Multimodal Taxonomic Conditioning for Generative Plankton ImageryDaniela Ivanova, Ozgu Goksu, Nicolas Pugeault · arXiv · Sep 10, 2026
Automated plankton imaging produces severely long-tailed datasets, where the rare taxa of greatest ecological interest have too few images to train or evaluate classifiers reliably. We generate synthetic plankton imagery conditioned on taxo…
- OmniKVQuant: KV Cache Quantization for Omni-LLMsSuho Yoo, Hyunjong Ok, Jongmin Choi, Jihoo Jung et al. · arXiv · Sep 10, 2026
As Omni-modal large language models (Omni-LLMs) take in audio, video and text together, their KV cache memory cost grows. KV cache quantization is the de facto approach in text-only LLMs, but its application to Omni-LLMs remains unexplored.…
- Prototype Matters: Modality-unified Prototype Self-distillation for Unsupervised Visible-infrared Person Re-identificationMenglin Wang, Xiaojin Gong · arXiv · Sep 10, 2026
Estimating reliable cross-modality association is crucial to unsupervised visible-infrared person re-ID. While optimal transport is shown to be a practical solution for cross-modality association, it suffers from the rigidness of hard label…
- FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow EstimationVladislav Bargatin, Alexander Yakovenko, Khaled Abud, Dmitriy Vatolin · arXiv · Sep 10, 2026
Optical flow methods typically rely on task-specific inductive biases, such as correlation volumes, feature warping, and iterative refinement, among others, to reach high accuracy. While effective, such biases constrain the model to predefi…
- Show-Harness: Just a VLM Agent Can Play RobotsYanzhe Chen, Zechen Bai, Zhijun Cao, Wenzheng Zeng et al. · arXiv · Sep 9, 2026
Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to "play" robots t…
- Studying Image Tokenizers as Visual Languages in Unified Multimodal ModelsSiting Li, Zhengyang Wang, Simon Shaolei Du, Xi Chen et al. · arXiv · Sep 8, 2026
Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through isolated metrics or generation-/understanding-only evaluations. These evaluations do not fully capture how visual tokens behave w…
- Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMsXiaofu Chen, Stella Frank, Yova Kementchedjhieva · arXiv · Sep 8, 2026
Visual encoders construct a representation of the image input for Vision-Language models. How much conceptual, as opposed to immediately visible, information does this representation contain? We use canonical color as a controlled test case…
- DXPR: Depth-Based Vision-LiDAR Cross-Modal Place Recognition Using Vision Foundation ModelsYungsoo Han, Youngseok Jang, Seungwon Roh, Jeongyeon Seo et al. · arXiv · Sep 8, 2026
We present DXPR, a depth-based cross-modal place recognition (CMPR) framework that uses vision foundation models (VFMs) to match monocular camera queries against a LiDAR map without modality-specific encoders. This enables robots and autono…
- EgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV ReasoningJingpu Yang, Fengxian Ji, Mingxuan Cui, Yilin Sun et al. · arXiv · Sep 8, 2026
UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimodal models receive no explicit, stable reference for that separation. We present EgoSIS, a pose-free adapter that converts RGB-der…
- Evolution of Multimodal Question Answering: From Modality-Adaptive Extraction to Unified Language RepresentationAbdullah Al Shafi · arXiv · Sep 8, 2026
The rapid growth of multimodal data has intensified the need for question answering (QA) systems capable of reasoning across heterogeneous sources such as text, tables, and images. In this paper, we present a comprehensive methodological co…
- Think-Verify-Revise: Neuro-Symbolic Visual Reasoning with Vision-Language Models and Dynamic Logic Tensor NetworksHomayoun Afshari, Pietro Basci, Alessandro Russo, Lia Morra · arXiv · Sep 4, 2026
Visual reasoning tasks require a system to jointly perceive visual content and apply formal relational constraints---a combination that neither pure neural nor purely symbolic approaches handle well in isolation. This paper proposes a Neuro…
- Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action ManipulationVivek Chavan, Yahuan Shi, Oliver Heimann, Kevin Haninger et al. · arXiv · Sep 4, 2026
Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon procedures requiring persistent task state, dependency-aware reasoning, conditional decisions, and reliable grounding. We investig…
- MEOX: Compact Multimodal Mixture-of-Experts for Earth ObservationMohanad Albughdadi · arXiv · Sep 4, 2026
Recent advances in Earth Observation representation learning accommodate heterogeneous sensors and missing observations, often through larger architectures. We present MEOX (Multimodal Earth Observation with eXperts), a multimodal masked au…
- RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan et al. · arXiv · Sep 4, 2026
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight in…
- Cross-Domain Tracker Adaptation Without Target-Domain Labels via Vision-Language AgentsDaniel Davila, Ravikumar Balakrishnan, Mike Cochran · arXiv · Sep 4, 2026
We present a system that uses a Vision-Language Model (VLM) as a diagnostic agent for adapting a detect-to-track pipeline to a new target domain without access to target-domain labels. Rather than optimizing against annotated metrics, the V…
- First Things First: Teaching LLM-Based Agents to Prioritize Must-Haves before Nice-to-HavesTianjie Ju, Xinyue Xu, Wanxuan Sun, Lingxiao Diao et al. · arXiv · Sep 4, 2026
Recent progress in multimodal large language models (MLLMs) has fueled significant enthusiasm in their potential to act as autonomous agents for real-world tasks. However, scenarios requiring agents to fulfill users' complex, structured req…
- Puffin-World: Scaling a Unified Multimodal Model with Native 3D World StatesKang Liao, Yihang Luo, Xiao-Ming Wu, Linyi Jin et al. · arXiv · Sep 3, 2026
We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules. To reliably construct and interac…
- Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video CaptioningYe-Chan Kim, Seunghee Choi, SeungJu Cha, Si-Woo Kim et al. · arXiv · Sep 3, 2026
Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos given only an ordered set of event-level captions per video. Recent work synthesizes auxiliary transition captions via LLM to provide…
- Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video UnderstandingHongyu Qu, Guangming Yao, Ling Xing, Xiaobin Hu et al. · arXiv · Sep 3, 2026
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual inputs and respond to user queries under strict causality and bounded memory. Existing approaches typically compress historical obs…
- Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp SynthesisSixu Yan, Shikang Wang, Binhua Huang, Xuanlai Tang et al. · arXiv · Sep 3, 2026
This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generalizable grasp synthesis across different robotic hands. Unlike existing VLG methods that tightly couple foundation models with end-to-…
- Continuous Actions from Discrete Minds: Latent-Aligned Planning for End-to-End Autonomous DrivingRuoyu Yao, Yusen Xie, Qingzhao Liu, Pei Liu et al. · arXiv · Sep 3, 2026
Bridging the gap between the discrete reasoning of Vision-Language Models and the continuous, physics-constrained nature of autonomous driving remains a significant challenge. In this work, we introduce LaPla, a unified Vision-Language-Acti…
- ShallowStream: Index Shallow then Answer Deep for Streaming Video UnderstandingJitai Hao, Ke Yang, Qiang Huang, Jun Yu · arXiv · Sep 2, 2026
Streaming video understanding is a critical capability for real-world applications, including embodied intelligence, autonomous driving, industrial monitoring, surveillance and early warning, and wearable assistants. However, processing con…