Latest Video Understanding Research Papers
The newest Video Understanding papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Video Understanding so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Video Understanding papers in your inbox — free →Recent papers
- A Framework for Understanding Moving Image Music Concerts: Repertoire, Reception and RecontextualisationElizabeth Hunt · University of Liverpool · Jan 1, 2029
This thesis investigates the phenomenon of moving image music concerts, focussing on the live orchestral presentation of music from film, television, and video games. With orchestras increasingly turning to such concerts to attract new audi…
- Ethical agency and worker precarity amid polycrisis: a protest narrative analysisZhangzhong Huang, Zaheer Abbas, Talib Hussain, Seung Won Lee · Scientific Electronic Libra... · Oct 1, 2026
Abstract: This study explores the ethical dimensions of worker precarity and protest agency in the context of ongoing polycrisis conditions in Gilgit-Baltistan, Pakistan. Drawing upon publicly available protest videos from various occupatio…
- Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video UnderstandingWeitong Cai, Hang Zhang, Yukai Huang, Yiqiao Xie et al. · arXiv · Sep 10, 2026
Long-video understanding on edge devices must reason over hours of content under tight compute and bandwidth budgets. Subsampling visual tokens loses temporal structure, while text-only video memories lose fine-grained visual attributes. We…
- Single-Stream Multi-Feature Fusion with Temporal Robustness for Gait Emotion RecognitionShirong Lyu, Silu Quan, Yixuan Ding, Chengpeng Wang · arXiv · Sep 10, 2026
3D skeleton-based gait emotion recognition faces high annotation costs, data scarcity, and poor generalization on heterogeneous data. This paper proposes SV-GCN, a single-stream multi-feature fusion framework with temporal invariance. We in…
- Vidu S2: Real-Time Interactive, Editable, and Spatial Video GenerationJintao Zhang, Kai Jiang, Jintao Chen, Xu Wang et al. · arXiv · Sep 10, 2026
We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. Moreover, we explore the feasibility of real-time spatial video generation for both V…
- MMGait: Benchmarking and Unifying Gait Recognition across Heterogeneous ModalitiesSaihui Hou, Chenye Wang, Qingyuan Cai, Aoqi Li et al. · arXiv · Sep 10, 2026
Gait recognition is commonly studied using RGB videos or their derived silhouettes and poses. Yet human walking produces heterogeneous photometric, geometric, and motion cues that cannot be systematically examined with RGB-centered benchmar…
- OmniKVQuant: KV Cache Quantization for Omni-LLMsSuho Yoo, Hyunjong Ok, Jongmin Choi, Jihoo Jung et al. · arXiv · Sep 10, 2026
As Omni-modal large language models (Omni-LLMs) take in audio, video and text together, their KV cache memory cost grows. KV cache quantization is the de facto approach in text-only LLMs, but its application to Omni-LLMs remains unexplored.…
- World in World: Explore the World with World ModelsChenxi Song, Yanming Yang, Chi Zhang · arXiv · Sep 10, 2026
Autoregressive video world models enable interactive, long-horizon exploration, but flexible control remains challenging. Exploring a source video from new viewpoints requires the generated rollout to remain synchronised with the recorded e…
- Harnessing Intrinsic Subject-Aware Attention for Controllable Multi-Subject Video GenerationNiange Yu, Ye Tian, Biaolong Chen, Miao Lu et al. · arXiv · Sep 10, 2026
Multi-subject video generation faces two key challenges: uncontrollable fidelity strength and potential semantic drift. We address these by analyzing the internal mechanisms of Diffusion Transformers (DiTs). We found that certain attention …
- FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow EstimationVladislav Bargatin, Alexander Yakovenko, Khaled Abud, Dmitriy Vatolin · arXiv · Sep 10, 2026
Optical flow methods typically rely on task-specific inductive biases, such as correlation volumes, feature warping, and iterative refinement, among others, to reach high accuracy. While effective, such biases constrain the model to predefi…
- Programmable World ModelZheng-Hui Huang, Guixu Lin, Jiacheng Lin, Yi-Chuan Huang et al. · arXiv · Sep 9, 2026
Recent video world models generate increasingly realistic and interactive visual experiences, yet lack reliable mechanisms for maintaining persistent world state and enforcing programmable rules over extended interactions. We introduce Prog…
- Field Converter: Geometry-Initialized Temporal Residual Refinement for World-Grounded Player Pose Estimation from Soccer BroadcastsSimon Khan, Laurent Gajny, Jennyfer Lecompte, Sébastien Laporte · arXiv · Sep 9, 2026
Recovering 3D human pose from monocular sports broadcasts remains challenging when players must be localized in a shared metric world coordinate system rather than only reconstructed relative to their own body. We introduce Field Converter,…
- Beyond Weak Labels: Prompt-Guided Local Refinement for Weakly Supervised Water Segmentation in High-Resolution Multispectral ImageryMuhammad Farhan Humayun, Mohammad Imangholiloo, Afifah Shah, Tomi Westerlund et al. · arXiv · Sep 9, 2026
High-resolution water mapping supports environmental monitoring and related applications, but accurate pixel-level labels are difficult and costly to produce. Official hydrographic vectors provide scalable weak supervision, but they contain…
- Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking RolloutZhuoran Zhao, Shengju Qian, Tongtong Liang, Xianghao Kong et al. · arXiv · Sep 8, 2026
Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods distill pretrained bidirectional video diffusion models into causal AR students through Distribution Matching Distillation (…
- PIC: Revisiting INR for Image Coding with Fast Encoding and Sub-Millisecond DecodingXiang Liu, Jinxiang Wang, Bin Chen, Zimo Liu et al. · arXiv · Sep 8, 2026
Implicit neural representation (INR) has achieved remarkable progress in novel view synthesis and image/video coding in recent years.Compared to conventional end-to-end image codecs, INR-based compressors demonstrate significant advantages …
- Concentrate After Imagination: Text-Conditioned Evidence Grounding for Partially Relevant Video RetrievalShuaiqi Cheng, Siyu You, Yanbi Wu, Yuxi Chen et al. · arXiv · Sep 8, 2026
Partially Relevant Video Retrieval (PRVR) retrieves untrimmed videos when queries describe only short moments. Although recent methods improve local representations, uncertainty modeling, and global context, final ranking often still trusts…
- EgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV ReasoningJingpu Yang, Fengxian Ji, Mingxuan Cui, Yilin Sun et al. · arXiv · Sep 8, 2026
UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimodal models receive no explicit, stable reference for that separation. We present EgoSIS, a pose-free adapter that converts RGB-der…
- WorldSculpt: Generating Compositional Worlds from Grounded VideosMuyao Niu, Jixuan He, Ruihan Yu, Lian Fu et al. · arXiv · Sep 4, 2026
We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds of objects. The goal is to represent the scene as a collection of individual object meshes placed in a shared world frame, as requ…
- Few-Shot Video Recognition via Hierarchical Metric LearningJiaxin Zhang, Haoran Gao, Xizhan Gao, Zihao Dong et al. · arXiv · Sep 4, 2026
Few-shot action recognition (FSAR) aims to recognize unseen action categories with only a small number of annotated video samples. Recent works typically apply single-prototype supervision at the network output and fail to sufficiently expl…
- From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-MakingDavide Testa, Hugh Mee Wong, Alessandro Lenci, Bernardo Magnini et al. · arXiv · Sep 4, 2026
Vision-Language Models are commonly evaluated through their final predictions, but understanding whether these decisions are grounded in visual evidence requires tracing how visual information contributes to language-based decisions. With t…
- Temporal Self-Distillation: Learning Visual State Tracking in Videos Without SupervisionShravan Venkatraman, Wenshuai Zhao, Mohammad Hassan Vali, Arno Solin · arXiv · Sep 3, 2026
We introduce S$^3$T (Self-Supervised Self-Distillation over Time), which, to the best of our knowledge, is the first fully self-contained framework for continuous video state tracking. Our method treats temporal sampling density as privileg…
- Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D ReconstructionChin-Yang Lin, Yang-Che Sun, Cheng Sun, Fu-En Yang et al. · arXiv · Sep 3, 2026
Online 3D reconstruction models perform poorly on long videos. This happens because regressing poses relative to a fixed first-frame anchor forces extrapolation far beyond the training distribution. Small drifts accumulate and amplify into …
- Principia: Relational Physics Tests for Video ModelsVarun Varma Thozhiyoor, Shivam Tripathi, Venkatesh Babu Radhakrishnan, Anand Bhattad · arXiv · Sep 3, 2026
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate, object scale, and camera calibration, all of which are often ambiguous or unavailable in generated video. We propose a dif…
- One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video EditingAdheesh Sunil Juvekar, Onkar Kishor Susladkar, Kiet A. Nguyen, Muntasir Wahed et al. · arXiv · Sep 3, 2026
Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing within a single unified framework remains challenging. We introduce EditVid, a training-free framework combining sparse …
- Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video CaptioningYe-Chan Kim, Seunghee Choi, SeungJu Cha, Si-Woo Kim et al. · arXiv · Sep 3, 2026
Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos given only an ordered set of event-level captions per video. Recent work synthesizes auxiliary transition captions via LLM to provide…
- Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video UnderstandingHongyu Qu, Guangming Yao, Ling Xing, Xiaobin Hu et al. · arXiv · Sep 3, 2026
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual inputs and respond to user queries under strict causality and bounded memory. Existing approaches typically compress historical obs…
- BooM-VVT: Boosting Mask-Free Video Virtual Try-On with Image-Level Pseudo DataWei Zhang, Xin Li, Peishu Shi, Jialin Gao et al. · arXiv · Sep 3, 2026
Video virtual try-on (VVT) aims to generate realistic videos of a person wearing a target garment. Recent methods leverage a keyframe-driven video generation paradigm to improve in-the-wild performance, yet they still rely on masks to local…
- The Shape of Time: Video-Token Contrast for Temporal Understanding in VideoLMsYumeng Shi, Quanyu Long, Yin Wu, Wenya Wang · arXiv · Sep 3, 2026
Seeing frames in order does not mean representing time. Modern VideoLMs receive ordered video streams, yet their main supervision acts on generated text rather than video-token representations where event dynamics should first emerge. This …
- DSAQuant: Denoising-Stage-Aligned Quantization-Aware Training for Video GenerationShuaiting Li, Zelin Gao, Haibin Shen, Yujun Shen et al. · arXiv · Sep 3, 2026
Video diffusion models (VDMs) have achieved impressive progress in text-to-video generation, but their high memory and computational costs hinder practical deployment. Quantization-aware training (QAT) is an effective solution for compressi…
- WorldReward: Reward Modeling for Camera-Conditioned World ModelsYibin Wang, Zehan Wang, Junshu Tang, Zhimin Li et al. · arXiv · Sep 3, 2026
Camera-conditioned world models generate interactive videos in which commanded actions should induce the expected scene changes while appearance, geometry, and temporal dynamics remain coherent. Existing rewards assess these requirements se…