Latest Video Generation Research Papers
The newest Video Generation papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Video Generation so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Video Generation papers in your inbox — free →Recent papers
- Vidu S2: Real-Time Interactive, Editable, and Spatial Video GenerationJintao Zhang, Kai Jiang, Jintao Chen, Xu Wang et al. · arXiv · Sep 10, 2026
We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. Moreover, we explore the feasibility of real-time spatial video generation for both V…
- Harnessing Intrinsic Subject-Aware Attention for Controllable Multi-Subject Video GenerationNiange Yu, Ye Tian, Biaolong Chen, Miao Lu et al. · arXiv · Sep 10, 2026
Multi-subject video generation faces two key challenges: uncontrollable fidelity strength and potential semantic drift. We address these by analyzing the internal mechanisms of Diffusion Transformers (DiTs). We found that certain attention …
- FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow EstimationVladislav Bargatin, Alexander Yakovenko, Khaled Abud, Dmitriy Vatolin · arXiv · Sep 10, 2026
Optical flow methods typically rely on task-specific inductive biases, such as correlation volumes, feature warping, and iterative refinement, among others, to reach high accuracy. While effective, such biases constrain the model to predefi…
- Uncertainty DMD: Restoring Diversity in Few-Step Autoregressive Video DistillationZixuan Duan, Xunzhi Xiang, Yabo Chen, Xin Zhang et al. · arXiv · Sep 10, 2026
Few-step distillation improves the efficiency of autoregressive (AR) video generation, but often causes diversity collapse: under the same prompt, different noise samples tend to produce highly similar videos with weakened motion dynamics. …
- Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking RolloutZhuoran Zhao, Shengju Qian, Tongtong Liang, Xianghao Kong et al. · arXiv · Sep 8, 2026
Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods distill pretrained bidirectional video diffusion models into causal AR students through Distribution Matching Distillation (…
- Design and Validation of a Multimodal AI Conversational System for Automated Travel Itinerary Generation and Promotional Video SynthesisPablo Vicente-Martínez, Carlos Ferrer-Baixauli, Emilio Soria-Olivas, Antonio Fernández-Baldera et al. · Applied Sciences · Sep 7, 2026
The manual creation of personalized travel itineraries remains a labor-intensive process that requires travel agents to consolidate heterogeneous information from multiple sources, including natural language interactions, booking confirmati…
- BooM-VVT: Boosting Mask-Free Video Virtual Try-On with Image-Level Pseudo DataWei Zhang, Xin Li, Peishu Shi, Jialin Gao et al. · arXiv · Sep 3, 2026
Video virtual try-on (VVT) aims to generate realistic videos of a person wearing a target garment. Recent methods leverage a keyframe-driven video generation paradigm to improve in-the-wild performance, yet they still rely on masks to local…
- DSAQuant: Denoising-Stage-Aligned Quantization-Aware Training for Video GenerationShuaiting Li, Zelin Gao, Haibin Shen, Yujun Shen et al. · arXiv · Sep 3, 2026
Video diffusion models (VDMs) have achieved impressive progress in text-to-video generation, but their high memory and computational costs hinder practical deployment. Quantization-aware training (QAT) is an effective solution for compressi…
- SolarWM: Open Data and Scalable Training for Long-Horizon Video World ModelsJunchao Huang, Guian Fang, Shengju Qian, Xianghao Kong et al. · arXiv · Sep 2, 2026
We introduce SolarWM, a fully open foundation for building interactive video world models from data preparation through long-horizon inference. Training across heterogeneous data sources and video backbones is challenging: datasets differ i…
- H3-World: Turning Language Understanding into World ControlDanze Chen, Zeqing Wang, Ziyue Lin, Xingyi Yang et al. · arXiv · Sep 1, 2026
We present H3-World, an efficient framework that turns the 33B MiniMax-H3 video generator into an interactive world model. Our key finding is that, as large video generators become more capable, language is emerging as a natural interface f…
- DreamX-Creator: Democratizing Native Audio-Video Generation at 2K ResolutionJiashu Zhu, Yanhao Zheng, Ruitian Tian, Rujing Dang et al. · arXiv · Aug 31, 2026
Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We present DreamX-Creator 1.0, a compact native joint audio-video generation system centered…
- Video Generative Models as Geometry LearnerHaosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu et al. · arXiv · Aug 28, 2026
Recent generative approaches to geometry estimation adapt pretrained image diffusion models and treat the task as image-conditioned generation. Leveraging off-the-shelf image diffusion models, they either (i) train task-specific geometry mo…
- LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video GenerationYixuan Ding, Jiahao Kong, Wei Huang, Ruijie Quan et al. · arXiv · Aug 28, 2026
Autoregressive video diffusion enables scalable long-video generation by producing chunks from a bounded recent context. While recency-based caching preserves local continuity, it evicts historical cues needed when subjects, objects, scenes…
- How Far Can 5,500 Hours of Driving Take You? A Scaling Law Analysis of Video Diffusion ModelsVictor Besnier, Anh-Quan Cao, Elias Ramzi, Spyros Gidaris et al. · ECCV 2026 · Aug 28, 2026
Video generation for autonomous driving cannot follow the web-scale route: driving data is expensive to collect, bound by privacy requirements, and cannot be scraped at will, so models must make the most of a fixed corpus. We present a syst…
- CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical SimulatorsKechen Liu, Ola Shorinwa · arXiv · Aug 27, 2026
State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physic…
- PAWBench: How Far Are We from Probabilistically Aligned World Modeling?Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes et al. · arXiv · Aug 27, 2026
Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also the distribution of p…
- GTA: Advancing Image-to-3D World Generation via Geometry Then Appearance Video DiffusionHanxin Zhu, Cong Wang, Peiyan Tu, Jiayi Luo et al. · International Journal of Co... · Aug 27, 2026
- TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency DistillationXiaoda Yang, Yuxiang Liu, Kaiwen Zheng, Yuan Liu et al. · arXiv · Aug 25, 2026
Joint text-to-video-audio generation produces synchronized visual and acoustic content, but the long sampling trajectories and heterogeneous multimodal computation of large models make inference prohibitively expensive. We present TurboT2VA…
- Deep Academic Survey: Stateful Agentic Closed-Loop Paradigm for Academic Survey AutomationZhikai Xu, Zhucun Xue, Teng Hu, Yabiao Wang et al. · arXiv · Aug 18, 2026
Academic surveys play a central role in organizing rapidly expanding scholarly literature, yet their construction requires extensive paper analysis, coherent knowledge organization, fine-grained citation support, and reliable manuscript ass…
- LinCa: Accelerating Diffusion Models via Learnable Decomposed Feature CachingJinshan Liu, Haoran Qin, Xiaobing Tu, Jiacheng Liu et al. · arXiv · Aug 18, 2026
Diffusion models have achieved remarkable success in image and video generation, yet the high computational cost of iterative sampling remains a critical bottleneck for practical deployment. Feature caching has emerged as a promising accele…
- PersonaShot: Benchmarking Person-Centric Narrative Continuity in Multi-Shot Video GenerationYuji Wang, Yuheng Chen, Teng Hu, Ran Yi et al. · arXiv · Aug 17, 2026
Video generation is rapidly evolving from single-shot clips to multi-shot narratives, where the human character serves as the core narrative anchor. However, existing benchmarks mainly assess character appearance or individual-shot quality,…
- Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social DisseminationShuo Liang, Yixing Ma, Pengfei Zhou, Xingyan Chen et al. · arXiv · Aug 14, 2026
Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation. Existing benchmarks, however, provide limited evidence on detector…
- V-RAE: Rethinking Video Latent Spaces for GenerationMinghui Guo, Shengqiong Wu, Hao Fei · arXiv · Aug 13, 2026
Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-le…
- SNM-VFI: Symmetric Nonlinear Motion-Guided Generative Video Frame InterpolationJisoo Jeong, Hong Cai, Jamie Menjay Lin, Hanno Ackermann et al. · arXiv · Aug 13, 2026
We propose Symmetric Nonlinear Motion-guided Generative Video Frame Interpolation (SNM-VFI), a training-free framework for motion-controllable generative video frame interpolation with pre-trained optical flow and video diffusion models. Un…
- StateFlow: Building, Evolving, and Accessing 3D World States for PrevisualizationYuyang Yin, Zixiang Li, Longxuan Deng, Hongkai Li et al. · arXiv · Aug 12, 2026
Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design. It lets creators iteratively refine scenes, actions, cameras, and spatial-temporal dynamics. Yet existing generative meth…
- GeoFlow: Efficient Driving Video Generation via Geometry-Aligned PriorsJiazheng Liu, Hang Li, Jiawei Zhang, Jiahe Li et al. · arXiv · Aug 12, 2026
Generative models like Diffusion Models and Flow Matching have demonstrated remarkable capabilities in synthesizing high-fidelity driving videos, but are severely constrained by high inference latency due to the requirement of extensive sam…
- VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video ForensicsBowei Liu, Zheng Lu, Yuhan Bian, Xinchen Zhang et al. · arXiv · Aug 11, 2026
Recent advances in video generation models have significantly improved the realism of synthetic videos, blurring the boundary between generated and authentic content and raising concerns about misinformation. Existing MLLM-based detectors m…
- SimWAM: A Simple World Action Model for End-to-End Autonomous DrivingZongchuang Zhao, Xin Zhou, Tianyang Xu, Zhengyang Sun et al. · arXiv · Aug 7, 2026
World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods require costly future generation at inference. We present SimWAM, a simple yet effective WAM t…
- MirrorWorld: Taming Video Diffusion Models for Mirror Reflection GenerationYoujun Zhao, Alex Warren, Gary K. L. Tam, Rynson W. H. Lau · arXiv · Aug 7, 2026
Recent advances in video diffusion models (VDMs) have enabled high-fidelity video synthesis. However, generating mirror reflections remains challenging because the content within a mirror must remain consistent with the surrounding scene. E…
- EmoWorld: A Decoupled Affective Field for Controllable Emotional Video GenerationBingyuan Wang, Baistan Zhyldyzbekov, Kunyu Feng, Zeyu Wang · arXiv · Aug 6, 2026
Emotion shapes how viewers interpret a scene, yet existing video generators entangle global atmosphere, affect-bearing semantic cues, and temporal progression within a single text condition. We present EmoWorld, a framework that decouples t…