Latest Diffusion Models Research Papers
The newest Diffusion Models papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Diffusion Models so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Diffusion Models papers in your inbox — free →Recent papers
- Guided Super-Resolution of Digital Elevation Models with Diffusion-Based Image GeneratorsArmand Mihai Nicolicioiu, Dominik Narnhofer, Nando Metzger, Daniel Panangian et al. · arXiv · Sep 10, 2026
High-resolution digital surface models (DSMs) play an important role in urban analysis, 3D building reconstruction, and infrastructure monitoring, yet their availability remains limited due to the high cost and complexity of data acquisitio…
- Harnessing Intrinsic Subject-Aware Attention for Controllable Multi-Subject Video GenerationNiange Yu, Ye Tian, Biaolong Chen, Miao Lu et al. · arXiv · Sep 10, 2026
Multi-subject video generation faces two key challenges: uncontrollable fidelity strength and potential semantic drift. We address these by analyzing the internal mechanisms of Diffusion Transformers (DiTs). We found that certain attention …
- FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow EstimationVladislav Bargatin, Alexander Yakovenko, Khaled Abud, Dmitriy Vatolin · arXiv · Sep 10, 2026
Optical flow methods typically rely on task-specific inductive biases, such as correlation volumes, feature warping, and iterative refinement, among others, to reach high accuracy. While effective, such biases constrain the model to predefi…
- Advanced Brain Tissue Imaging with Data-Consistent Diffusion Priors in Laminographic X-Ray NanoimagingWenxuan Fang, Abraham L. Levitan, Ana Diaz, Carles Bosch et al. · arXiv · Sep 9, 2026
Nanoscale imaging of mammalian brains is critical for connectomics. X-ray laminography enables high-throughput imaging of extended, plate-like biological specimens. However, the tilted acquisition geometry leads to incomplete Fourier-space …
- SceneHI: High-Resolution 3D-Consistent Scene Texturing with Controllable IlluminationAthanasios Tragakis, Marco Aversa, Daniela Ivanova, Chaitanya Kaul et al. · arXiv · Sep 9, 2026
SceneHI is a framework that lifts high-resolution, illumination-aware priors from 2D diffusion models to perform 3D texture synthesis. It is the first to demonstrate that high-resolution textures, previously limited to 2D synthesis, can be …
- Turing Pattern Formation in the Chlorine Dioxide-Iodine-Malonic Acid Reaction-Diffusion SystemSima Setayeshgar · OpenAlex · Sep 9, 2026
The formation of localized structures in the chlorine dioxide-idodine-malonic acid (CDIMA) reaction-diffusion system is investigated numerically using a realistic model of this system. We analyze the one-dimensional patterns formed along th…
- Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking RolloutZhuoran Zhao, Shengju Qian, Tongtong Liang, Xianghao Kong et al. · arXiv · Sep 8, 2026
Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods distill pretrained bidirectional video diffusion models into causal AR students through Distribution Matching Distillation (…
- Positivity and long-term behaviour of a diffusion model with measure-valued nonlocal reaction term : applications in bioscience and engineeringXiao Yang, Qiyao Peng, SC Hille · Lancaster EPrints (Lancaste... · Sep 5, 2026
The behaviour is investigated of solutions to a diffusion equation on the real line with nonlocal and singular reaction term, i.e., given by a Dirac source or sink at the origin. It gives a simplified representation of for example a control…
- Reflection-aware Generative Novel View SynthesisGeonU Kim, Shin Dong-Yeon, Tae-Hyun Oh · arXiv · Sep 4, 2026
We propose Ref-GeNVS, a training-free, reflection-aware method for generative novel view synthesis (NVS) in mirror scenes. Existing multi-view diffusion models often fail to recognize the mirror in the scene and cannot exploit reflected con…
- Measured Sliders: Learning Continuous Controls from Differentiable Image MeasurementsYijia Chen, Boyu Wei, Xuanhua Yin · arXiv · Sep 4, 2026
Continuous sliders are useful only when coefficient changes produce predictable image changes. Yet most diffusion sliders derive their axes from text or learned representations, leaving their scales disconnected from observable image proper…
- Editable Visual DesignJunyan Ye, Wei Liu, Dongzhi Jiang, Zichen Wen et al. · arXiv · Sep 3, 2026
While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-to-end generation inherently yields flattened bitmaps with error-prone text, precluding layer-wise post-editing. Conversely,…
- DSAQuant: Denoising-Stage-Aligned Quantization-Aware Training for Video GenerationShuaiting Li, Zelin Gao, Haibin Shen, Yujun Shen et al. · arXiv · Sep 3, 2026
Video diffusion models (VDMs) have achieved impressive progress in text-to-video generation, but their high memory and computational costs hinder practical deployment. Quantization-aware training (QAT) is an effective solution for compressi…
- Artificial Intelligence and Future-Oriented Aesthetic Education: Design Thinking in Middle School Art CurriculumYijun Feng · Intelligent & Human Futures · Sep 3, 2026
Middle school art curricula in many systems still center on skill reproduction, which sits uneasily with the rapid diffusion of generative AI. This study develops a conceptual instructional model that embeds generative AI within the five-st…
- GDB-Reward: From Evaluation Metrics to Training Rewards for Graphic DesignAdrienne Deganutti, Purvanshi Mehta, Simon Hadfield, Andrew Gilbert · arXiv · Sep 2, 2026
Text-to-image models excel at natural image synthesis but struggle with graphic design, where success depends on satisfying precise constraints on typography, layout, color, and visual communication. While prompt optimization offers an attr…
- DualDiff3D: Dual Structure-Appearance Diffusion Priors for Reliability-Enhanced 3D Gaussian SplattingQian Wang, Yu Wang, Weiqi Li, Xinhua Cheng et al. · arXiv · Sep 1, 2026
While 3D Gaussian Splatting (3DGS) has revolutionized 3D reconstruction and novel-view synthesis, scenarios with limited input views often lead to poor reconstruction quality and artifacts in rendered novel views. Recent efforts attempt to …
- Gaussian Core LoRA: Distribution-Aware Dynamic Adaptation for Broad Concept ErasureQinghui Gong, Xunlei Chen, Yu-Xuan Zhang, Hua Meng et al. · arXiv · Sep 1, 2026
Concept erasure aims to suppress unsafe, privacy-sensitive, or undesirable generations in text-to-image diffusion models while preserving benign semantics, visual quality, and deployment efficiency. Existing adapter-based methods, such as L…
- Exploring the Link Between Intravoxel Incoherent Motion Measured Brain Diffusivity During Wakefulness and Sleep Macrostructure in the Elderly.Lena Meinhold, Misha P T Kaandorp, Ziyao Shang, Medea Häuselmann et al. · Open Access CRIS of the Uni... · Sep 1, 2026
Intravoxel incoherent motion (IVIM) is a diffusion-weighted magnetic resonance imaging (MRI) method that models slow (D, tissue diffusivity) and fast (D*, microvascular perfusion) signal components (f, signal fraction). This observational s…
- Extreme scenario generation for drought-flood abrupt alternation based on improved conditional diffusion modelPengxin CHEN, Fuming Yao, Weibin HUANG, Yanmei ZHU et al. · DOAJ (DOAJ: Directory of Op... · Sep 1, 2026
Under the backdrop of global warming,abrupt transitions between drought and flood events have been rather frequent,posing severe challenges to flood-control,drought-resistance,and water-resource management in river basins. The scarce availa…
- First-principles study of helium migration in stishovite in the Earth's mantleYu Huang, Hang Ren · DOAJ (DOAJ: Directory of Op... · Sep 1, 2026
The migration mechanisms of helium (He) in anhydrous and hydrous stishovite under Earth's mantle conditions are studied using density functional theory (DFT) and climbing image nudged elastic band (CI-NEB) transition state calculations. In …
- DreamX-Creator: Democratizing Native Audio-Video Generation at 2K ResolutionJiashu Zhu, Yanhao Zheng, Ruitian Tian, Rujing Dang et al. · arXiv · Aug 31, 2026
Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We present DreamX-Creator 1.0, a compact native joint audio-video generation system centered…
- Identity-Conditioned Latent Consistency Distillation for Face SynthesisTiago Kienen Chaves, Bernardo Biesseck, David Menotti · arXiv · Aug 31, 2026
Diffusion models have achieved strong results in high-fidelity image synthesis, but their iterative sampling process makes large-scale generation computationally expensive. This limitation is especially relevant when generating synthetic fa…
- Video Generative Models as Geometry LearnerHaosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu et al. · arXiv · Aug 28, 2026
Recent generative approaches to geometry estimation adapt pretrained image diffusion models and treat the task as image-conditioned generation. Leveraging off-the-shelf image diffusion models, they either (i) train task-specific geometry mo…
- Learning the Target Priors Before Image Translation: A Decoupled Training Paradigm for Cross-Modal Image Translation in Remote SensingKeyan Hu, Mingtao Wang, Ziyu Zhou, Tiandong Shi et al. · arXiv · Aug 28, 2026
Cross-modal image translation in remote sensing must preserve source-observed content while matching the target-domain distribution. Existing methods jointly learn the target prior and cross-modal dependence from scarce paired data, overloo…
- LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video GenerationYixuan Ding, Jiahao Kong, Wei Huang, Ruijie Quan et al. · arXiv · Aug 28, 2026
Autoregressive video diffusion enables scalable long-video generation by producing chunks from a bounded recent context. While recency-based caching preserves local continuity, it evicts historical cues needed when subjects, objects, scenes…
- How Far Can 5,500 Hours of Driving Take You? A Scaling Law Analysis of Video Diffusion ModelsVictor Besnier, Anh-Quan Cao, Elias Ramzi, Spyros Gidaris et al. · ECCV 2026 · Aug 28, 2026
Video generation for autonomous driving cannot follow the web-scale route: driving data is expensive to collect, bound by privacy requirements, and cannot be scraped at will, so models must make the most of a fixed corpus. We present a syst…
- Denoising-Aware Temporal Point Cloud Completion for 3D Crop Architecture Recovery and Phenotypic Trait ExtractionMrudul Mittal, Soumyashree Kar · arXiv · Aug 28, 2026
High-throughput phenotyping depends on accurate 3D reconstruction of plants across growth stages, yet the development and evaluation of temporal completion methods are limited by the lack of datasets with complete geometric ground truth. To…
- FUSED: Forensic-Semantic Mixture-of-Experts for AI Inpainting Detection and LocalizationAnton Nuzhdin, Marcel Worring, Ivona Najdenkoska · arXiv · Aug 28, 2026
Diffusion-based inpainting models modify only a localized part of an image, while many AI-image detectors rely on global artifacts and do not localize. These artifacts vary across generators, limiting detector transfer under distribution sh…
- Physics-Guided Flow Matching for CT Image ReconstructionDavide Evangelista · arXiv · Aug 28, 2026
Deep generative models have recently emerged as powerful priors for solving ill-posed inverse problems in CT, with diffusion-based approaches achieving state-of-the-art reconstruction performance. However, diffusion models typically rely on…
- Score-Based Ideal Observer Approximation via Denoising Score Matching for Signal-Known-Exactly Detection TasksWeimin Zhou · arXiv · Aug 25, 2026
The Bayesian Ideal Observer (IO) establishes the theoretical upper bound on task performance for binary detection tasks. However, analytical computation of the IO test statistic is generally intractable. Numerical approaches based on Markov…
- TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency DistillationXiaoda Yang, Yuxiang Liu, Kaiwen Zheng, Yuan Liu et al. · arXiv · Aug 25, 2026
Joint text-to-video-audio generation produces synchronized visual and acoustic content, but the long sampling trajectories and heterogeneous multimodal computation of large models make inference prohibitively expensive. We present TurboT2VA…