Latest Image Generation Research Papers
The newest Image Generation papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Image Generation so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Image Generation papers in your inbox — free →Recent papers
- A Comprehensive Study on the Synthesis, Nonlinear Optical Properties, and Biological Applications of Substituted Piperazine DerivativesSahaya Infant Lasalle B, Senthil Pandian Muthu, P. Ramasamy · DOAJ (DOAJ: Directory of Op... · Dec 1, 2026
Piperazine is an important organic heterocyclic compound featuring a six-membered ring with two nitrogen atoms positioned opposite each other and four carbon atoms. This moiety is present in numerous widely recognized drugs with diverse the…
- Guided Super-Resolution of Digital Elevation Models with Diffusion-Based Image GeneratorsArmand Mihai Nicolicioiu, Dominik Narnhofer, Nando Metzger, Daniel Panangian et al. · arXiv · Sep 10, 2026
High-resolution digital surface models (DSMs) play an important role in urban analysis, 3D building reconstruction, and infrastructure monitoring, yet their availability remains limited due to the high cost and complexity of data acquisitio…
- BrainTaskonomy: Learning How to Pretrain and What to Transfer in fMRI Foundation ModelsJunfeng Xia, Wenhao Ye, Junxiang Zhang, Jiayu Zuo et al. · arXiv · Sep 9, 2026
fMRI foundation models increasingly aggregate heterogeneous data across brain states, cohorts, and acquisition settings, yet pretraining domains are commonly treated as a flat mixture and downstream tasks are adapted independently. We study…
- Scalable Detection of Fossil Palynomorphs in Multifocal Digital Microscopy ImagesAbbas Shaikh, Praise Mayor, Patrick Ainlay-Vazquez, Aditya Viswanathan et al. · arXiv · Sep 4, 2026
Palynomorphs (microscopic, organic-walled fossils such as pollen, spores, and dinoflagellates) are important high-resolution records of past climates and are critical to the study of ancient ecosystems. Existing methods rely on manual analy…
- Thinking in Pictures: A Systematic Benchmark for Reasoning-driven Image GenerationYutong Liu, Nan Huang, Xu Cao, James M. Rehg · arXiv · Sep 2, 2026
Recent advancements in unified generative models (UGMs) and world simulators have achieved unprecedented results in visual perception and synthesis. However, these models primarily rely on surface-level event alignment, leaving the capacity…
- PlantC2USeg: Cross-Scale Consistent Pre-Training for Few-Shot Unified Plant Point Cloud SegmentationYu Tian, Xintong Jiang, Jan Franklin Adamowski, Shiv O. Prasher et al. · arXiv · Sep 2, 2026
Modern crop breeding demands precise organ-level analysis for trait quantification, making plant point cloud segmentation (PPCS) increasingly important. However, conventional deep learning approaches rely heavily on densely annotated datase…
- GDB-Reward: From Evaluation Metrics to Training Rewards for Graphic DesignAdrienne Deganutti, Purvanshi Mehta, Simon Hadfield, Andrew Gilbert · arXiv · Sep 2, 2026
Text-to-image models excel at natural image synthesis but struggle with graphic design, where success depends on satisfying precise constraints on typography, layout, color, and visual communication. While prompt optimization offers an attr…
- Genesis: A Generative Engine for Hierarchical Satellite Image SynthesisSubash Khanal, Yangzhi Cui, Daniel Cher, Eric Xing et al. · arXiv · Sep 2, 2026
Earth observation is fundamentally multi-scale; geospatial tasks span varied resolutions, and satellite imagery is organized into cascading tile pyramids that nest fine detail within wide coverage. Current generative models of satellite ima…
- SpatialGuard: Harness-Guided Verifiable Spatial Reasoning for Text-to-Image GenerationZiyun Qian, Zizhi Chen, Yizhou Liu, Mingyang Sun et al. · arXiv · Sep 1, 2026
Complex 3D spatial text to image generation requires models to convert natural language into stable visual geometry, not merely semantic appearance. Existing prompt-driven or layout-conditioned methods improve controllability, but often lac…
- What, Where, and How: Probing Spatiotemporal Representations in Video Foundation ModelsSharon S. Musa, Fereshteh Forghani, Harrish Thasarathan, Sonia Joseph et al. · arXiv · Sep 1, 2026
Self-supervised video foundation models learn rich spatiotemporal representations, yet it remains unclear what visual concepts these representations encode, where they emerge across transformer layers, and how they are geometrically organiz…
- BRF-GS: Hyperspectral Bidirectional Reflectance Factor Modeling and Image Generation Based on 3D Gaussian SplattingYiling Yao, Wenjuan Zhang, Bowen Wang, Bocheng Li et al. · arXiv · Aug 31, 2026
The bidirectional reflectance factor (BRF) characterizes the directional radiative properties of terrestrial surfaces. However, existing three-dimensional (3D) radiative transfer models require complex scene construction and computationally…
- LISynSeg: Data-Centric Label-to-Image Synthesis for Cross-Modality Whole-Heart SegmentationJiacheng Wang, Ivana Isgum, Ipek Oguz · arXiv · Aug 31, 2026
Whole-heart segmentation (WHS) in computed tomography (CT) and magnetic resonance imaging (MRI) is affected by acquisition shifts and heterogeneous cardiac annotations. Existing WHS systems combine architectural design, transfer learning, a…
- Identity-Conditioned Latent Consistency Distillation for Face SynthesisTiago Kienen Chaves, Bernardo Biesseck, David Menotti · arXiv · Aug 31, 2026
Diffusion models have achieved strong results in high-fidelity image synthesis, but their iterative sampling process makes large-scale generation computationally expensive. This limitation is especially relevant when generating synthetic fa…
- Knowledge Distillation Driven Semantic NOMA with GAN Refinement for 6G Robotic Vehicle NetworksQifei Wang, Zhen Gao, Li Qiao, Ziwei Wan et al. · arXiv · Aug 27, 2026
To achieve sustainable intelligent mobility, 6G-empowered robotic vehicles (RVs) require high-fidelity visual perception under stringent bandwidth and energy constraints. Semantic communication offers a spectral-efficient solution but suffe…
- TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency DistillationXiaoda Yang, Yuxiang Liu, Kaiwen Zheng, Yuan Liu et al. · arXiv · Aug 25, 2026
Joint text-to-video-audio generation produces synchronized visual and acoustic content, but the long sampling trajectories and heterogeneous multimodal computation of large models make inference prohibitively expensive. We present TurboT2VA…
- WithEveryone: Unified Planning and Identity Grounding for Group Image GenerationHengyuan Xu, Qixun Wang, Yiji Cheng, Miles Yang et al. · arXiv · Aug 20, 2026
Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time…
- Swift-Image: Exploring the Performance Frontier of Compact Unified Image Generation ModelsTaihang Hu, Zhao Wang, Zuan Gao, Tao Liu et al. · arXiv · Aug 20, 2026
We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi-image editing. Our goal is to explore how far a relatively small visual generator can be pushed through systematic training engine…
- From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image GenerationXingjian Wang, Zhao Wang, Taihang Hu, Jun Zheng et al. · arXiv · Aug 18, 2026
Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate…
- Deep Academic Survey: Stateful Agentic Closed-Loop Paradigm for Academic Survey AutomationZhikai Xu, Zhucun Xue, Teng Hu, Yabiao Wang et al. · arXiv · Aug 18, 2026
Academic surveys play a central role in organizing rapidly expanding scholarly literature, yet their construction requires extensive paper analysis, coherent knowledge organization, fine-grained citation support, and reliable manuscript ass…
- TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image GenerationHaoran Wang, Chaofan Ma, Ran Yi, Lizhuang Ma · arXiv · Aug 17, 2026
Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., "subject composition"), which are ill-suited to this combinatorial setting a…
- GenRouter: Unified Workflow Routing for Agentic Image GenerationHarold Haodong Chen, Zhiyu Hou, Wen-Jie Shu, Weilin Ruan et al. · arXiv · Aug 17, 2026
The rapid evolution of text-to-image (T2I) generation models has effectively solved the foundational challenge of raw pixel synthesis, shifting the community's focus toward fulfilling increasingly intricate user requests. While recent agent…
- A diffusion-based multi-style image generation framework with adaptive weighted style fusionHaoran Gong · Discover Artificial Intelli... · Aug 14, 2026
- V-RAE: Rethinking Video Latent Spaces for GenerationMinghui Guo, Shengqiong Wu, Hao Fei · arXiv · Aug 13, 2026
Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoder architectures have evolved substantially, their latent spaces are still optimized primarily for pixel-le…
- TabSOM: A tabular-to-image encoding method based on self-organizing mapsDavid Chushig-Muzo, María Ángeles Rodríguez de Cara, Eva Milara, Francisco J. Lara-Abelenda et al. · arXiv · Aug 13, 2026
Tabular-to-image methods have emerged as novel approaches to leverage the high predictive performance of convolutional neural networks and vision transformers. They convert tabular data into image representations, mapping each feature at a …
- Evaluation of Clinically Steerable Retinal Image Generation from Foundation Model Latent SpacesZuzanna A. Wakefield-Skórniewska, Bartłomiej W. Papież · arXiv · Aug 13, 2026
Medical foundation models learn latent representations of clinically meaningful phenotypes, yet their ability to support controllable image generation remains largely unexplored. We evaluate four retinal foundation models within the represe…
- XYZFlow:Scaling Multi dimensional Shortcut Flows for Efficient Generative ModelingJinxiu Liu, Xuanming Liu, Kangfu Mei, Yandong Wen et al. · arXiv · Aug 12, 2026
High-fidelity image generation faces a trade-off between speed and quality. Diffusion models produce strong visuals but require costly iterative sampling. Existing efficient methods mainly distill pretrained models into few-step samplers, a…
- HSTGFormer: Hyper Spatial-Temporal Graph Transformer for 3D Human Pose EstimationRuochen Li, Shuang Chen, Wenke E, Farshad Arvin et al. · arXiv · Aug 12, 2026
Transformer-based methods have achieved strong performance in monocular 3D human pose estimation, but most existing approaches organise spatial and temporal reasoning as separate stages, which may weaken unified spatial-temporal interdepend…
- Understanding ChatGPT Adoption for AI Image Generation Among Visual Communication Design StudentsFreya Enggrayni, Anita Wulansari, Tri Puspa Rinjeni · bit-Tech · Aug 10, 2026
The use of Artificial Intelligence Generated Content (AIGC) for image creation is growing rapidly among Visual Communication Design (VCD) students, yet concerns over copyright, creative originality, and professional relevance persist. This …
- Chiral Bifacial Indacenodithiophene‐Based Hole‐Transport Materials With Chirality‐Induced Spin Selectivity: Chirality‐Spin Polarity Correspondence and Perovskite PassivationShuang Li, Fumitaka Ishiwari, Ryosuke Nishikubo, Akinori Saeki · Small · Aug 8, 2026
Chirality-induced spin selectivity (CISS) is emerging as a key element for spin-dependent functions in organic electronics. We previously developed a chiral bifacial indacenodithiophene (IDT) backbone that shows strong CISS in both π-conjug…
- PRISM: Distribution-Gated Flow Matching for Controllable Unpaired Image TranslationElad Yoshai, Natan T. Shaked · arXiv · Aug 6, 2026
Unpaired image-to-image translation must decide, per image, what to change and what to preserve without paired supervision. Many diffusion-based unpaired translators control preservation through a single global noise or guidance value appli…