Latest Text-to-Image Research Papers
The newest Text-to-Image papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Text-to-Image so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Text-to-Image papers in your inbox — free →Recent papers
- FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow EstimationVladislav Bargatin, Alexander Yakovenko, Khaled Abud, Dmitriy Vatolin · arXiv · Sep 10, 2026
Optical flow methods typically rely on task-specific inductive biases, such as correlation volumes, feature warping, and iterative refinement, among others, to reach high accuracy. While effective, such biases constrain the model to predefi…
- GDB-Reward: From Evaluation Metrics to Training Rewards for Graphic DesignAdrienne Deganutti, Purvanshi Mehta, Simon Hadfield, Andrew Gilbert · arXiv · Sep 2, 2026
Text-to-image models excel at natural image synthesis but struggle with graphic design, where success depends on satisfying precise constraints on typography, layout, color, and visual communication. While prompt optimization offers an attr…
- Genesis: A Generative Engine for Hierarchical Satellite Image SynthesisSubash Khanal, Yangzhi Cui, Daniel Cher, Eric Xing et al. · arXiv · Sep 2, 2026
Earth observation is fundamentally multi-scale; geospatial tasks span varied resolutions, and satellite imagery is organized into cascading tile pyramids that nest fine detail within wide coverage. Current generative models of satellite ima…
- SpatialGuard: Harness-Guided Verifiable Spatial Reasoning for Text-to-Image GenerationZiyun Qian, Zizhi Chen, Yizhou Liu, Mingyang Sun et al. · arXiv · Sep 1, 2026
Complex 3D spatial text to image generation requires models to convert natural language into stable visual geometry, not merely semantic appearance. Existing prompt-driven or layout-conditioned methods improve controllability, but often lac…
- Gaussian Core LoRA: Distribution-Aware Dynamic Adaptation for Broad Concept ErasureQinghui Gong, Xunlei Chen, Yu-Xuan Zhang, Hua Meng et al. · arXiv · Sep 1, 2026
Concept erasure aims to suppress unsafe, privacy-sensitive, or undesirable generations in text-to-image diffusion models while preserving benign semantics, visual quality, and deployment efficiency. Existing adapter-based methods, such as L…
- LISynSeg: Data-Centric Label-to-Image Synthesis for Cross-Modality Whole-Heart SegmentationJiacheng Wang, Ivana Isgum, Ipek Oguz · arXiv · Aug 31, 2026
Whole-heart segmentation (WHS) in computed tomography (CT) and magnetic resonance imaging (MRI) is affected by acquisition shifts and heterogeneous cardiac annotations. Existing WHS systems combine architectural design, transfer learning, a…
- Identity-Conditioned Latent Consistency Distillation for Face SynthesisTiago Kienen Chaves, Bernardo Biesseck, David Menotti · arXiv · Aug 31, 2026
Diffusion models have achieved strong results in high-fidelity image synthesis, but their iterative sampling process makes large-scale generation computationally expensive. This limitation is especially relevant when generating synthetic fa…
- Vision-centric generative AI models: A software-hardware perspectiveEleni Tselepi, Cristian Sestito, Shady Agwa, Themis Prodromakis · arXiv · Aug 27, 2026
Vision generative artificial intelligence (AI) has emerged as one of the most rapidly advancing areas of deep learning. The explosion of multimodal models has made them widely associated with text-to-image applications running on large data…
- TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency DistillationXiaoda Yang, Yuxiang Liu, Kaiwen Zheng, Yuan Liu et al. · arXiv · Aug 25, 2026
Joint text-to-video-audio generation produces synchronized visual and acoustic content, but the long sampling trajectories and heterogeneous multimodal computation of large models make inference prohibitively expensive. We present TurboT2VA…
- Swift-Image: Exploring the Performance Frontier of Compact Unified Image Generation ModelsTaihang Hu, Zhao Wang, Zuan Gao, Tao Liu et al. · arXiv · Aug 20, 2026
We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi-image editing. Our goal is to explore how far a relatively small visual generator can be pushed through systematic training engine…
- From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image GenerationXingjian Wang, Zhao Wang, Taihang Hu, Jun Zheng et al. · arXiv · Aug 18, 2026
Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate…
- Deep Academic Survey: Stateful Agentic Closed-Loop Paradigm for Academic Survey AutomationZhikai Xu, Zhucun Xue, Teng Hu, Yabiao Wang et al. · arXiv · Aug 18, 2026
Academic surveys play a central role in organizing rapidly expanding scholarly literature, yet their construction requires extensive paper analysis, coherent knowledge organization, fine-grained citation support, and reliable manuscript ass…
- An Empirical Study of Training Pixel-Space Text-to-Image Diffusion ModelsDengyang Jiang, Ruoyi Du, Zhennan Chen, Dongyang Liu et al. · arXiv · Aug 17, 2026
This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Although numerous studies have explored this topic, most focus on small-scale or class-conditional settings. Consequently, a pract…
- Unlocking the Potential of Image Editing via Concept Scaling and Dense SupervisionLong Cui, Xiaoqian Liu, Qi Qin, Yi Xin et al. · arXiv · Aug 17, 2026
Existing image editing frameworks predominantly follow the training paradigm of text-to-image diffusion models. However, extending this paradigm to image editing highlights two inherent discrepancies, specifically, the insufficient attentio…
- PixRestore: Unified Image Restoration via Pixel Diffusion TransformerLingchen Sun, Rongyuan Wu, Xiangtao Kong, Jixin Zhao et al. · arXiv · Aug 17, 2026
Unified image restoration (UIR) aims to recover high-quality (HQ) content from low-quality (LQ) images with different degradations using a single model. Most recent methods adapt large pretrained text-to-image (T2I) latent diffusion models …
- GenRouter: Unified Workflow Routing for Agentic Image GenerationHarold Haodong Chen, Zhiyu Hou, Wen-Jie Shu, Weilin Ruan et al. · arXiv · Aug 17, 2026
The rapid evolution of text-to-image (T2I) generation models has effectively solved the foundational challenge of raw pixel synthesis, shifting the community's focus toward fulfilling increasingly intricate user requests. While recent agent…
- CAPEval: A Decoupled Caption Evaluation across Understanding and GenerationZhipeng Liu, Haochen Wang, Zhaoxiang Zhang · arXiv · Aug 3, 2026
Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the caption quality as a single scalar objective, which conflates two distinct properties: (1…
- TIGA: Trajectory-Injected Generative Attack against Black-box AIGC DetectorsXia Du, Zhuosen Bao, Zheng Lin, Jizhe Zhou et al. · arXiv · Jul 28, 2026
Recent diffusion models have achieved remarkable realism in facial image synthesis, posing growing challenges to artificial intelligence-generated content (AIGC) forensic detectors.Existing evasion methods typically perturb pre-generated im…
- MicroZoom: Structure-Preserving Detail Synthesis at Extreme ScaleHuy Huynh, Jingwei Ma, Brian Curless, Ira Kemelmacher-Shlizerman et al. · arXiv · Jul 27, 2026
We introduce MicroZoom, a generative framework for gigapixel image synthesis at the microscopic scale. Given a standard photograph and a sparse set of consumer-grade microscope close-ups, MicroZoom synthesizes a seamless, gigapixel-resoluti…
- ERUnderstand: Evaluating Vision-Language Models on Structured ER DiagramsAli Ansari, Yasmin Mohammadi, Farnoush Nili, Parsa Esmaeilkhani et al. · arXiv · Jul 27, 2026
Entity-Relationship Diagrams (ERDs) are central to conceptual database design, yet they are typically available only as rendered images rather than machine-readable schemas, limiting AI-assisted database engineering. We introduce ERUndersta…
- Appearance Pointers -- Multimodal Region Control of Diffusion TransformersRahul Sajnani, Yulia Gryaditskaya, Radomír Měch, Srinath Sridhar et al. · arXiv · Jul 21, 2026
Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alo…
- ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual SynthesisYuan Wang, Yongchao Du, Mengting Chen, Jinsong Lan et al. · arXiv · Jul 21, 2026
Recent advances in multimodal generative models have enabled instruction-based image generation to move beyond semantic manipulation to knowledge-driven visual reasoning. However, these methods focus on explicit commonsense reasoning, shall…
- Text Template Tokens Are Implicit Semantic Registers in Diffusion TransformersMaohua Li, Qirui Li, Yanke Zhou, Yiduo Li et al. · arXiv · Jul 21, 2026
Text-to-image diffusion transformers (DiTs) jointly process text and image tokens, yet their internal computation during denoising remains poorly understood. We introduce a causal interpretability framework for modern large-scale DiTs that …
- GenEval 2: Addressing Benchmark Drift in Text-to-Image EvaluationAmita Kamath, Kai-Wei Chang, Ranjay Krishna, Luke S. Zettlemoyer et al. · arXiv.org · Dec 18, 2025
Automating Text-to-Image (T2I) model evaluation is challenging; a judge model must be used to score correctness, and test prompts must be selected to be challenging for current T2I models but not the judge. We argue that satisfying these co…
- FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive BenchmarkRongyao Fang, Aldrich Yu, Chengqi Duan, Linjiang Huang et al. · arXiv.org · Sep 11, 2025
The advancement of open-source text-to-image (T2I) models has been hindered by the absence of large-scale, reasoning-focused datasets and comprehensive evaluation benchmarks, resulting in a performance gap compared to leading closed-source …
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?Ouxiang Li, Yuan Wang, Xinting Hu, Huijuan Huang et al. · arXiv.org · Sep 3, 2025
Text-to-image (T2I) generation aims to synthesize images from textual prompts, which jointly specify what must be shown and imply what can be inferred, which thus correspond to two core capabilities: \textbf{\textit{composition}} and \textb…
- Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement LearningYibin Wang, Zhimin Li, Yuhang Zang, Yujie Zhou et al. · arXiv.org · Aug 28, 2025
Recent advancements highlight the importance of GRPO-based reinforcement learning methods and benchmarking in enhancing text-to-image (T2I) generation. However, current methods using pointwise reward models (RM) for scoring generated images…
- T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image GenerationKaiyue Sun, Rongyao Fang, Chengqi Duan, Xian Liu et al. · arXiv.org · Aug 24, 2025
We propose T2I-ReasonBench, a benchmark evaluating reasoning capabilities of text-to-image (T2I) models. It consists of four dimensions: Idiom Interpretation, Textual Image Design, Entity-Reasoning and Scientific-Reasoning. We propose a two…
- T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive GenerationChieh-yun Chen, Min Shi, Gong Zhang, Humphrey Shi · IEEE International Conference on Computer Vision · Jul 28, 2025
Text-to-Image (T2I) generative models have revolutionized content creation but remain highly sensitive to prompt phrasing, often requiring users to repeatedly refine prompts multiple times without clear feedback. While techniques such as au…
- The Illusion of Unlearning: The Unstable Nature of Machine Unlearning in Text-to-Image Diffusion ModelsNaveen George, Karthik Nandan Dasaraju, Rutheesh Reddy Chittepu, Konda Reddy Mopuri · Computer Vision and Pattern Recognition · Jun 10, 2025
Text-to-image models such as Stable Diffusion, DALL•E, and Midjourney have gained immense popularity lately. However, they are trained on vast amounts of data that may include private, explicit, or copyrighted material used without permissi…