Latest Synthetic Data Research Papers
The newest Synthetic Data papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Synthetic Data so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Synthetic Data papers in your inbox — free →Recent papers
- Physics-Informed Deep Learning for False Ventricular Tachycardia Alarm Reduction in the ICUAthanasios Papastathopoulos-Katsaros, Alexandra Stavrianidi, Zhandong Liu · arXiv · Sep 8, 2026
False ventricular tachycardia (VT) alarms are a leading contributor to alarm fatigue in intensive care units. We propose a deep learning framework combining a 1D SE-ResNet with ICU-realistic data augmentations and a physics-informed auxilia…
- LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style UpdatesDmitrii Andriianov, Andrey Veprikov, Aleksandr Beznosikov · arXiv · Sep 2, 2026
Low-rank adaptation (LoRA) is the standard way to fine-tune large models, yet when its two factors are trained independently, the update ignores the geometry of the low-rank weight change it induces. We introduce LoRA-TSD, an optimizer that…
- Learning a Size-Weight Frontier for Synthetic-Augmented InferenceChengpiao Huang, Kaizheng Wang · arXiv · Aug 28, 2026
Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic samples as real data can introduce bias and lead to unreliable inference. We develop a general framework for synthetic-augmented infe…
- CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge BasesSil Hamilton, Albert Yu Sun, Oscar J. Romero, Carl-Leander Henneking et al. · arXiv · Aug 27, 2026
LLMs are increasingly able to answer complex questions about enterprise-scale document collections. But evaluation is hard: companies don't want to share internal communications, and synthetic datasets have been overly simple. We present Co…
- Across-Design Uncertainty in Short Pricing Panels: Evidence from Simulated Price TrajectoriesPedro Cadahia Delgado · arXiv · Aug 21, 2026
Short observational pricing panels can contain many observations while offering only a small number of distinct price movements. This paper studies the inferential consequences of that distinction in a synthetic data-generating process cali…
- On the Fragility of Self-Improving Agents: Variance, Task Order, and UnderspecificationQinyuan Ye, Yu Li, Yada Pruksachatkun, Jiaxin Zhang et al. · arXiv · Aug 18, 2026
Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank--have shown great promise in recent literature. However, the reliability aspects of these methods…
- Boosting Data Augmentation with Stochastic Weight AveragingLongde Huang, Axel Flinth, Jan E. Gerken · arXiv · Aug 14, 2026
The symmetries of a learning task have become an important factor in designing modern deep learning solutions. Data augmentation is a straightforward and effective way of incorporating symmetries into a generic neural network. Recent result…
- OTLesMix: Wasserstein Barycenter and Optimal Transport Map for Synthetic Lesion Generation with Diverse Shapes and LocationsRobin Trombetta, Carole Lartizien · arXiv · Aug 6, 2026
The development of deep learning over the past decade has revolutionized medical imaging segmentation, allowing the extraction of precise descriptors from large volumes to characterize pathologies. Data augmentation is a technique widely re…
- Assessment of Conditional Diffusion Model for Synthetic Histopathology Image GenerationSeyed Kahaki, Shijie Li, Weijie Chen, Nicholas Petrick · arXiv · Aug 4, 2026
Synthetic histopathology image generation has emerged as an approach that may address data scarcity in computational pathology, yet current evaluation methodologies may not fully assess synthetic data quality for medical applications. This …
- Synthetic data generation framework for quality control automation in gravure printingKorota Arsène Coulibaly, Mohamed Hamlich, Khalid Hmali, Andrea Trombin · arXiv · Jul 23, 2026
Quality control in printing, particularly in rotogravure printing, still depends on slow, costly, and subjective manual inspection. Automated surface defect detection is critical for maintaining high-quality standards in rotogravure printin…
- ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language ModelsKrishna Teja Chitty-Venkata, Murali Emani · ICLR 2026 Workshop DATA-FM · Mar 2, 2026
We develop ImageNet-Think-250K, a multimodal reasoning dataset designed to aid the development of Vision Language Models (VLMs) with explicit reasoning capabilities. Our dataset is built on 250,000 images from ImageNet-21k dataset, providin…
- ESDAE: Evaluating Synthetic Data for Agent EvaluationShuaiqi Wang, Aadyaa Maddi, Zinan Lin, Giulia Fanti · ICLR 2026 Workshop DATA-FM · Mar 2, 2026
Agent evaluation is often performed on static datasets of execution trajectories, but real traces may be sensitive, proprietary, or too small to support comprehensive testing. Practitioners may therefore replace or augment real datasets wit…
- SynQuE: Estimating Synthetic Dataset Quality Without AnnotationsArthur Chen, Victor Zhong · ICLR 2026 Workshop DATA-FM · Mar 2, 2026
We introduce and formalize the Synthetic Dataset Quality Estimation (SYNQUE) problem: ranking synthetic datasets by their expected real-world task performance using only limited unannotated real data. This addresses a critical and open chal…
- Less is More: Adaptive Coverage Sampling for Synthetic Training DataSasan Tavakkol, Max Springer, Mohammadhossein Bateni, Vincent Cohen-Addad et al. · ICLR 2026 Workshop DATA-FM · Mar 2, 2026
Large Language Models (LLMs) enable rapid generation of synthetic training data for downstream classifiers, offering a solution when human-labeled data is costly, scarce, or time-sensitive. However, synthetic datasets suffer from systematic…
- EPSVec: Efficient and Private Synthetic Data Generation via Dataset VectorsAmin Banayeeanzade, Qingchuan Yang, Deqing Fu, Spencer Hong et al. · ICLR 2026 Workshop DATA-FM · Mar 2, 2026
High-quality data is essential for modern machine learning, yet many valuable corpora are sensitive and cannot be freely shared. Synthetic data offers a practical substitute for downstream development, and large language models (LLMs) have …
- Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term ConvergenceBingji Yi, Qiyuan Liu, Yuwei Cheng, Haifeng Xu · ICLR 2026 Workshop DATA-FM · Mar 2, 2026
Synthetic data has been increasingly used to train frontier generative models. However, recent study raises key concerns that iteratively retraining a generative model on its self-generated synthetic data may keep deteriorating model perfor…
- Learning from Synthetic Data Improves Multi-hop ReasoningAnmol Kabra, Yilun Yin, Albert Gong, Kamilė Stankevičiūtė et al. · ICLR 2026 Workshop DATA-FM · Mar 2, 2026
Reinforcement Learning (RL) has been shown to significantly boost reasoning capabilities of large language models (LLMs) in math, coding, and multi-hop reasoning tasks. However, RL fine-tuning requires abundant high-quality verifiable data,…
- PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation ModelsVignesh Kothapalli, Rishabh Ranjan, Valter Hudovernik, Vijay Prakash Dwivedi et al. · ICLR 2026 Workshop DATA-FM · Mar 2, 2026
Relational Foundation Models (RFMs) facilitate data-driven decision-making by learning from complex multi-table databases. However, the diverse relational databases needed to train such models are rarely public due to privacy constraints. W…
- Motion Capture is Not the Target Domain: Scaling Synthetic Data for Learning Motion RepresentationsFiras Darwish, George Nicholson, Aiden Doherty, Hang Yuan · ICLR 2026 Workshop DATA-FM · Mar 2, 2026
Synthetic data offers a compelling path to scalable pretraining when real-world data is scarce, but models pretrained on synthetic data often fail to transfer reliably to deployment settings. We study this problem in full-body human motion,…
- WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement LearningKuan Li, Zhongwang Zhang, Huifeng Yin, Rui Ye et al. · arXiv.org · Sep 16, 2025
Transcending human cognitive limitations represents a critical frontier in LLM training. Proprietary agentic systems like DeepResearch have demonstrated superhuman capabilities on extremely complex information-seeking benchmarks such as Bro…
- BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale PretrainingPratyush Maini, Vineeth Dorna, Parth Doshi, Aldo Carranza et al. · arXiv.org · Aug 14, 2025
Recent advances in large language model (LLM) pretraining have shown that simply scaling data quantity eventually leads to diminishing returns, hitting a data wall. In response, the use of synthetic data for pretraining has emerged as a pro…
- A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks AlignmentJean-Philippe Corbeil, Amin Dada, Jean-Michel Attendu, Asma Ben Abacha et al. · Annual Meeting of the Association for Computational Linguistics · May 15, 2025
High computation costs and latency of large language models such as GPT-4 have limited their deployment in clinical settings. Small language models (SLMs) offer a cost-effective alternative, but their limited capacity requires biomedical do…
- Synthetic Data Generation & Multi-Step RL for Reasoning & Tool UseAnna Goldie, Azalia Mirhoseini, Hao Zhou, Irene Cai et al. · arXiv.org · Apr 7, 2025
Reinforcement learning has been shown to improve the performance of large language models. However, traditional approaches like RLHF or RLAIF treat the problem as single-step. As focus shifts toward more complex reasoning and agentic tasks,…
- Scaling Laws of Synthetic Data for Language ModelsZeyu Qin, Qingxiu Dong, Xingxing Zhang, Li Dong et al. · arXiv.org · Mar 25, 2025
Large language models (LLMs) achieve strong performance across diverse tasks, largely driven by high-quality web data used in pre-training. However, recent studies indicate this data source is rapidly depleting. Synthetic data emerges as a …
- Synthetic Data Generation Using Large Language Models: Advances in Text and CodeMihai Nadǎş, Laura Dioşan, Andreea Tomescu · IEEE Access · Mar 18, 2025
This survey reviews how large language models (LLMs) are transforming synthetic training data generation in both natural language and code domains. By producing artificial but task-relevant examples, these models can significantly augment o…
- Synthetic data generation: a privacy-preserving approach to accelerate rare disease researchJorge M. Mendes, Aziz Barbar, Marwa Refaie · Frontiers Digit. Health · Mar 18, 2025
Rare disease research faces significant challenges due to limited patient data, strict privacy regulations, and the need for diverse datasets to develop accurate AI-driven diagnostics and treatments. Synthetic data—artificially generated da…
- Empowering Time Series Analysis with Synthetic Data: A Survey and Outlook in the Era of Foundation ModelsXu Liu, Taha İbrahim Aksu, Juncheng Liu, Qingsong Wen et al. · arXiv.org · Mar 14, 2025
Time series analysis is crucial for understanding dynamics of complex systems. Recent advances in foundation models have led to task-agnostic Time Series Foundation Models (TSFMs) and Large Language Model-based Time Series Models (TSLLMs), …
- Unlocking Post-hoc Dataset Inference with Synthetic DataBihe Zhao, Pratyush Maini, Franziska Boenisch, Adam Dziedzic · ICLR 2025 Workshop Data Problems Poster · Mar 6, 2025
The remarkable capabilities of large language models stem from massive internet-scraped training datasets, often obtained without respecting data owners' intellectual property rights. Dataset Inference (DI) enables data owners to verify una…
- Synthetic Data is an Elegant GIFT for Continual Vision-Language ModelsBin Wu, Wuxuan Shi, Jinqiao Wang, Mang Ye · Computer Vision and Pattern Recognition · Mar 6, 2025
Pre-trained Vision-Language Models (VLMs) require Continual Learning (CL) to efficiently update their knowledge and adapt to various downstream tasks without retraining from scratch. However, for VLMs, in addition to the loss of knowledge p…
- Improved YOLOv12 with LLM-Generated Synthetic Data for Enhanced Apple Detection and Benchmarking Against YOLOv11 and YOLOv10Ranjan Sapkota, Manoj Karkee · IFAC-PapersOnLine · Feb 26, 2025
This study evaluated the performance of the YOLOv12 object detection model, and compared against the performances YOLOv11 and YOLOv10 for apple detection in commercial orchards based on the model training completed entirely on synthetic ima…