Latest Text-to-Speech Research Papers
The newest Text-to-Speech papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Text-to-Speech so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Text-to-Speech papers in your inbox — free →Recent papers
- Continuous-Time Acoustic Modelling with Neural Controlled Differential EquationsMattias Cross, Minghui Zhao, Anton Ragni · arXiv · Sep 10, 2026
Text-to-speech (TTS) models commonly address text--speech alignment by expanding phone-level encoder states to frame-level decoder inputs using predicted durations. While this length-regulation step resolves alignment structurally, this use…
- Complex-Text Robustness Evaluation and Failure Diagnosis for Low-Resource Multilingual Text-to-SpeechTianlun Zuo, Ziyu Zhang, Tingzhi Mao, Zhonghua Fu et al. · arXiv · Sep 10, 2026
Low-resource multilingual text-to-speech (TTS) systems have expanded language coverage, but their robustness under complex text inputs remains insufficiently diagnosed. Existing evaluations mainly focus on naturalness, speaker similarity, a…
- Post-Training Zero-Shot TTS for Fine-Grained Emotion and Duration Control via Natural LanguageLianru Gao, Yujie Guo, Yong Qin · arXiv · Sep 10, 2026
Audiobook narration, conversational agents, and audiovisual dubbing require speech that conveys changing emotions and adapts its pacing within a single utterance. But most existing TTS systems typically rely on utterance-level style conditi…
- Deterministic Prompting for Speaker-Stable Low-Resource Greek TTSGeorgios Syllas, Efthymios Georgiou, Kosmas Kritsis, Alexandros Potamianos · arXiv · Sep 9, 2026
Modern TTS systems approach human quality for high-resource languages but degrade when clean speech data is scarce. Modern Greek exemplifies this, lacking the curated corpora behind state-of-the-art synthesis. We propose a data curation rec…
- Integrating human linguistic insights into AI: theory-driven representation for multilingual text-to-speechCong Zhang, Huinan Zeng, Huang Liu, Jiewen Zheng · Phonetica · Aug 26, 2026
This paper explores the integration of human linguistic insights into multilingual text-to-speech (TTS) systems by evaluating the Featurally Underspecified Lexicon (FUL) as a theory-driven input representation. Unlike data-intensive end-to …
- PERANCANGAN ANNOUNCEMENT SYSTEM BERBASIS WEB DENGAN FITUR TEXT TO SPEECH UNTUK AKSESIBILITAS INFORMASI PRAKIRAAN CUACA PERAIRAN SELATAN JAWARosyid - · Jurnal Informatika dan Tekn... · Aug 13, 2026
Kawasan pesisir selatan Jawa memiliki potensi bahaya gelombang tinggi dan angin kencang, sementara akses publik terhadap informasi prakiraan cuaca perairan dari BMKG masih terbatas. Penelitian ini bertujuan merancang dan mengimplementasikan…
- MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow MatchingXingwei Sun, Heinrich Dinkel, Gang Li, Jiahao Mei et al. · arXiv · Aug 12, 2026
Generating coherent audio scenes that simultaneously blend speech, music, and sound effects remains a significant challenge. Current approaches typically rely on a disjointed pipeline where a frozen, decoupled text encoder feeds a separate …
- Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech TokenizationPeijie Chen, Zhuanling Zha, Zhipeng Nie, Weijie Wu et al. · arXiv · Aug 12, 2026
In current zero-shot text-to-speech systems, conventional semantic tokenizers are typically optimized using supervised automatic speech recognition or self-supervised learning objectives. However, due to the inherent nature of speech, seman…
- Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker EncoderHuaxuan Wang, Huimin Wang, Ruiyu Zhang, Yingjie Li et al. · arXiv · Aug 12, 2026
Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero-shot TTS systems still depend on audio prompt transcripts at inference time. This dependency limits …
- Luna-TTS Family Technical ReportFeng Yin, Shuai Shi, Junjie Zheng, Kechenying Zhou et al. · arXiv · Aug 12, 2026
Modern text-to-speech (TTS) is dominated by autoregressive (AR) codec language models, whose left-to-right decoding brings latency that grows with utterance length, error accumulation along the committed prefix, and an artificial generation…
- Comprehensive review of traditional and deep learning approaches to text to speech synthesisTien-Dung Do, Dinh-Tuan Nguyen, Trong-Khoi Nguyen · Discover Artificial Intelli... · Aug 12, 2026
- Voice to Empower: A Multilingual AI Solution for Tribal InclusionProf. Meghna K, Mohammed Rizvin M K, Sahwa K, Wafa et al. · Zenodo (CERN European Organ... · Aug 11, 2026
The Voice to Empower project introduces a multilingual artificial intelligence platform designed to bridge the communication barriers faced by India's tribal communities. Many tribal groups lack access to digital services because most platf…
- Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded DimensionsOluwanifemi Bamgbose, Simon Rosen, Jash Shah, Lindsay Devon Brin et al. · arXiv · Aug 10, 2026
Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspect…
- MADBench: A Benchmark for Modality-Aware Audio Deepfake DetectionYanqiu Li, Yang Xiao, Jisheng Bai, Bin Chen et al. · arXiv · Aug 10, 2026
Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated o…
- ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow MatchingHan Zhu, Wei Kang, Zengwei Yao, Liyong Guo et al. · Automatic Speech Recognition & Understanding · Jun 16, 2025
Existing large-scale zero-shot text-to-speech (TTS) models deliver high speech quality but suffer from slow inference speeds due to massive parameters. To address this issue, this paper introduces ZipVoice, a high-quality flow-matching-base…
- EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text PromptingGuanrou Yang, Chen Yang, Qian Chen, Ziyang Ma et al. · ACM Multimedia · Apr 17, 2025
Human speech goes beyond the mere transfer of information; it is a profound exchange of emotions and a connection between individuals. While Text-to-Speech (TTS) models have made huge progress, they still face challenges in controlling the …
- Model architectures to extrapolate emotional expressions in DNN-based text-to-speechKatsuki Inoue, Sunao Hara, Masanobu Abe, Nobukatsu Hojo et al. · Speech Commun. 2021 · Jan 1, 2021
Highlights•We propose model architectures to synthesize emotional speech in extrapolation.•The target speaker borrows emotional expressions from the data of other speakers.•Neural Network is trained with multi-speaker and multi-emotional sp…
- Normal-to-Lombard adaptation of speech synthesis using long short-term memory recurrent neural networksBajibabu Bollepalli, Lauri Juvela, Manu Airaksinen, Cassia Valentini-Botinhao et al. · Speech Commun. 2019 · Jan 1, 2019
In this article, three adaptation methods are compared based on how well they change the speaking style of a neural network based text-to-speech (TTS) voice. The speaking style conversion adopted here is from normal to Lombard speech. The s…
- Quantitative intonation modeling of interrogative sentences for Mandarin speech synthesisYa Li, Jianhua Tao, Wei Lai, Xiaoying Xu · Speech Commun. 2017 · Jan 1, 2017
Previous intonational research on Mandarin has mainly focused on the prosody modeling of statements or the prosody analysis of interrogative sentences. To support related speech technologies, e.g., Text-to-Speech, the quantitative modeling …
- The Romanian speech synthesis (RSS) corpus: Building a high quality HMM-based speech synthesis system using a high sampling rateAdriana Stan, Junichi Yamagishi, Simon King, Matthew P. Aylett · Speech Commun. 2011 · Jan 1, 2011
This paper first introduces a newly-recorded high quality Romanian speech corpus designed for speech synthesis, called “RSS”, along with Romanian front-end text processing modules and HMM-based synthetic voices built from the corpus. All of…
- Analysis of statistical parametric and unit selection speech synthesis systems applied to emotional speechRoberto Barra-Chicote, Junichi Yamagishi, Simon King, Juan Manuel Montero et al. · Speech Commun. 2010 · Jan 1, 2010
We have applied two state-of-the-art speech synthesis techniques (unit selection and HMM-based synthesis) to the synthesis of emotional speech. A series of carefully designed perceptual tests to evaluate speech quality, emotion identificati…
- Statistical parametric speech synthesisHeiga Zen, Keiichi Tokuda, Alan W. Black · Speech Commun. 2009 · Jan 1, 2009
This review gives a general overview of techniques used in statistical parametric speech synthesis. One instance of these techniques, called hidden Markov model (HMM)-based speech synthesis, has recently been demonstrated to be very effecti…
- Training intonational phrasing rules automatically for English and Spanish text-to-speechJulia Hirschberg, Pilar Prieto · Speech Commun. 1996 · Jan 1, 1996
We describe a procedure for acquiring intonational phrasing rules for text-to-speech synthesis automatically, from annotated text, and some evaluation of this procedure for English and Spanish. The procedure employs decision trees generated…