Latest Dialogue Systems Research Papers
The newest Dialogue Systems papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Dialogue Systems so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Dialogue Systems papers in your inbox — free →Recent papers
- SWRouter: Similarity-Contractive Window Routing for Multi-Turn Large Language Model ConversationsYu Wang, Yuchen Li, Rui Kong, Xinran Chen et al. · arXiv · Sep 10, 2026
Large language models exhibit complementary strengths, motivating routing methods that dispatch each query to the most suitable model. Although existing routers are effective in single-turn settings, they do not directly transfer to multi-t…
- SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQHuy Hoang Le, Long-Bao Nguyen, Minh Tri Dao · arXiv · Sep 10, 2026
This paper describes our system for Task~2 of the second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge. We adapt Qwen3-Omni-30B-A3B-Instruct with a segment-evidence-aware data and post-training pipeline. A language m…
- ProMediConv: Benchmarking Proactive Conversational Agents in Legal Dispute MediationZesheng Wei, Mengfan Li, Wenhao Liu, Yixin Zhang et al. · arXiv · Sep 10, 2026
Dispute mediation is essential for maintaining social harmony and resilience, yet developing skilled mediators is costly and time-consuming. Existing LLM-based mediation research remains limited by unrealistic task formulations, low-fidelit…
- Using Semantic Uncertainty to Estimate Transition Relevance in Turn-takingMuhammad Umair, Jan P. de Ruiter · arXiv · Sep 10, 2026
Turn-taking is a fundamental mechanism that governs when interlocutors speak and listen. Although Spoken Dialogue Systems (SDS) exploit a range of linguistic, acoustic, and non-verbal cues, they produce ill-timed responses in unscripted int…
- Rosetta at AlexandriaX-2026: LoRA-Adapted NileChat for Context-Aware Dialectal Arabic Dialogue TranslationNada Esmaeil, Fathima Rena, Sibi Subhash, Osama Elgendy et al. · arXiv · Sep 9, 2026
This paper describes the Rosetta system for Subtask 1 (Context-Aware English-to-Dialectal Arabic Dialogue Translation) of the AlexandriaX shared task, participating in both constrained and unconstrained tracks. The approach fine-tunes a LoR…
- SocialRL: Refining LLMs' Social Intelligence through Multi-turn Reinforcement Learning and Reward DesignJianing Wang, Xintao Wang, Aili Chen, Jie Shi et al. · arXiv · Sep 9, 2026
Social intelligence enables agents to read social context, infer intent, and adapt over sustained dialogue. As language models become autonomous collaborators, it is central to building effective and trustworthy human-AI interaction. Existi…
- ConversationalVoice: Full-Duplex Speech Data from Real Conversations through Source-Faithful Reconstruction and Conversation-Grounded ExpansionRichard Yucheng He, Baodong Cao, Chen Xu, Yihang Liu et al. · arXiv · Sep 8, 2026
Full-duplex speech models require training data that preserves turn-taking, overlap, interruption, and backchannel behavior, yet these signals are entangled across speakers in noisy real-world recordings. We present Conversational Voice, a …
- Controlling and Assessing Appropriate Persona Use in LLM-based Dialogue GenerationJongkyung Shin, Inkyu Lee, Chiehyeon Lim · arXiv · Sep 4, 2026
In persona-based dialogue generation (PDG), LLMs often overuse persona attributes by incorporating them regardless of dialogue context, resulting in unnatural responses. Despite its practical significance, the underlying causes remain unexp…
- TRILOGUE: A Trilingual Spoken Dialogue Fact-Checking Benchmark with Evidence and Paired AudioChaewan Chun, Meruyert Aristombayeva, Jiyoung Choi, Mahjabin Nahar et al. · arXiv · Sep 3, 2026
Modern misinformation is often heard before it is read, yet fact-checking systems are still evaluated mainly on clean written claims. Spoken dialogue remains different even when systems operate on transcripts: claims may be distributed acro…
- Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue SynthesisSanyuan Chen, Min-Jae Hwang, Sho Inoue, Anna Sun et al. · arXiv · Sep 3, 2026
We present Alignment-Free Text-Audiobox (Text-AB), a unified framework for high-quality voice dubbing and full-duplex dialogue synthesis. Building on a Diffusion Transformer trained with a flow-matching objective, Text-AB departs from the A…
- FiMI Banking: A Sovereign Model for Indian Retail BankingNPCI AI Research Team, Aman Kumar, Asit Desai, Chandra Bhushan et al. · arXiv · Sep 3, 2026
Banks need conversational systems that can answer product questions, assist customers with account-related requests, and operate safely within strict operational and regulatory constraints. General-purpose language models do not reliably me…
- Beyond Majority Vote: Multi-Perspective Adjudication for Medical Hallucination DetectionJoe Cecil, Marjorie Freedman · arXiv · Sep 3, 2026
Understanding the frequency of factual errors in chatbot-generated text and evaluating systems that detect these errors is critical for determining chatbot safety. Yet factual-error detection is often treated as a single-pass, single-annota…
- RuleMem: Active Rule Memory for Long-Term Conversational AgentsXingyuan Zeng, Zuohan Wu, Quanming Yao, Yue Wang et al. · arXiv · Sep 3, 2026
Question answering agents in long-term conversations must reason over massive, temporally dispersed dialogue histories. However, existing memory mechanisms primarily treat past information as \textit{passively} stored facts, leading to sema…
- When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational AgentsWen-Yu Chang, Yun-Nung Chen · arXiv · Sep 3, 2026
Large language models (LLMs) are increas- ingly deployed as long-horizon conversational agents, motivating growing interest in mem- ory systems. However, existing benchmarks primarily evaluate memory through QA-style probing rather than in-…
- FPCO-Dialog: A Multi-Turn False-Premise Benchmark for Correction and Cooperation in Vision-Language ModelsJiayuan Ma, Yuqi Lu, Weiyang Guo, Chenrui Wang et al. · arXiv · Sep 3, 2026
Vision-language models (VLMs) are increasingly deployed in multi-turn settings where users may describe visual content with incorrect assumptions. Yet existing evaluations rarely isolate how models respond when the same visually grounded fa…
- Decoupling Turn-Taking from Semantics: A Decoupled Data Approach for Finite-State-Machine-Based Full-Duplex DialogueYihang Li, Chenhui Chu · arXiv · Sep 3, 2026
The Neural Finite State Machine (NFSM) framework offers a pragmatic path to full-duplex dialogue by serializing turn-taking control and response generation onto a single causal tape under the standard next-token prediction objective, thereb…
- Predictors of Loneliness in Older Adults Using Multimodal Analysis of Speech and LanguageVinmay Khandode, Sai Karthik Kosuri, Neil K. R. Sehgal, Adam Greene et al. · arXiv · Sep 2, 2026
Loneliness is a critical public health issue among older adults, linked to higher risks of depression, cognitive decline, and mortality. Scalable, objective methods for its detection remain limited, particularly in natural conversational co…
- PragAlign: Feedback-Guided Pragmatic Alignment for Controlled Synthetic Dialogue GenerationSmitha Muthya Sudheendra, Jaideep Srivastava · arXiv · Sep 2, 2026
Synthetic dialogue generation can support research in privacy-restricted service settings, but generated conversations must preserve communicative intent, affective meaning, and natural dialogue flow. We introduce PragAlign, a feedback-guid…
- Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn JailbreakingSiyu Chen, Haoran Wang, Xiaojian Li, Yao Huang et al. · arXiv · Sep 2, 2026
Multi-turn jailbreak attacks demonstrate that harmful intent can be distributed across dialogue, yet existing methods obscure what conversational mechanisms drive vulnerability. We introduce BLUEPRINT, a safety-evaluation framework separati…
- A Tri-Agent Framework for Evaluating and Aligning Question Clarification Capabilities of Large Language ModelsYikai Zhao, Saurabh Pandey, Pradeep Kumar Misra · arXiv · Sep 2, 2026
Large Language Models (LLMs) are increasingly deployed in interactive systems where understanding user intent precisely is paramount. A key capability for such systems is effective question clarification, especially when user queries are am…
- Threat and Remedy: AI’s Technological and Normative Role in Democratic Discourse and Counter SpeechDiana Rieger, Mario Haim · Cogitatio (Cogitatio) · Sep 2, 2026
Artificial intelligence (AI) has fundamentally transformed online discourse, serving simultaneously as a source of and a potential solution to threats to democracy. This article conceptually examines AI’s multifaceted role in digital public…
- Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture VideosS M Masrur Ahmed, Jaspal Subhlok · arXiv · Sep 1, 2026
Recorded lecture videos, often enhanced with search and summarization features, are a standard study resource. However, students cannot easily ask course specific questions or verify answers against an instructor's lecture. We report a seme…
- SDARE-Bench: Evaluating Large Language Models on Conversational Stigma Detection and Response in Dyadic and Group DialogueStephanie Fong, Yiwen Jiang, Zimu Wang, Hongxi Yang et al. · arXiv · Sep 1, 2026
Large Language Models (LLMs) are increasingly used in advice seeking and decision making that may affect social judgements. Despite stigma's profound effects on people and communities, benchmarks remain scarce. Existing general-domain evalu…
- FinLifeBench: Exhaustive Life-Event History and Financial-State Reconstruction from Longitudinal Banking DialogueHangyeul Lee, Juyoung Oh, Jaeyong Ko, Sunmin Kim et al. · arXiv · Sep 1, 2026
Repeated banking interactions require assistants to maintain complete, current, and traceable customer records as life changes emerge incidentally in routine requests. Existing benchmarks emphasize question answering, bounded episodes, or t…
- PersuaRL: Reinforcement Learning-Driven Multi-Expert Selection for Persuasive Dialogue Generation in InsuranceRohan Kirti, Akash Ghosh, Aryan Vats, Niladri Ghosh et al. · arXiv · Sep 1, 2026
Large Language Models (LLMs) are revolutionizing digital communication by powering conversational agents deployed across domains such as customer service, digital sales, and insurance. These agents, built on LLMs, can understand user input,…
- ClinTraceBench: Source-Verifiable Longitudinal Clinical Reasoning over EHR-Derived DialoguesHuimin Wang, Zhengyi Zhao, Yutian Zhao · arXiv · Sep 1, 2026
Clinical LLM assistants must reason over multi-visit patient trajectories, yet whether the compact history representations used to scale them---retrieval, structured timelines, LLM summaries, agentic memory---preserve the longitudinal signa…
- TEIDAN: A Multilingual Multiparty Dialogue CorpusTaiga Mori, Koji Inoue, Mikey Elmers, Divesh Lala et al. · arXiv · Sep 1, 2026
Multi-party interaction is a central setting for human communication and a necessary target for human-agent interaction systems that must participate in group conversation. Yet available corpora often focus on meetings, task-oriented intera…
- Mind the Gap: Theory-of-Mind-Grounded Friction for Epistemic AlignmentYifan Zhu, Kyeongmin Rim, James Pustejovsky · arXiv · Aug 31, 2026
Productive dialogue alignment requires distinguishing \emph{surface coordination} (acknowledgments and smooth task progression) from \emph{epistemic alignment} (convergence of belief states); standard preference-based methods typically opti…
- UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational MemoryPeijun Qing, Fobo Shi, Soroush Vosoughi · arXiv · Aug 31, 2026
Long-term memory is increasingly important for conversational agents, yet existing benchmarks primarily measure memory through pointwise factual recall: whether a system can recover isolated facts or event-level details from prior interacti…
- Stranger, Fan, or Peer? A Systematic Study on the Role of Interlocutor in Persona-Based Dialogue GenerationDaniela Occhipinti, Malvina Nissim, Marco Guerini · arXiv · Aug 28, 2026
Persona-based dialogue systems are usually conditioned on speaker biography, but dialogues involve at least two participants, and who has access to whose biography can vary across training, inference, and evaluation. Prior work often neglec…