Latest Speech Recognition Research Papers
The newest Speech Recognition papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Speech Recognition so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Speech Recognition papers in your inbox — free →Recent papers
- SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQHuy Hoang Le, Long-Bao Nguyen, Minh Tri Dao · arXiv · Sep 10, 2026
This paper describes our system for Task~2 of the second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge. We adapt Qwen3-Omni-30B-A3B-Instruct with a segment-evidence-aware data and post-training pipeline. A language m…
- Automatic Lyric Transcription for Greek Songs: Scaling and Task Composition Effects in Whisper AdaptationMaria Frangiadaki, Dimitrios Damianos, Kosmas Kritsis, Vassilis Katsouros · arXiv · Sep 10, 2026
Automatic Lyric Transcription (ALT) remains substantially more challenging than speech recognition due to melodic variability, rhythmic irregularity, and accompaniment interference. This is heightened in low-resource languages like Greek, w…
- Xiaomi-CocktailASR-1 Technical ReportYiru Zhang, Hang Su, Lichun Fan, Ying Zeng et al. · arXiv · Sep 10, 2026
Recently, large language model (LLM) based ASR models have achieved significant progress, yet they generally lack support for multi-speaker scenarios, where the cocktail party problem remains a critical bottleneck for further advancing ASR.…
- Do speech foundation models really learn words?Robin Huo, Ewan Dunbar · arXiv · Sep 9, 2026
Self-supervised speech foundation models are now used in a wide array of downstream applications, including traditional speech recognition and as the basis for tokens in speech-aware language models. Attempts to understand their usefulness …
- Pushing the Boundaries of Streaming Multi-Speaker ASR: A Systematic Study of Architectural Trade-offsTaejin Park, Ivan Medennikov, Kunal Dhawan, Weiqing Wang et al. · arXiv · Sep 9, 2026
Streaming multi-speaker ASR is a challenging task that must balance accuracy, latency, and efficiency while handling overlapping speech and maintaining coherent long-context modeling over extended conversations in an online fashion. We pres…
- Orukeet: Multilingual ASR with Frozen Gabor KernelsNathan Roll, Irene Yi, Büşra Marşan, Vianney Grenez et al. · arXiv · Sep 9, 2026
Orukeet replaces half of an adapted Parakeet encoder's temporal filters with 12,288 fitted Gabor kernels, freezes these replacements, and trains the remaining parameters on multilingual and multi-accent data. Final adaptation and checkpoint…
- Source-Adaptive Data Curation for Bilingual NVV-Aware ASRYuang Cao, Qirui Zhan, Jingbin Hu, Ziyu Zhang et al. · arXiv · Sep 9, 2026
Nonverbal vocalizations (NVVs), such as laughter, sighs, breaths, and coughs, convey affective and interactional information that conventional automatic speech recognition (ASR) systems often discard. We present a bilingual Mandarin-English…
- Adaptive dual stream neural encoder with privacy aware non-identifying data augmentation pipeline for low latency cross lingual speech recognitionM. Ramkumar, M. Marimuthu, R. Lakshminarayanan, V Sumanth · Scientific Reports · Sep 9, 2026
Abstract Cross-lingual speech recognition for real-time applications is constrained by the joint requirements of low latency, high recognition accuracy, and strict privacy over conversational audio. This research work addresses these constr…
- Influence of factors related to electrode array placement on the early speech recognition of cochlear implant recipients with and without functional hearing preservationMargaret T. Dillon, Madison E. Broome, Margaret E. Richter, Nicholas J. Thompson et al. · Frontiers in Audiology and ... · Sep 9, 2026
Cochlear implant (CI) and electric-acoustic stimulation (EAS) device users vary in their early speech recognition. This variability may be due in part to factors related to the electrode array placement, including frequency-to-place mismatc…
- Extending Epuskesmas With A Domain-Specific Large Language Model For Orthopaedic Documentation of Osteoarthritis And Osteoporosis: A Proof-Of-Concept Study Toward Indonesian Primary Health Care StrengtheningAhmad Azmul A. Irfan, Nur Ahmad Khatim, Achmad Zaki, Bisatyo Mardjikoen et al. · International Journal of He... · Aug 29, 2026
Background. Clinical documentation is a major driver of workload in primary care, and Indonesia's mandated transition to electronic medical records has increased the recording burden on community health centres (Puskesmas). A recent proof-o…
- Lightweight LLM-based Speech Recognition via KAN AdaptersYuxi Li, Yan Wang · Applied and Computational E... · Aug 23, 2026
In recent years, the combination of large language model (LLM) and pre-trained voice encoder has shown great potential in the field of automatic speech recognition (ASR). However, bridging the modal communication between acoustic characteri…
- Automatic speech recognition underperforms with diverse accents as measured by Word Error Rate and a semantic similarity measureAndrea Urqueta Alfaro, Karina A. Roundtree, Mark S. Pfaff, Nathan Bos et al. · Disability and Rehabilitati... · Aug 22, 2026
- Advances in automated speech transcription during cognitive testingDavid L. Woods, Michael Blank, Geraci Kristin, Isabella Jaramillo et al. · Zenodo (CERN European Organ... · Aug 22, 2026
We evaluated transcription accuracy of individual automatic speech recognition (ASR) engines and a consensus ASR (CASR) method using approximately 416,000 manually reviewed words from 453 older healthy participants completing California Cog…
- Advances in automated speech transcription during cognitive testingDavid L. Woods, Michael Blank, Geraci Kristin, Isabella Jaramillo et al. · Zenodo (CERN European Organ... · Aug 22, 2026
We evaluated transcription accuracy of individual automatic speech recognition (ASR) engines and a consensus ASR (CASR) method using approximately 416,000 manually reviewed words from 453 older healthy participants completing California Cog…
- A Bidirectional Indian Sign Language Translation System Using Speech Recognition and a 3D AvatarMackenzie Rebelo, Vaishnavi Vasant Lotulkar, V. Gaonkar, Athary Prabhu et al. · International Journal of Co... · Aug 20, 2026
- Poem meter classification of spoken Arabic poetry: integrating high-resource systems for a low-resource taskMaged S. Al-Shaibani, Zaid Alyafeai, Irfan Ahmad, Abdulkareem Saleh Alzahrani · PeerJ Computer Science · Aug 17, 2026
Arabic poetry is a cornerstone of Arab cultural and linguistic heritage, yet the computational analysis of spoken Arabic poetry remains critically underexplored. An open question is: how can the meter of a spoken Arabic poem be automaticall…
- The Correlation Between Exposure to English-Language Content on Instagram and Listening Comprehension among EFL Students at UIN Datokarama PaluMiftahul Inayah, Ruslin Ruslin, Hijrah Syam, Nur Asmawati et al. · Journal of General Educatio... · Aug 15, 2026
Listening comprehension is a foundational receptive skill in English as a Foreign Language (EFL) learning, yet learners often face challenges due to limited access to authentic spoken English. With the widespread use of social media, platfo…
- Extending a Formal Model of Predictive Coding to Human Spoken Word RecognitionBryce Rogers, Monica Yi-Chen Li, Thomas Hannagan, James S. Magnuson et al. · Neural Computation · Aug 14, 2026
Abstract There is a general consensus in theories of human speech recognition that humans engage in predictive processing during online speech processing. There are also claims that predictive processing is indicative of the operation of a …
- IMPLEMENTASI DEEP LEARNING HYBRID RESNET50V2-BIGRU PADA APLIKASI WEB VISUAL SPEECH RECOGNITION BAGI KOMUNITAS TULISalma Nurfauziah · Jurnal Informatika dan Tekn... · Aug 13, 2026
Visual Speech Recognition (VSR) merupakan teknologi komunikasi krusial bagi komunitas Tuli, terutama di lingkungan bising di mana sistem berbasis audio gagal berfungsi. Pengembangan VSR untuk bahasa Indonesia menghadapi tantangan kelangkaan…
- AI Based Voice Authentication Banking SystemSiva Pavin .R, Mayarudhran. S, MS.C.S.Selin Chandra · International Journal of Sc... · Aug 13, 2026
This project presents the development of a secure voice authentication system for banking applications using biometric speech recognition technology. The system enables users to perform secure login, registration, and financial transactions…
- Tibetan-PASEM: Phonology-Aware Speech Evidence Matching for Low-Resource Tibetan Written-Query Keyword SpottingYanze Guo, Xingmeng Guo, Zengguang Li, Jiaxin Song et al. · Sensors · Aug 12, 2026
Low-resource written-query keyword spotting detects a text-specified target in speech without spoken enrollment or full automatic speech recognition. We present Tibetan Phonology-Aware Speech Evidence Matching (Tibetan-PASEM), a method that…
- Voice to Empower: A Multilingual AI Solution for Tribal InclusionProf. Meghna K, Mohammed Rizvin M K, Sahwa K, Wafa et al. · Zenodo (CERN European Organ... · Aug 11, 2026
The Voice to Empower project introduces a multilingual artificial intelligence platform designed to bridge the communication barriers faced by India's tribal communities. Many tribal groups lack access to digital services because most platf…
- Development of SIBI and BISINDO Fingerspelling Translation Using Cloud-Based Voice Recognition: An Iterative ApproachMuhammad Fazli Ramadhani Sukma, Chandra Kusuma Dewa · bit-Tech · Aug 10, 2026
Communication barriers between the hearing public and the Deaf community remain a significant problem, partly because public knowledge of sign language is limited and conventional static media offer little interactivity for learning fingers…
- Artificial Intelligence in English Language Learning: Redefining Teaching Methods and Student PerformanceC. Shabharishwaran, Mr. P. KavinKumar, Mr B.Manojkumar · Stanzaleaf International Jo... · Aug 10, 2026
Artificial Intelligence is increasingly influencing the methods through which English is taught, practised, assessed, and learned. The emergence of generative AI, intelligent tutoring systems, automated writing evaluation, adaptive learning…
- Artificial Intelligence in Language Learning and Communication: Opportunities, Challenges, and Pedagogical ImplicationsNeetha A.J. · Zenodo (CERN European Organ... · Aug 9, 2026
Artificial Intelligence (AI) is rapidly transforming language learning and communication by offering adaptive, personalized, and context-aware learning experiences. AI-powered tools such as intelligent tutoring systems, natural language pro…
- Artificial Intelligence in Language Learning and Communication: Opportunities, Challenges, and Pedagogical ImplicationsNeetha A.J. · Zenodo (CERN European Organ... · Aug 9, 2026
Artificial Intelligence (AI) is rapidly transforming language learning and communication by offering adaptive, personalized, and context-aware learning experiences. AI-powered tools such as intelligent tutoring systems, natural language pro…
- Empowering Language Learners Through AI-Based Tools: A Pedagogical PerspectiveSanthosh Kumar R. · Zenodo (CERN European Organ... · Aug 9, 2026
Artificial Intelligence (AI) has become a powerful enabler in contemporary language education, offering innovative tools that support learner autonomy, engagement, and communicative competence. AI-based tools such as intelligent tutoring sy…
- Empowering Language Learners Through AI-Based Tools: A Pedagogical PerspectiveSanthosh Kumar R. · Zenodo (CERN European Organ... · Aug 9, 2026
Artificial Intelligence (AI) has become a powerful enabler in contemporary language education, offering innovative tools that support learner autonomy, engagement, and communicative competence. AI-based tools such as intelligent tutoring sy…
- Benchmarking Commercial Speech Recognition and Multimodal Large Language Models on Dysarthric Speech: Severity‐Stratified Baselines and Architecture‐Specific Prompting EffectsA. Alsayegh, Tariq Masood · International Journal of Intelligent Systems · Dec 19, 2025
Voice‐based human‐machine interaction has become a primary means of accessing intelligent systems, yet individuals with dysarthria are systematically excluded by persistent gaps in recognition accuracy. Whilst automatic speech recognition (…
- Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ LanguagesOmnilingual Asr team Gil Keren, Artyom Kozhevnikov, Yen Meng, C. Ropers et al. · arXiv.org · Nov 12, 2025
Automatic speech recognition (ASR) has advanced in high-resource languages, but most of the world's 7,000+ languages remain unsupported, leaving thousands of long-tail languages behind. Expanding ASR coverage has been costly and limited by …