Latest Machine Translation Research Papers
The newest Machine Translation papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Machine Translation so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Machine Translation papers in your inbox — free →Recent papers
- Human-Based Machine Translation Evaluation: A Multi-Dimensional Approach to Sentiment, Emotion, and Argumentation Preservation in Chinese-English TranslationJingshi Zhou · University of Liverpool · Jan 1, 2028
The research landscape of Machine Translation Evaluation (MTE) has traditionally been dominated by automated metrics that, while computationally efficient, often fail to capture the nuanced aspects of translation quality paramount to human …
- Nuha-Speech: Building General-Purpose Arabic Speech-LLMsYingzhi Wang, Reem Alhazzani, Muhammad Alqurishi · arXiv · Sep 10, 2026
As Speech Large Language Models (speech-LLMs) become increasingly multilingual, Arabic remains significantly underrepresented, highlighting the need for dedicated infrastructure to train and evaluate Arabic speech-LLMs. To address this gap,…
- Component-Aware Differential Privacy for Federated Multilingual Speech-LLMsJordi Luque, Fernando López, Aleix Sant · arXiv · Sep 10, 2026
Per-layer differential privacy (DP) clipping improves gradient fidelity in federated learning by allocating per-matrix clipping budgets proportional to parameter count. We show that this recipe breaks for speech large language models (speec…
- The Eloquence submission for Task 2 of the Interspeech 2026 MLC-SLM challengeJordi Luque, Lorenzo Concina, Marco Matassoni, Alessio Brutti et al. · arXiv · Sep 10, 2026
This paper details the Eloquence team's approach to Task 2 of the 2nd MLC-SLM challenge at Interspeech 2026, which involves multilingual Multiple-Choice Question Answering (MCQA) across 21 languages. Three approaches are explored. First, we…
- Complex-Text Robustness Evaluation and Failure Diagnosis for Low-Resource Multilingual Text-to-SpeechTianlun Zuo, Ziyu Zhang, Tingzhi Mao, Zhonghua Fu et al. · arXiv · Sep 10, 2026
Low-resource multilingual text-to-speech (TTS) systems have expanded language coverage, but their robustness under complex text inputs remains insufficiently diagnosed. Existing evaluations mainly focus on naturalness, speaker similarity, a…
- Structural priors for data-efficient language learningYana Veitsman, Jonas Mayer Martins, Jonathan Lautenschlager, Lisa Beinborn · arXiv · Sep 10, 2026
Efficient language learning requires methods to reduce the reliance on large data and computational resources. We investigate structural transfer: First training models on non-language data to induce useful priors for natural language. This…
- Cross-Lingual Clinical Annotation Projection as Constrained Text Generation: A Six-Language StudyÁlvaro Rey-Blanes, Francisco J. Moreno-Barea, Francisco J. Veredas · arXiv · Sep 10, 2026
Background: To determine whether cross-lingual clinical annotation projection can be formulated as a text-preserving, document-level generative task that produces verifiable character-level annotations for multilingual clinical corpus const…
- TransClean: A Benchmark for Detecting and Extracting Clean Translations from Large Language Model OutputsShenbin Qian, Yves Scherrer · arXiv · Sep 10, 2026
Large language models (LLMs) are increasingly used for machine translation, yet their outputs often contain additional text beyond the translation itself, such as language labels, explanations or bilingual repetitions, which we term transla…
- SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQHuy Hoang Le, Long-Bao Nguyen, Minh Tri Dao · arXiv · Sep 10, 2026
This paper describes our system for Task~2 of the second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge. We adapt Qwen3-Omni-30B-A3B-Instruct with a segment-evidence-aware data and post-training pipeline. A language m…
- Beyond Solver Verdicts: Generative Reward Models for AutoformalizationVikash Singh, Debargha Ganguly, Aman Goel, Ali Torkamani et al. · arXiv · Sep 10, 2026
Neurosymbolic systems rely on mathematical solvers to guarantee reasoning correctness, yet solvers are fundamentally blind to whether a formal translation maintains strict reference-equivalence to a designated formalization. We formalize th…
- Rosetta at AlexandriaX-2026: LoRA-Adapted NileChat for Context-Aware Dialectal Arabic Dialogue TranslationNada Esmaeil, Fathima Rena, Sibi Subhash, Osama Elgendy et al. · arXiv · Sep 9, 2026
This paper describes the Rosetta system for Subtask 1 (Context-Aware English-to-Dialectal Arabic Dialogue Translation) of the AlexandriaX shared task, participating in both constrained and unconstrained tracks. The approach fine-tunes a LoR…
- NOPE-HYPE: A Structured Simulation Workflow for Robust Speech-to-Text Across Diverse Acoustic EnvironmentsNiramay M. Patel, Bibek Behera, Raksha Sharma · arXiv · Sep 9, 2026
Robust speech-to-text translation systems should perform reliably across diverse acoustic conditions, yet practical pipelines lack controllable tools for systematic environment exploration. Large speech models remain sensitive to unseen aco…
- SalamandraTA at WMT 2026 Terminology Shared Task: Hard Examples Are Better TeachersXixian Liao, Maite Melero · arXiv · Sep 9, 2026
Terminology-aware translation asks for more than a correct translation: the output must use the exact terms a glossary prescribes. The standard recipe, fine-tuning on glossary-annotated translation pairs, hides an inefficiency: for most exa…
- Multi-Functional Embedding Models for Funder Name Disambiguation in Scientific Publication RecordsKanyao Han, Zhiwen You, Jinseok Kim, Jana Diesner · arXiv · Sep 9, 2026
Understanding the historical allocation and distribution of research funding advances our knowledge of how scientific research is supported across fields, institutions, and regions. However, large-scale analyses are hindered by the lack of …
- Improving Cross-Lingual Token Representations by Adding a Pinch of SALTGuillem Ramírez · arXiv · Sep 9, 2026
Cross-lingual sentence encoders enable scalable transfer across hundreds of languages, powering applications such as translation mining and zero-shot learning in low-resource settings. Although trained for sentence-level alignment, they are…
- BuzzASR: A Swarm of 100+ Monolingual Speech Recognition ModelsShivam Singh, Aditya Yadavalli, Catherine Arnett, Alex Warstadt · arXiv · Sep 9, 2026
We introduce BuzzASR, a collection of language-specialized fine-tuned Whisper models adapted for automatic speech recognition (ASR) in 102 languages. Large end-to-end Transformer-based ASR models such as Whisper have revolutionized ASR, but…
- From Fixed Keys to Readable Schemas: Small Language Models for Vehicle Agent Function CallsHamed Jafarzadeh Asl, Yuanhao Yu, Vahid Partovi Nia · arXiv · Sep 8, 2026
In-vehicle assistants must translate natural-language requests into accurate vehicle function calls under strict memory and latency constraints, making small language models (SLMs) attractive for on-device deployment. For such models, a key…
- SWORD: Wikidata-based Distortions Reveal Hidden Cross-Lingual Inconsistencies in LLM Factual Error RejectionSanghyeok Park, Minji Kang, Hosung Kwak, Jinhyuk Yun · arXiv · Sep 8, 2026
Modern LLMs demonstrate impressive multilingual performance, yet standard benchmarks primarily reward selecting correct answers rather than evaluating genuine factual understanding. We introduce Systematic Wikidata-based Object-Relation Dis…
- Improving Term Evaluation in Machine Translation: Variation MattersNicolas Dahan, Ziqian Peng, François Yvon, Rachel Bawden · arXiv · Sep 8, 2026
Terminology evaluation in machine translation (MT) usually assumes a single correct target form per source term. However, human translators routinely introduce variation that current metrics penalize as inconsistency. We examine how to acco…
- Same Values, Different Languages? From Multilingual Probing to Steering LLMs Toward Chinese Social ValuesYuemei Xu, Kexin Xu, Jian Zhou, Haoyu Lu et al. · arXiv · Sep 8, 2026
As Large Language Models (LLMs) are increasingly integrated into human society, aligning them with pluralistic social values has become a critical priority. However, whether LLMs exhibit consistent value preferences across languages remains…
- Compositional Multilingual and Behavioral Attribute SteeringHyun Gu Kang, Daniil Gurgurov, Tanja Baeumel, Josef van Genabith et al. · arXiv · Sep 8, 2026
This study examines the compositionality of steering vectors for language and behavioral control in large language models. Focusing on language, jailbreak, and conciseness, we investigate whether additive, training-free composition of attri…
- Structural Jailbreaks Generalize but Do Not Compound: A cross-provider and multilingual study of Involuntary In-Context LearningTejasvi C. Addagada · arXiv · Sep 8, 2026
Aligned language models fail under two independent pressures: the structural jailbreak class recently formalized as Involuntary In-Context Learning (IICL), which reframes a harmful request as the final missing cell of a data-labeling task c…
- Tracing Stereotypes from Representation to Output in Multilingual LLMsAriun-Erdene Tumurchuluun, Yusser Al Ghussin, Pinzhen Chen, Josef van Genabith et al. · arXiv · Sep 8, 2026
Multilingual LLMs show stereotype-related behavior that varies across languages, but behavioral scores do not show where the relevant information is represented or how it affects the output. To investigate these internal mechanisms, we comp…
- EviSI: An Evaluation Agent for Simultaneous InterpretingBen Yan, Zongyao Li, Daimeng Wei, Weidong Liu et al. · arXiv · Sep 8, 2026
Simultaneous speech-to-speech translation requires understanding, translation and spoken delivery while the source stream continues. To support timely delivery and limit accumulated delay, systems adopt reformulation and summarization, whic…
- When Metrics Reward the Worst Translations: Internalizing Cultural Reasoning for Social Media Translation EvaluationYiwen Qiu, Linjuan Wu, Dingming Li, Yizhou Liu et al. · arXiv · Sep 8, 2026
Automatic translation quality metrics trained on general-domain corpora systematically fail on social media content, where communicative intent is encoded in culturally loaded expressions (internet slang, homophonic ciphers, and platform-sp…
- IGT @ FinMMEval 2026 Task 2: Question-Type Prompting with Targeted Extraction for Multilingual Financial QAYuwen Chiu · arXiv · Sep 8, 2026
We present the IGT system for PolyFiQA Task 2 of the FinMMEval Lab at CLEF 2026, a multilingual financial question answering task over English SEC filings and multilingual news articles (English, Chinese, Japanese, Spanish, Greek) for four …
- Rethinking Sign Language Translation: The Impact of Signer Dependence on Model EvaluationKeren Artiaga, Sabyasachi Kamila, Haithem Afli, Conor Lynch et al. · EMNLP 2025 · Sep 7, 2026
Sign Language Translation has advanced with deep learning, yet evaluations remain largely signer-dependent, with overlapping signers across train/dev/test. This raises concerns about whether models truly generalise or instead rely on signer…
- EuroAlpaca: Task-Preserving Localisation of Instruction Data for European LanguagesAleix Sant, Jordi Luque, Carlos Escolano · arXiv · Sep 4, 2026
Machine translation (MT) offers a scalable way to extend English instruction-tuning data to multiple languages, but it can distort task-critical constraints and required outputs, creating corrupted training examples and degrading models tra…
- Discourse Dependency: A Continuous Criterion for Translation DifficultyAhrii Kim, Chanjun Park, Seong-heum Kim · arXiv · Sep 4, 2026
Recent calls for harder machine translation benchmarks have not clarified what difficulty should mean. We argue that one meaningful and currently unmeasured axis is referential reach, the distance a segment must look back into its document …
- MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical DomainSourav Malakar, Harshit Nigam, Akash Ghosh, Sriparna Saha et al. · arXiv · Sep 4, 2026
Time-series data in clinical settings is crucial for capturing dynamic changes in a patient's health over time, enabling timely diagnosis, personalized treatment, and early detection of critical events. However, the development of clinicall…