Latest Machine Translation Research Papers
The newest Machine Translation papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Machine Translation so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Machine Translation papers in your inbox — free →Recent papers
- Human-Based Machine Translation Evaluation: A Multi-Dimensional Approach to Sentiment, Emotion, and Argumentation Preservation in Chinese-English TranslationJingshi Zhou · Open MIND · Jan 1, 2028
The research landscape of Machine Translation Evaluation (MTE) has traditionally been dominated by automated metrics that, while computationally efficient, often fail to capture the nuanced aspects of translation quality paramount to human …
- Inference-Time Steering for Cross-Lingual Factual Consistency in LLMsAlexander Manev · arXiv · Jul 21, 2026
Although Large Language Models (LLMs) demonstrate remarkable multilingual fluency, their internal knowledge representations remain disproportionately biased toward high-resource languages. This leads to cross-lingual factual inconsistency, …
- The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine TranslationMichael Jungo, Aixiu An · arXiv · Jul 21, 2026
Reinforcement learning with verifiable rewards (RLVR) has been established as a viable paradigm for the post-training of Large Language Models (LLMs), including downstream tasks, such as Neural Machine Translation (NMT). With the latest res…
- Reasoning Before Translation: Enhancing Legal Machine Translation with Structured ReasoningAixiu An, Michael Jungo, Eloi Eynard, Mark Drenhaus et al. · arXiv · Jul 21, 2026
Neural machine translation (NMT) in the legal domain is a linguistically and conceptually demanding task, primarily due to the complexity of legal language and the high level of precision it requires. The recent emergence of reasoning-capab…
- Translation as Augmentation: Effect of Translated Data on Assessment of DifficultyYiheng Wu, Jue Hou, Roman Yangarber · arXiv · Jul 21, 2026
Reliable Text Difficulty Assessment is a prerequisite for valid text simplification workflows and personalized learning applications. However, the development of robust assessment models is severely hindered by a critical bottleneck: the sc…
- From a Multilingual Streaming ASR Backbone to Kenyan-Language Systems: Data-Centric Adaptation of Nemotron 3.5 for Kikuyu, Dholuo, and KalenjinMark Gatere · arXiv · Jul 21, 2026
Automatic speech recognition (ASR) for African languages is constrained by orthographic inconsistency, annotation artifacts, missing audio, speaker and domain imbalance, and evaluation procedures that differ from deployment. We present an e…
- AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM AgentsKunlun Zhu, Xuyan Ye, Zhiguang Han, Yuchen Zhao et al. · arXiv · Jul 21, 2026
LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay execution traces but provide little support for identifying the root cause or transl…
- Is EEG-to-Text Feasible in Real-World Scenarios? An In-Depth Analysis Using a Neuropsychology-Inspired BenchmarkZihan Zhang, Yu Bao, Xiao Ding, Tianyi Jiang et al. · arXiv · Jul 21, 2026
Translating brain signals into text could restore communication for people with severe paralysis, yet practically usable systems to date rely on invasive electrocorticography (ECoG). Electroencephalography (EEG) offers a non-invasive altern…
- LatentMT: Machine Translation with Latent ReasoningWei-Rui Chen, Samar M. Magdy, Chiyu Zhang, Wenhui Zhu et al. · arXiv · Jul 21, 2026
Latent-reasoning looped language models (LoopLMs) offer a different scaling path for machine translation (MT): instead of increasing parameter count or emitting explicit chain-of-thought tokens, they spend additional recurrent computation i…
- Building a European Multilingual Evaluation Dataset: The MMLU Localisation Project within the EMT NetworkPilar Sánchez-Gijón, Susana Valdez, Sofía Calvo Del Barrio, Florence Bellemont et al. · arXiv · Jul 20, 2026
This paper reports on a collaboration between the Directorate-General for Translation (DGT) and the European Master's in Translation (EMT) to localise the MMLU dataset into 11 European languages. Beyond creating a more inclusive benchmark f…
- Enabling Multilingual Privacy Policy Audits: Large-Scale Analysis of Spanish Mobile AppsMarcos Moran, David Rodriguez, Luka Nenadic, Norman Sadeh et al. · arXiv · Jul 20, 2026
Automated analyses of privacy policies enable large-scale assessments of transparency in digital ecosystems, yet existing auditing pipelines remain predominantly English-centric. This limits their ability to systematically evaluate multilin…
- Expanding the Lexicon of Ge'ez Based African Languages: A Comparative Study of Amharic and TigrinyaHailay Kidu Teklehaymanot, Debela Desalegn Yadeta, Wolfgang Nejdl · arXiv · Jul 16, 2026
Multilingual pre-trained language models (PLMs) exhibit degraded performance on low-resource, non-Latin-script languages, driven by high out-of-vocabulary (OOV) rates and excessive subword fragmentation that result from Latin-script-centric…
- LLM Evaluators are Biased across LanguagesEj Zhou, Lucas Resck, Zheng Hui, Anna Korhonen · arXiv · Jul 16, 2026
LLM evaluators (trained reward models and prompted LLM-as-a-Judge) are routinely validated via pairwise accuracy. In a multilingual setting, this operates under the premise that high pairwise accuracy implies reliable, language-neutral scor…
- Can an Old Dog Be Taught New Tricks? Taking LLMs Beyond Sentence Level TranslationAlaina Brandt · arXiv · Jul 15, 2026
Automatic translation systems, from CAT tools to MT, overwhelmingly treat translation as a sentence-by-sentence act. This paper asks whether LLMs can be moved beyond that paradigm through whole-document, corpus-informed translation. We pres…
- DeltaMerge-LowRes: Composing Language and Task Deltas for Low-Resource AdaptationSon Ha Xuan, Xuan-Bach Le, Phat T. Tran-Truong · arXiv · Jul 15, 2026
Adapting a multilingual encoder to a new language \emph{and} a new task with only a few hundred gold examples is a common low-resource NLP setting, yet the two axes are usually fused via an expensive language--task fine-tuning run. We ask w…
- High-Order Question Generation in a Multilingual Educational ContextSuna-Şeyma Uçar, Itziar Aldabe, Nora Aranberri, Orphée De Clercq · arXiv · Jul 15, 2026
Critical thinking is a fundamental skill that helps learners move beyond simple memorization. One way to develop this skill is through high-order questioning. However, crafting such questions remains a challenge for educators, and classroom…
- The Test Oracle Problem in Synthetic LLM-as-Judge Corpora: Disappearance, Distortion and a Validation ProtocolSerkan Ballı · arXiv · Jul 15, 2026
Studies of bias in LLM-as-judge systems typically build synthetic corpora by prompting an LLM to generate a hallucinated answer to pair with a factual one, then presenting both to a judge. We report a case in which this generation step sile…
- MET: Theory-Grounded and Culture-Aware Multilingual Moral ReasoningAyoung Lee, Ryan Kwon, Yunxiang Zhang, Yuxuan Liu et al. · arXiv · Jul 13, 2026
Language models are increasingly used for moral decision-making across diverse linguistic and cultural contexts, yet existing work overlooks multilinguality on three aspects: 1) multilingual evaluation benchmarks use direct translation, fai…
- STEP: Career-Path Recommendation via Temporal and Educational Trajectory ModelingIman Johary, Guillaume Bied, Alexandru C. Mara, Tijl De Bie · arXiv · Jul 13, 2026
Career paths encode decades of skill acquisition, role transitions, and educational investment, and understanding them at scale underpins workforce planning, labor market policy, and job recommendation. Resumes are a rich source of informat…
- Direct Image-to-Modern Vietnamese Translation of Han-Nom Manuscripts via Multimodal RLHF Preference AlignmentThi Kim Trang Vo, Nghia Hieu Nguyen, Ha Minh Tan · arXiv · Jul 13, 2026
Translating Han-Nom manuscripts into modern Vietnamese is challenging because historical pages are often degraded, the script contains rare logographic characters, and parallel supervision is limited. We propose a multimodal RLHF preference…
- Q-BridgeNet: A Quantization Network for Cross-Lingual Sign Language TranslationLiqian Feng, Lintao Wang, Xiaochen Liu, Anusha Withana et al. · arXiv · Jul 13, 2026
Most sign language translation (SLT) methods focus on isolated native sign-spoken pairs (e.g., American Sign Language - English). Extending language-specific SLT models to multilingual translation would improve accessibility by enabling com…
- Unified Gradient Projection: Language-Balanced Continual Learning for Multilingual Low-Resource ASRZiang Ren, Guodong Lin, Yuchen Ai, Kaize Tan et al. · arXiv · Jul 13, 2026
Large-scale pretrained ASR models such as Whisper exhibit strong multilingual capabilities. However, fine-tuning on low-resource languages often causes catastrophic forgetting. Although continual learning mitigates this issue, existing meth…
- The Nuts and Bolts of Natural Language to SQL Translation: A Systematic Analysis of Model Pipeline Optimisation Approaches and their InteractionsFilip Klubicka, Vasudevan Nedumpozhimana, Sneha Rautmare, Bora Caglayan et al. · arXiv · Jul 12, 2026
In the age of large language models, Natural Language to SQL (NL2SQL) translation remains an open problem with many useful applications. We explore interactions between several NL2SQL pipeline extensions to inspire development of more light…
- Toward Real-Time Sentence-Level Sign Language TranslationThanh-Hoang Nguyen Doan · arXiv · Jul 10, 2026
Most sign language understanding systems operate at the level of isolated signs, limiting their usefulness in natural communication. We study sentence-level sign language translation (SLT) with the primary goal of real-time deployment rathe…
- Test-Time Scaling for Small VLMs on Multilingual Visual MCQSpiros Baxevanakis, Peng-Jian Yang · arXiv · Jul 10, 2026
Test-time scaling (TTS) reliably improves reasoning in large language models, but whether it transfers to small open vision-language models remains unclear. We examine this on EXAMS-V, a multilingual visual multiple-choice benchmark, compar…
- VTaMo: Video-Text Alignment Model for Sign Language TranslationJunyi Hu, Zhewen He, Haomian Huang, Aoxiang Yang et al. · arXiv · Jul 10, 2026
Sign language translation (SLT) converts continuous sign videos into spoken language text. Gloss-free approaches leverage pre-trained visual encoders and language models but rely on implicit cross-modal alignment from translation supervisio…
- Echoes Across Vietnam's Highlands, Delta, and Coast: A Multilingual Corpus for Cham, Khmer, and Tay-NungAnh Trac Duc Dinh, Khang Nhat Hoang Vo, Vinh Cong Doan, Tai Tien Ta et al. · arXiv · Jul 9, 2026
Vietnam's ethnic minority languages are almost absent from the field of Natural Language Processing (NLP), and the challenge goes beyond data scarcity: Cham, Khmer, and Tay-Nung differ sharply in script, Vietnamese contact, and standardizat…
- Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational SpeechHao Wu, RongQi Han, Zhen Wang, Wei Liang et al. · arXiv · Jul 9, 2026
This paper describes our self-designed system for Task 1 of the MLC-SLM 2026 Challenge for multilingual two-speaker conversational speech. The system combines a modular speaker diarization front end with a challenge-adapted Qwen3-ASR-1.7B r…
- SynthAVE: Scalable Synthetic Labeling for E-Commerce with LLM-Arena ValidationAndrea Scarinci, Virginia Negri, Brayan Impata, Suleiman Khan et al. · arXiv · Jul 8, 2026
Fine-tuning large language models (LLMs) for e-commerce attribute extraction requires labeled data representative across thousands of product types, attributes, and multiple languages. This combinatorial scale translates to millions of anno…
- Automated Compliance Mapping in Cloud Security with Domain-Adapted Sentence TransformersJohn Bianchi, Luca Petrillo, Fabio Martinelli, Marinella Petrocchi · arXiv · Jul 7, 2026
Mapping cloud security controls to technical metrics is currently a manual process. This paper proposes domain adaptation of Sentence Transformer models to automate it. We build a training corpus of 3,499 semantic pairs from five European s…