Latest Large Language Models Research Papers
The newest Large Language Models papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Large Language Models so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Large Language Models papers in your inbox — free →Recent papers
- Constraint-Based PhysicalismKendall Taylor · Zenodo (CERN European Organ... · Feb 17, 2028
DISCLOSURE:This paper presents the author's own original philosophical framework, refined through hudreds of iterative exchanges and adversarial critiques, with the author directing each stage of revision. The final text was generated with …
- Large Language Model-Based Data Querying and Analysis for Manufacturing IndustrySadeer Beden, Zuha Shahid, Cinzia Giannetti, Arnold Beckmann · Cronfa (Swansea University) · Jan 1, 2027
- Large language models accurately identify decision reasons in verbal reportsK. Fuławka, R. ; https://orcid.org/0000-0002-9908-9556 Hertwig, D. ; https://orcid.org/0000-0002-4008-8022 Wulff · MPG.PuRe (Max Planck Society) · Dec 31, 2026
- The Sociolinguistics of Machine Identity: LLM Personality and Ideology PropagationGuangni Li · OpenAlex · Dec 31, 2026
Do large language models (LLMs) possess a measurable "personality," and how do the linguistic properties of training corpora shape their cognitive style and downstream reasoning? This paper approaches these questions from a sociolinguistic …
- The Sociolinguistics of Machine Identity: LLM Personality and Ideology PropagationGuangni Li · Knowledge Commons (Lakehead... · Dec 31, 2026
Do large language models (LLMs) possess a measurable "personality," and how do the linguistic properties of training corpora shape their cognitive style and downstream reasoning? This paper approaches these questions from a sociolinguistic …
- Moral Consistency Variance: A Pilot Benchmark for Decision Stability under Moral Prompt Perturbations in Large Language ModelsQiao Liang · OpenAlex · Dec 31, 2026
Static moral question-answering benchmarks do not test whether a model's decision distribution remains stable when the same dilemma is rephrased without changing the underlying facts. This paper introduces Moral Consistency Variance (MCV), …
- SIP-AI-08 Ontological Non-Determinism and the Geometry of Coherence CollapseJohn Richard Smith, SHAI / HATI3 · Zenodo (CERN European Organ... · Dec 4, 2026
AbstractThis paper presents a unified mathematical framework for understanding AI coherence collapse, synthesisedthrough a novel multi-AI research methodology. We treat the non-determinism of large language models asan ontological property …
- SIP-AI-08 Ontological Non-Determinism and the Geometry of Coherence CollapseJohn Richard Smith, SHAI / HATI3 · Zenodo (CERN European Organ... · Dec 4, 2026
AbstractThis paper presents a unified mathematical framework for understanding AI coherence collapse, synthesisedthrough a novel multi-AI research methodology. We treat the non-determinism of large language models asan ontological property …
- Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated DataAtindra Jha, Margaret Li, Jure Leskovec, Percy Liang et al. · arXiv · Sep 10, 2026
As the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activated Transformers, but the effects of data repetition remains …
- Nuha-Speech: Building General-Purpose Arabic Speech-LLMsYingzhi Wang, Reem Alhazzani, Muhammad Alqurishi · arXiv · Sep 10, 2026
As Speech Large Language Models (speech-LLMs) become increasingly multilingual, Arabic remains significantly underrepresented, highlighting the need for dedicated infrastructure to train and evaluate Arabic speech-LLMs. To address this gap,…
- Domain-Specific Hallucination Detection in Large Language ModelsVarun Teja Chundru, Debasmita Biswas · arXiv · Sep 10, 2026
Large language models generate fluent text that can contain unfaithful claims -- a phenomenon known as hallucination. We present a multi-signal detection pipeline combining fine-tuned DeBERTa-v3 classification, Monte Carlo (MC) Dropout unce…
- The Last AI Built by Humans: Toward Genuine Recursive Self-ImprovementYi Duan, Ying Liu, Zirui Tang, Haodong Chen et al. · arXiv · Sep 10, 2026
Recursive self-improvement (RSI) enables AI systems to turn experience and feedback into persistent changes that improve both their capabilities and the process of future improvement. We first use the Headroom-Closed Index (HCI) to reveal t…
- Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language ModelLisa Bylinina · arXiv · Sep 10, 2026
A language model normally begins training with random word embeddings: whatever 'banana' means must be learned from training corpora. I implement St. Augustine's picture of word learning, meaning by ostension, for a small masked language mo…
- RetroThinker: Enabling Retrospective Thinking in Speech LLMsYi-Jen Shih, Puyuan Peng, Abdelrahman Mohamed, David Harwath · arXiv · Sep 10, 2026
Speech large language models (SpeechLLMs) offer reduced latency and retain paralinguistic nuances that are typically lost in cascaded automatic speech recognition (ASR) and text-based LM architectures. However, they continue to lag behind t…
- SpecGuard: Inference-Time Backdoor Detection For FreeRui Wen, Ahmed Salem, Andrew Paverd, Mark Russinovich et al. · arXiv · Sep 10, 2026
Large language models are often fine-tuned, shared, or downloaded from third parties, so a deployed model may carry a hidden backdoor that behaves normally on benign inputs but switches to attacker-controlled behavior when a secret trigger …
- Beyond Word Error Rate: A Switch Aware Evaluation of ASR and Audio Language Models on English Yoruba Code-Switched SpeechChibuzor Okocha, Christan Earl Grant · arXiv · Sep 10, 2026
Automatic speech recognition (ASR) systems and audio language models (audio LMs) now report low error rates on monolingual benchmarks, but their behavior on code switched speech in low resource, diacritic rich languages remains poorly chara…
- Whisper-Based Speech Transcription from Videos Across Multiple Languages for Cross-Cultural UnderstandingMichael Picheny · arXiv · Sep 10, 2026
Cross-cultural understanding has become increasingly important in today's highly connected, cross-national world. The success of LLM-based technologies is now driving the development of automated tools to aid understanding for nonnative peo…
- The widening evaluation gap in medical large language model research 2023 to 2026Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif · arXiv · Sep 10, 2026
Large language models are superseded every few quarters; clinical evidence takes years. We asked whether medical research is keeping pace with the systems it evaluates. PubMed returned 11,628 records for January 2023 to June 2026 across fou…
- Recognizing Is Not Reversing: A Controlled Inversion Test of Fact-Preserving News FramingYi Liu · arXiv · Sep 10, 2026
Large language models (LLMs) are increasingly used to analyze and rewrite news, yet current framing studies mainly evaluate generation, detection, or whether rewritten text appears more neutral. They do not directly show whether a model can…
- Component-Aware Differential Privacy for Federated Multilingual Speech-LLMsJordi Luque, Fernando López, Aleix Sant · arXiv · Sep 10, 2026
Per-layer differential privacy (DP) clipping improves gradient fidelity in federated learning by allocating per-matrix clipping budgets proportional to parameter count. We show that this recipe breaks for speech large language models (speec…
- RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM SafetyAdithiyan Rajan Indira Saravanan, Kathleen C. Fraser · arXiv · Sep 10, 2026
Allowing large language models (LLMs) to retrieve information from a set of trusted documents can increase reliability and reduce hallucination. However, recent work has demonstrated that retrieval-augmented generation (RAG) can have uninte…
- LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language GenerationDongfang Zhao · arXiv · Sep 10, 2026
Large language model serving costs scale directly with output sequence length, yet standard preference alignment often inflates response verbosity without improving utility. We study whether the parameterization of post-training updates aff…
- Why Does Post-Training Quantization Work?Yuxiang Chen, Michael Beyer, Jun Zhu, Jianfei Chen · arXiv · Sep 10, 2026
Post-training quantization compresses large language models (LLMs) by storing their weights at reduced precision, and each quantized weight introduces an error into the hidden states. Naively, these errors should accumulate with depth and c…
- Negative Self-Distillation: Learning to Reason by Avoiding FlawsRongcan Pei, Zhepei Wei, Shuyao Xu, Xinyu Zhu et al. · arXiv · Sep 10, 2026
On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However,…
- Structured Transforms for Low-Overhead Quantization of Language ModelsDaria Cherniuk, Alexander Rudikov, Boris Kashin, Ivan Oseledets · arXiv · Sep 10, 2026
We revisit Kashin-decomposition-based weight quantization for large language models and propose an improved algorithm with stronger convergence properties and structured, efficient orthogonal transforms. The method retains the core factoriz…
- A Training-Free, Alignment-Free Approach to Corporate Intelligence: Application to SEC FilingsJean-François Delpech · arXiv · Sep 10, 2026
High-dimensional dense text embeddings and large language models face real obstacles in financial-disclosure analysis: context-window limits, hallucination risk, high computational cost, and the arbitrary rotation of vector spaces across in…
- Structural priors for data-efficient language learningYana Veitsman, Jonas Mayer Martins, Jonathan Lautenschlager, Lisa Beinborn · arXiv · Sep 10, 2026
Efficient language learning requires methods to reduce the reliance on large data and computational resources. We investigate structural transfer: First training models on non-language data to induce useful priors for natural language. This…
- Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy OperationalizationAyan Majumdar, Shounak Paul, Pushpdeep Singh, Ines Abdelaziz et al. · arXiv · Sep 9, 2026
The growing complexity of content moderation policies presents a critical challenge for their consistent operationalization. While foundation models possess the basic capabilities needed to confront this challenge, whether they can reliably…
- Retrofitting Code Using LLMs to Support Exceptional BehaviorLinghan Zhong, Jiyang Zhang, Jayanth Srinivasa, Junyi Jessy Li et al. · arXiv · Sep 9, 2026
Exception Related Code (ERC), which includes throw statements, conditions (if statements) that guard those throw statements, and try/catch blocks, is an essential component of software systems, allowing developers to detect and handle excep…
- Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMsKillian Steunou, Yannis Tevissen, Mounîm A. El Yacoubi · arXiv · Sep 9, 2026
Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their strong performance o…