Latest Large Language Models Research Papers
The newest Large Language Models papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Large Language Models so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Large Language Models papers in your inbox — free →Recent papers
- Constraint-Based PhysicalismKendall Taylor · Zenodo (CERN European Organ... · Feb 17, 2028
DISCLOSURE:This paper presents the author's own original philosophical framework, refined through hudreds of iterative exchanges and adversarial critiques, with the author directing each stage of revision. The final text was generated with …
- Defying the Catholic Secondary School Enrollment Decline: A Case Study Exploring Strategic Enrollment Management Practices in an All-Boys Catholic High SchoolTraci A Koval · Seton Hall University eRepo... · May 15, 2027
Catholic secondary schools in the United States continue to face persistent enrollment decline, closures, and consolidations, creating an urgent need for sustainable and mission-centered strategies. This qualitative case study examined Hill…
- How Large Language Models Are Reshaping Skills and Job Requirements for Public Health Professionals in Saudi ArabiaMulfi Alkhinjar · Scholarship @ Claremont (Th... · Jan 1, 2027
Context: Large Language Models (LLMs) such as ChatGPT, Gemini, and DeepSeek are transforming professional work across sectors by enhancing information processing and decision support. In public health, these technologies offer the potential…
- The Sociolinguistics of Machine Identity: LLM Personality and Ideology PropagationGuangni Li · Knowledge Commons (Lakehead... · Dec 31, 2026
Do large language models (LLMs) possess a measurable "personality," and how do the linguistic properties of training corpora shape their cognitive style and downstream reasoning? This paper approaches these questions from a sociolinguistic …
- LKValues: Aligning Large Language Models with Sri Lankan Societal ValuesNethmi Muthugala, Supryadi, Surangika Ranathunga, Nisansa de Silva et al. · arXiv · Jul 22, 2026
Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms. This results in the mishandling of local values in multilingual societies such as Sri Lanka that have their unique cultural dynamic…
- Notes to Self: Can LLMs Benefit from Experiential Abstractions?Chang Liu, Xinyu Li, Artur Dubrawski · arXiv · Jul 22, 2026
Humans distill experience into reusable abstractions, e.g., strategies and cautionary reminders, and apply them to gradually solve problems more effectively. We study whether Large Language Models (LLMs) can similarly benefit from such expe…
- Test-Time Training for Modality Order Consistency in Vision-Language ModelsAditi Gupta, Yossi Gandelsman · arXiv · Jul 22, 2026
We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented. Across three models and three benchmarks, image first prompting consistently outperform…
- PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative InferenceNiqi Lyu, Pengtao Shi, Wei Qiu, Jianlin Zhong et al. · arXiv · Jul 22, 2026
Large language models (LLMs) provide strong reasoning capabilities but are expensive to serve at scale, whereas small language models (SLMs) are cheaper but less reliable on difficult problems. We introduce PyroDash, a cost-aware framework …
- The Blessing of Dimensionality: How Near-Orthogonality in High-Dimensional Spaces Explains Temporal PortabilityAbigail Woodring, Adrian Chan, Rana Muhammad Shahroz Khan, Sukwon Yun et al. · arXiv · Jul 22, 2026
Fine-tuning has been widely used to adapt large language models (LLMs) for domain-specific tasks. Parameter efficient fine-tuning (PEFT) methods such as low-rank adaptation (LoRA) are frequently used to reduce computational costs. PortLLM i…
- Sound Probabilistic Safety Bounds for Large Language ModelsMahdi Nazeri, Anne-Kathrin Schmuck, Sadegh Soudjani, Alessandro Abate · arXiv · Jul 22, 2026
We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given prompt. We study a new application of the Clopper-Pearson confidence intervals to obtain pro…
- Which Values Do LLMs Confuse? A Schwartz-Based Recognition StudyAndrei Chetvergov, Stepan Ukolov, Timofei Sivoraksha, Alexander Evseev et al. · arXiv · Jul 22, 2026
Large language models are increasingly evaluated through the values they endorse, but such evaluations presuppose that models can identify the value expressed in a concrete situation. We study this prerequisite as controlled top-1 recogniti…
- PoTRE: Test-Time Reasoning inspired by Cognitive HeterogeneityAnmol Kankariya, Sercan Ö. Arık · arXiv · Jul 22, 2026
While Large Language Models (LLMs) excel at many tasks, they frequently struggle with complex reasoning that requires long-horizon planning and iterative error correction. Furthermore, standard single-stream prompting proves brittle when mo…
- The Maskability Index: Predicting Task-Objective Alignment in Pretrained Language ModelsAhmad Pouramini, Mahsa Afsharzadeh · arXiv · Jul 22, 2026
Large-scale pretrained language models such as T5 and BERT have demonstrated strong capabilities for generating structured knowledge. However, their performance depends on how closely the prompting strategy matches the objectives used durin…
- Exposure is Optional: Learning Unlike Coordination in Language ModelsJiamu Luo, Shane Steinert-Threlkeld · arXiv · Jul 22, 2026
Coordination, a fundamental linguistic structure, remains a subject of intense debate, and its exact nature continues to elude theoretical linguistics. A common view holds that only same-category constituents can be conjoined, which has bee…
- On the Systematic Challenges of Culturally Loaded Machine Translation: Dream of the Red Chamber as the Cultural LensYiming Wang, Jiayuan Di · arXiv · Jul 22, 2026
Culturally loaded translation poses unique challenges for machine translation (MT), as meanings are deeply embedded in socio-cultural contexts beyond surface linguistic forms. Although large language models (LLMs) have enabled MT systems to…
- HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question AnsweringAbdessalam Bouchekif, Mohammed-En-Nadhir Zighem, Salah Eddine Bekhouche, Hichem Telli et al. · arXiv · Jul 22, 2026
Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify. Existing hallucination benchmarks often provide response-level labels, with limited support for i…
- surprisal is Not a TheoryAndrés Buxó-Lugo, Aniello De Santo, Morgan Grobol, Ryan J. Hubbard et al. · arXiv · Jul 22, 2026
Surprisal Theory is often characterized as a computational-level explanation per (Marr, 1982). We argue in this work that, even though a computational level narrative has been used to support "representation-agnostic research" within comput…
- Gotta Catch them all: the modes of SycophancyShreyans Jain, Alexandra Yost, Amirali Abdullah · arXiv · Jul 22, 2026
Large language models often align with users' beliefs at the expense of factual accuracy, a behavior known as sycophancy. Prior mechanistic studies largely treat sycophancy as a single behavioral dimension that can be uniformly amplified or…
- SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPODDongfang Li, Xiaodong Luo, Ruoyu Sun, Xuhui Chen et al. · arXiv · Jul 22, 2026
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficien…
- Back to Back with a Copy: A Computational Analysis of AI-Generated Visual Contemporary Art PastichesAnca Dinu, Andreiana Mihail, Andra-Maria Florescu, Claudiu Creanga et al. · arXiv · Jul 22, 2026
The aim of this paper is twofold. First, it investigates whether newer generative models are getting better at pastiching contemporary artworks. Second, it explores the consistency of the multidimensional nature of stylistic evaluation acro…
- OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party SkillsQiyuan Liu, Tingfeng Hui, Kun Zhan, Kaike Zhang et al. · arXiv · Jul 22, 2026
LLM-based agents leverage third-party skills to extend their capabilities in open-world scenarios. However, third-party skills can introduce extra security vulnerabilities, as seemingly harmless skills can contain latent safety risks that o…
- Understanding the Impact of Linguistic Realization Choices on LLM Stance with Causal TracingLangchen Huang, Sebastian Padó, Franziska Weeber · arXiv · Jul 22, 2026
Large language models (LLMs) are known to be sensitive to prompt and input formulations. However, existing studies have focused on lexical realization and largely ignored constructional choice. This paper studies whether linguistic construc…
- ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language ModelsKaran Goyal, Afreen Hossain, Debojyoti Das, Vishal Bhutani · arXiv · Jul 22, 2026
Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic a…
- Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval ResultsYanyu Chen, Yue Li, Yongyi Cui, Dongsheng Shi et al. · arXiv · Jul 22, 2026
Retrieval-augmented large language models frequently face contexts that interleave useful evidence with misleading statements or instruction-like content. Blanket refusal discards valid evidence, whereas uncritical adoption yields incorrect…
- The Two-Process Theory of Machine Self-ReportHubert Plisiecki, Filip Chmielewski, Kacper Dudzic, Anna Sterna et al. · arXiv · Jul 22, 2026
Language models are increasingly asked to self-report, informing safety evaluations, public understanding, and model-welfare debates. Yet their reports are elicited with human questionnaires never validated for models or ad hoc prompts of u…
- Solar Open 2 Technical ReportSungrae Park, Sanghoon Kim, Gyoungjin Gim, Jungho Cho et al. · arXiv · Jul 22, 2026
We present Solar Open 2, a 250B-A15B Mixture-of-Experts language model built for long-horizon agentic tasks, scaled up from Solar Open 1 (Solar Open 100B). To hold entire agent trajectories in a single context, Solar Open 2 reaches a 1M-tok…
- Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language ModelMarkus J. Buehler · arXiv · Jul 22, 2026
Large language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the governing physics. Here we show that materials science mechanism information in the open-weight google/gemm…
- Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language ModelsNetanel Eliav · arXiv · Jul 21, 2026
Practitioners make three prompt-design decisions with almost no controlled evidence behind them: how to format instructions and context (markdown, plain text, prose, or tabular), how many simultaneous instructions a system prompt can carry …
- Inference-Time Steering for Cross-Lingual Factual Consistency in LLMsAlexander Manev · arXiv · Jul 21, 2026
Although Large Language Models (LLMs) demonstrate remarkable multilingual fluency, their internal knowledge representations remain disproportionately biased toward high-resource languages. This leads to cross-lingual factual inconsistency, …
- AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion DraftersYu-Yang Qian, Hao-Cong Wu, Chen Chen, Jiacheng Sun et al. · arXiv · Jul 21, 2026
Speculative decoding, in which a lightweight draft model first generates a draft sequence that is then verified in parallel by the target model, has become a prevalent paradigm for accelerating large language model inference. Recent work su…