Latest Long-Context Modeling Research Papers
The newest Long-Context Modeling papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Long-Context Modeling so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Long-Context Modeling papers in your inbox — free →Recent papers
- Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing[object Object], [object Object], [object Object], [object Object] et al. · ACL (1) 2026 · Dec 31, 2026
Diffusion Large Language Models (dLLMs) deliver strong long-context processing capability in a non-autoregressive decoding paradigm. However, the considerable computational cost of bidirectional full attention limits the inference efficienc…
- Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing[object Object], [object Object], [object Object], [object Object] et al. · CoRR 2026 · Dec 31, 2026
Diffusion Large Language Models (dLLMs) deliver strong long-context processing capability in a non-autoregressive decoding paradigm. However, the considerable computational cost of bidirectional full attention limits the inference efficienc…
- Distance generalization in transformers: why bother with positional encoding?Daniel Henrik Nevermann, Claudius Gros · arXiv · Sep 10, 2026
Out-of-distribution length generalization, namely to extrapolate a task from short to longer context, has been studied intensively for transformers. Here we focus on distance generalization, which probes performance when inter-token distanc…
- A Training-Free, Alignment-Free Approach to Corporate Intelligence: Application to SEC FilingsJean-François Delpech · arXiv · Sep 10, 2026
High-dimensional dense text embeddings and large language models face real obstacles in financial-disclosure analysis: context-window limits, hallucination risk, high computational cost, and the arbitrary rotation of vector spaces across in…
- Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMsKillian Steunou, Yannis Tevissen, Mounîm A. El Yacoubi · arXiv · Sep 9, 2026
Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their strong performance o…
- Learning Length-Extrapolatable Recurrent ModelsHanwen Jiang · arXiv · Sep 8, 2026
Recurrent models provide a natural path to long-context modeling, yet models trained with backpropagation through time (BPTT) often fail beyond their training horizon. Classical analyses emphasize gradients that vanish or explode along temp…
- Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory PortabilityAnkit Goyal, Jaideep Ray · arXiv · Sep 4, 2026
Model upgrades are routine; memory migrations are not. An agent can keep the same memory store and still forget: a new model may interpret old notes differently, mixed embedding versions may break retrieval, and repair may fail without the …
- Trace as State: Reasoning Traces as Conditional States for Long-Context TransformersXu Zou, Jie Tang · arXiv · Sep 2, 2026
Transformers process information causally, but long-context reasoning may depend on task state discovered only later. We formalize this mismatch through conditional state update tasks. For causal state update processors, providing the condi…
- Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research WorkflowsMiao Liu, Zhizhe Liu · arXiv · Aug 25, 2026
Large language models (LLMs) are increasingly deployed as AI analysts to process financial disclosures and support AI-assisted investment decisions. Yet such systems are usually evaluated by what they can retrieve, not whether retrieved inf…
- On the Fragility of Self-Improving Agents: Variance, Task Order, and UnderspecificationQinyuan Ye, Yu Li, Yada Pruksachatkun, Jiaxin Zhang et al. · arXiv · Aug 18, 2026
Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank--have shown great promise in recent literature. However, the reliability aspects of these methods…
- Proteus: Incremental Memory Activation for Long-Context Sequence ModelingReza Bayat, Ali Behrouz, Vahab Mirrokni, Aaron Courville · arXiv · Aug 17, 2026
The quadratic cost of attention-based sequence models for long contexts has motivated a growing line of research on memory-based models that can compress context into a compact state. However, most existing memory models expose a static mem…
- SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context ReasoningHaonan He, Haodi Lei, Yun Luo, Haoran Zhang et al. · arXiv · Aug 14, 2026
On-policy distillation (OPD) offers a promising way to transfer reasoning capabilities from stronger teacher models, but applying it to long-context reasoning teachers and short-context students introduces practical challenges, including to…
- KV Cache Compression Through the Lens of Transform CodingHannah Laus, Claudio Mayrink Verdun, Hao Wang, Flavio du Pin Calmon et al. · arXiv · Aug 14, 2026
The key-value (KV) cache stores information from past tokens and is a major memory bottleneck in long-context inference. Existing quantization methods address this bottleneck by representing the KV cache uniformly with lower-precision data …
- Information Abundance Paradox: Long-Context Training Undermines Parametric KnowledgeArda Uzunoglu, Benjamin van Durme, Daniel Khashabi · arXiv · Aug 12, 2026
Large language models are increasingly trained and deployed with long contexts that span documents, code repositories, and interaction histories. This scaling reflects the implicit assumption that training on longer contexts will only help …
- CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAGGyuwan Kim, Cheoneum Park, Tao Yang · arXiv · Aug 7, 2026
Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retrieved contexts for higher efficiency, while significant information redundancy and noise still remain…
- ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache CompressionYuhang Zhan, Lisi Chen, Shuo Shang · arXiv · Jul 31, 2026
KV cache compression is essential for efficient long-context inference. Existing eviction methods permanently discard unselected tokens and consequently remove their aggregate contribution to attention. Merging-based alternatives preserve m…
- DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code SearchRaphaël Sourty, Antoine Chaffin, Paulo Roberto Moura Junior, Amélie Chatelain · arXiv · Jul 29, 2026
State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for training retrieval models and study how English supervision transfers to multilingual retr…
- Kimi K3: Open Frontier IntelligenceKimi Team, Tongtong Bai, Yifan Bai, Yiping Bao et al. · arXiv · Jul 27, 2026
We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which…
- Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language ModelsNetanel Eliav · arXiv · Jul 21, 2026
Practitioners make three prompt-design decisions with almost no controlled evidence behind them: how to format instructions and context (markdown, plain text, prose, or tabular), how many simultaneous instructions a system prompt can carry …
- Context Forcing: Consistent Autoregressive Video Generation with Long ContextShuo Chen, Cong Wei, Sun Sun, Tiancheng SHEN et al. · ICML 2026 regular · Apr 30, 2026
Recent approaches to real-time long video generation typically employ streaming tuning strategies, attempting to train a long-context student using a short-context (memoryless) teacher. In these frameworks, the student performs long rollout…
- QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement LearningFanqi Wan, Weizhou Shen, Shengyi Liao, Yingcheng Shi et al. · arXiv.org · May 23, 2025
Recent large reasoning models (LRMs) have demonstrated strong reasoning capabilities through reinforcement learning (RL). These improvements have primarily been observed within the short-context reasoning tasks. In contrast, extending LRMs …
- Clinical ModernBERT: An efficient and long context encoder for biomedical textSimon A. Lee, A. Wu, Jeffrey N. Chiang · arXiv.org · Apr 4, 2025
We introduce Clinical ModernBERT, a transformer based encoder pretrained on large scale biomedical literature, clinical notes, and medical ontologies, incorporating PubMed abstracts, MIMIC IV clinical data, and medical codes with their text…