Latest Text Summarization Research Papers
The newest Text Summarization papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Text Summarization so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Text Summarization papers in your inbox — free →Recent papers
- Text summarization via global structure awarenessJiaquan Zhang, Chaoning Zhang, Shuxu Chen, Yibei Liu et al. · CoRR 2026 · Dec 31, 2026
Text summarization is a fundamental task in natural language processing (NLP), and the information explosion has made long-document processing increasingly demanding, making summarization essential. Existing research mainly focuses on model…
- Beyond Top Words: MonoTM for Topic Modeling with Interpretable Monosemantic FeaturesUna Joh, Bei Yu · arXiv · Sep 9, 2026
Topic models summarize large text corpora, but top-ranked words often provide only a limited representation of topic semantics. Sparse autoencoders (SAEs) offer a way to move beyond word-level descriptors by extracting interpretable feature…
- EviSI: An Evaluation Agent for Simultaneous InterpretingBen Yan, Zongyao Li, Daimeng Wei, Weidong Liu et al. · arXiv · Sep 8, 2026
Simultaneous speech-to-speech translation requires understanding, translation and spoken delivery while the source stream continues. To support timely delivery and limit accumulated delay, systems adopt reformulation and summarization, whic…
- KnowVis: Knowledge-Centric Visual Summarization for Video LecturesYi Xu, Yifan Hou, Xiaoyu Zhang · arXiv · Sep 3, 2026
Video lectures are valuable educational resources, but their dense and lengthy formats often overwhelm novice learners. This difficulty stems from a fundamental pedagogical mismatch: while videos deliver transient information linearly, huma…
- SGD-KV: Summarization Guided KV Cache CompressionZeyu Liu, Woomin Song, Xuandi Fu, Sai Muralidhar Jayanthi et al. · arXiv · Sep 3, 2026
Large language models (LLMs) face severe memory bottlenecks in long-context inference due to the linearly growing size of key-value (KV) caches. Existing KV cache compression techniques typically rely on simple heuristics, overlooking the d…
- Developing a coastal hazard prediction system in ice-infested waters – Part 1: High-resolution regional wave modeling in the Estuary and Gulf of St. LawrenceJérémy Baudry, Dany Dumont, David Didier, Pascal Bernatchez et al. · Natural hazards and earth s... · Sep 3, 2026
Abstract. This study is the first of a two-part paper that summarizes the development of a prototype coastal hazard prediction system providing short-term (+48 h) forecasts of the total water level (TWL) at 50 m resolution for the province …
- Improving Health Literacy through Lay Summarization of Radiological Reports: An Evaluation of BioNER and Retrieval-Augmented GenerationEgecan Çelik Evgin, İlknur Karadeniz, Olcay Taner Yıldız · arXiv · Sep 2, 2026
Radiology reports are written primarily for clinicians, and their specialized terminology often makes them difficult for patients to interpret. As a result, many patients turn to publicly available Large Language Models (LLMs) to help expla…
- Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture VideosS M Masrur Ahmed, Jaspal Subhlok · arXiv · Sep 1, 2026
Recorded lecture videos, often enhanced with search and summarization features, are a standard study resource. However, students cannot easily ask course specific questions or verify answers against an instructor's lecture. We report a seme…
- Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization EvaluationHimil Vasava, Ming Jiang · arXiv · Sep 1, 2026
LLM-based evaluators of natural language generation (NLG) quality are widely deployed as scoring tools and as automated training signals, yet the internal procedure by which they assign a rating remains poorly understood. We investigate thi…
- What Does an Agentic Software Engineering Benchmark Measure? Profiling Task Demands and Agent Behaviour Beyond What Category Labels RevealRadin Shayanfar, Keheliya Gallaba, Ahmed E. Hassan · arXiv · Sep 1, 2026
Agentic software engineering benchmarks are typically summarized by nominal category labels such as "bug fix" or "feature implementation," yet benchmarks carrying the same label are built through very different curation pipelines. A label t…
- Speech-to-SOAP: End-to-End Summarization of Medical Dialogues: KIT@BeTraC 2026Enes Yavuz Ugan, Fabian Retkowski, Yuka Ko, Thai-Binh Nguyen et al. · arXiv · Aug 25, 2026
With the advent of Large Language Models and its instruction following capabilities a promising application is the task of summarization. Within this domain of task the extractive sub-task of clinical protocolling has emerged as a topic of …
- Benchmarking generative AI tools for literature retrieval and summarization in genomic variant interpretationAndrea Gazzo, Silvia Berardelli, Matteo Biancospino, Lorenzo Cuollo et al. · Genome biology · Aug 22, 2026
Abstract Background Generative AI is increasingly used to extract structured information across domains, but its reliability in academic and clinical research, where precision and accuracy are essential, remains largely unexplored. This stu…
- StreamSoccer: Event-Driven Memory for Streaming Soccer CommentaryChenxi Shao, Bozhong Wang, Jiaxin Huang, Zhao Liu et al. · arXiv · Aug 20, 2026
Streaming video understanding requires models to causally update state as video arrives and organize growing history into semantic units that can evolve, persist, and be recalled under bounded computation and memory. This challenge is prono…
- Automated Summarization of Financial News Using Large Language Models and Retrieval-Augmented Generation: An Early Empirical Study (Fall 2023)Pranav Chandaliya · arXiv · Aug 20, 2026
Stock market analysts and investors face a daily challenge: too much financial news, too little time. Manually reading and synthesizing hundreds of company-specific articles is impractical, yet missing key information can directly affect in…
- On the Fragility of Self-Improving Agents: Variance, Task Order, and UnderspecificationQinyuan Ye, Yu Li, Yada Pruksachatkun, Jiaxin Zhang et al. · arXiv · Aug 18, 2026
Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank--have shown great promise in recent literature. However, the reliability aspects of these methods…
- Information Satisfaction: A Reader-Centered Axis for Summarization EvaluationIsabel Cachola, William Walden, Reno Kriz, Mark Dredze · arXiv · Aug 14, 2026
The majority of work on summarization evaluation focuses on general summary quality (e.g., ROUGE, BERTScore) or specific desired properties (e.g., readability, factuality). However, these metrics fail to measure the utility of a summary to …
- ASSERT: A Measurement Pipeline for GenAI AuditsRiccardo Fogliato, Abhinav Palia, Xiawei Wang, Emily Sheng et al. · arXiv · Aug 14, 2026
Audits of generative AI (GenAI) systems often summarize behavior as a reported rate: how often the audited system complies with policy. Researchers and stakeholders use that rate to compare systems, track regressions, and gate deployment. A…
- LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level ConsolidationDongfang Li, Zixuan Liu, Junmai Wang, Jiahe Huang et al. · arXiv · Aug 13, 2026
Long-horizon LLM agents must preserve information from past interactions to support future tasks. Existing memory systems typically rely on eager consolidation, invoking LLMs after each interaction to extract, summarize, or update memories.…
- Mapping and Measuring the Behavioral Evolution of Large Language ModelsDong Qiao, Chris Ding, Jicong Fan · arXiv · Aug 11, 2026
Benchmark leaderboards summarize how well a language model performs, but not how its behavior relates to that of other models or changes across generations. We characterize the output behavior of 32 models from six families using their resp…
- Enabling Functional Independence: A Scoping Review of Upper Extremity Assistive Devices for Adults With Progressive Neuromuscular DiseasesKatherine M. Burke, S. Johnson, Kayla Chomko, Angela Escalante et al. · Muscle & Nerve · Aug 7, 2026
ABSTRACT Background Assistive technology offers an important means of compensating for lost upper extremity function in adults with progressive neuromuscular diseases (NMD), enabling participation in daily activities, supporting independenc…
- Decomposed Entailment for Factuality Checking and Hallucination DetectionAchir Oukelmoun, Nasredine Semmar, Gaël De Chalendar · arXiv · Aug 6, 2026
The reliability of Large Language Models (LLMs) is often compromised by factual inconsistencies, including hallucinations---cases where generated content is not supported by the underlying source. We present HallDetect, a lightweight, refer…
- GenGA: Editable and Data-Grounded Graphical Abstract Generation for Academic PapersTakuro Kawada, Shunsuke Kitada, Hitoshi Iyatomi · arXiv · Aug 5, 2026
Graphical Abstracts (GAs) visually summarize the key findings of academic papers, playing a crucial role in facilitating the understanding of research content. Recently, advancements in vision-language models and image generation models hav…
- Decoupling Generation and Selection for Budget-Constrained Faithful SummarizationZeyu Wang, Guanghua Wang, Meng Xu · arXiv · Aug 4, 2026
Abstractive summarization models remain vulnerable to factual inconsistency, redundancy, and weak length control. We propose a modular generation-and-selection framework for sentence-budget-constrained summarization. A pretrained generator …
- LiveMem: Maintaining Memory State Continuity in Long-Running LLM InferenceZhichen Liu, Ruihan Sun, Hengjie Yang, Zipeng Wu et al. · arXiv · Aug 3, 2026
Long-running assistants and agents consume interaction streams that eventually outgrow the context. Existing context retention, summarization, and retrieval preserve access to selected history, but do not provide a persistent state over the…
- Addressable Recall Compaction for Long Context-Window Control in AI AgentsThang Dang, Yuma Ichikawa, Sakina Fatima, Koichi Shirahata · arXiv · Jul 27, 2026
Long-horizon LLM agents accumulate reasoning traces, actions, and tool observations that can eventually exceed a model's fixed context window. Existing compaction methods address this limitation by discarding, summarizing, or retrieving ear…
- CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language ModelsDengzhe Hou, Lingyu Jiang, Fangzhou Lin, Kazunori D Yamada · arXiv · Jul 27, 2026
LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize beyond the models used to define them. We introduce CogAren…
- Transformer-Assisted LLM-Based Source Code Summarisation: to Enable More Secure Software DevelopmentJesse Phillips, Tracy Hall, Paul Rayson, Mo El-Haj · arXiv · Jul 23, 2026
Neural Source Code Summarisation (NSCS) aims to generate natural language summaries of source code to improve developers' and maintainers' understanding of code. Source code summaries are vital during the maintenance phase of the Secure Sof…
- When Does Knowledge Distillation Hurt? Reliability-Aware Distillation for Low-Resource Language SummarizationDipto Sumit, Ankan Kumar Roy Srizon, Sadia Khair Rodela, Atia Haque Asha et al. · arXiv · Jul 22, 2026
Knowledge distillation (KD) is a standard approach for compressing sequence-to-sequence models, but its per-sample effects are rarely examined. On the BanSum Bangla summarization benchmark, we find that standard KD improves ROUGE-L by only …
- Dialogue Summarization with Emotion Dynamics Using Topic- and Participant-Centric DecompositionLinyun Xiang, Mark Neerincx, Stephanie Tan · arXiv · Jul 16, 2026
Existing text summarization research has focused much on monologic information (e.g., newspaper articles, reports) without accounting for the interaction between speakers or authors. In contrast, dialogues are a rich communication channel w…
- Where Should RL Post-Training Compute Go? Model Size, Search, Learning, and FeedbackPatrick Wilhelm, Odej Kao · arXiv · Jul 15, 2026
Reinforcement Learning (RL) post-training is increasingly used to adapt foundation models for reasoning, planning, and feedback-driven robot-learning pipelines, but constrained post-training resources are often summarized by a single total …