Latest Chain-of-Thought Reasoning Research Papers
The newest Chain-of-Thought Reasoning papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Chain-of-Thought Reasoning so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Chain-of-Thought Reasoning papers in your inbox — free →Recent papers
- The Sociolinguistics of Machine Identity: LLM Personality and Ideology PropagationGuangni Li · OpenAlex · Dec 31, 2026
Do large language models (LLMs) possess a measurable "personality," and how do the linguistic properties of training corpora shape their cognitive style and downstream reasoning? This paper approaches these questions from a sociolinguistic …
- Multi-dimensional hierarchical temporal alignment for improved temporal commonsense reasoning in large language modelsGe Yan, Hai-Tao Yu, Lei Chao · Institutional Repositories ... · Nov 1, 2026
Benefiting from recent advances in generative AI and large language models (LLMs), current LLM-driven AI systems are rapidly reshaping how we reason, plan, and make decisions. Yet endowing these systems with human-like temporal intelligence…
- MindTopo: Can Foundation Models Reason in Topological Space?Yunfei Ge, Anbang Liu, Qineng Wang, Johnalbert Garnica et al. · arXiv · Sep 10, 2026
Spatial reasoning depends not only on metric properties such as distance, angle, and shape, but also on topological relations that remain invariant under continuous deformation. Cognitive science identifies these relations as foundational t…
- RetroThinker: Enabling Retrospective Thinking in Speech LLMsYi-Jen Shih, Puyuan Peng, Abdelrahman Mohamed, David Harwath · arXiv · Sep 10, 2026
Speech large language models (SpeechLLMs) offer reduced latency and retain paralinguistic nuances that are typically lost in cascaded automatic speech recognition (ASR) and text-based LM architectures. However, they continue to lag behind t…
- SIRF: A Spec-Internalized Risk Foundation Model for Industrial Content Risk ControlSuwan Wu, Yumeng Lin, Pengcheng Yuan, Xiaolong Jiang · arXiv · Sep 10, 2026
For industrial content risk control, the real deployment constraint is not average accuracy but how much risk can be auto-handled under high precision and second-level latency. We present SIRF (Spec-Internalized Risk Foundation Model), whic…
- Negative Self-Distillation: Learning to Reason by Avoiding FlawsRongcan Pei, Zhepei Wei, Shuyao Xu, Xinyu Zhu et al. · arXiv · Sep 10, 2026
On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However,…
- A Voice-Interactive Multi-Agent System for Smart Operating Rooms: Architecture Design and Key TechnologiesTianxiang Zhou · arXiv · Sep 10, 2026
This paper presents SurgicalRoomAgent, a voice-interactive multi-agent system for smart operating rooms based on large language models (LLMs). The system achieves natural language understanding, device control, intraoperative recording, and…
- Can LLMs Normalize Databases? A Benchmark and Multi-Agent Framework for Schema NormalizationDong-Jae Koh, Huisu Kim, SeongHwan Yoon, Lasse M. Jantsch et al. · arXiv · Sep 10, 2026
Large Language Models (LLMs) are increasingly used to generate structured outputs, but their reliability remains unclear when those outputs must satisfy database-level constraints. We study this issue through database normalization, involvi…
- Beyond Solver Verdicts: Generative Reward Models for AutoformalizationVikash Singh, Debargha Ganguly, Aman Goel, Ali Torkamani et al. · arXiv · Sep 10, 2026
Neurosymbolic systems rely on mathematical solvers to guarantee reasoning correctness, yet solvers are fundamentally blind to whether a formal translation maintains strict reference-equivalence to a designated formalization. We formalize th…
- BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question GenerationKarish Gupta, Matthew Alex, Alex Li, Yang Wu et al. · arXiv · Sep 9, 2026
Police body-worn camera (BWC) footage has emerged as a critical aspect of law enforcement that ensures legal transparency, officer accountability, and the protection of civil rights. However, effectively processing this data remains a signi…
- Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity LinkingParinthapat Pengpun, Simran Khanuja, Graham Neubig · arXiv · Sep 9, 2026
Multimodal entity linking grounds entity mentions in text and images to knowledge-base entries. These systems degrade on rare entities, but prior work measures rarity primarily through popularity-based metrics such as pageviews. We broaden …
- Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language ReasoningMehrnaz Mofakhami, Ananya Sahu, Alejandro R. Salamanca, Daniel D'souza et al. · arXiv · Sep 9, 2026
Reasoning language models have made substantial advances on a variety of complex tasks, yet their capabilities remain overwhelmingly English-centric: models primarily reason in English regardless of the language they are prompted in. This i…
- ConvMem: Convolutional Memory for Long-Context ReasoningHongming Zhang, Zhaozhen Gu, Fengshuo Bai, Ming Hao et al. · arXiv · Sep 9, 2026
While Large Language Models (LLMs) have demonstrated impressive capabilities, they often struggle with extremely long contexts due to fixed context limits. To address this, sequential approaches like MemAgent extend the effective context by…
- From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric ReasoningWeichen Dai, Rafael Medeiros Cabral, Ziyi Shou, Yan Cao et al. · arXiv · Sep 9, 2026
Plane geometry remains a significant challenge in AI, requiring the integration of visual perception and mathematical reasoning. While Large Multimodal Models (LMMs) naturally handle visuo-linguistic inputs, they are often computationally i…
- LiteRAG: Cost-Efficient Graph-Based Retrieval-Augmented GenerationDaniel Alejandro Coll Tejeda, Pedro García López, Daniel Barcelona-Pons · arXiv · Sep 9, 2026
Graph-based retrieval can improve multi-hop question answering, but existing approaches often incur high query-time costs and produce diffuse, oversized contexts that reduce generation efficiency. We present LiteRAG, a graph-based retrieval…
- $Φ$-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?Leilei Ding, Shumin Wang, Yuting Huang, Fanqi Wan et al. · arXiv · Sep 9, 2026
Large language models (LLMs) have demonstrated remarkable capabilities in reasoning and code generation, raising the prospect that they could assist in developing and optimizing the very infrastructure that powers them. However, existing be…
- Data-Centric Post-Training for Financial Reasoning: Mining, Distillation, and Verifiable LearningZhirayr Hayrapetyan, Andrei Kalmykov, Denis Kokosinskii, Dmitry Stanishevskii et al. · arXiv · Sep 9, 2026
Financial text, textbooks, and question-answer pairs are abundant, but only a small fraction is directly usable for reasoning-focused post-training. Existing QA pairs often lack explicit reasoning, sufficient context, or reliably verifiable…
- Stable Answers, Unfinished Reasoning: Why Self-Consensus Is Not a Safe Early-Exit SignalYunxiang Mo, Donghao Zhao, Hejia Geng · arXiv · Sep 9, 2026
A natural way to cut reasoning-model inference cost is to repeatedly probe a single partial trajectory for its current answer and stop once probes agree -- self-consensus. We ask whether any such rule is both safe and token-saving, and whet…
- VLX-VR: An Agentic-Aware Video Reasoning ModelSheng Li, Peng Liu, Qianqian Zhang, Tiancheng Zhao · arXiv · Sep 9, 2026
Real-world video understanding requires integrating visual, audio, textual, and temporal evidence distributed across a video. Yet many pipelines use a fixed video context and single-pass inference, limiting adaptive evidence acquisition whe…
- LogiScope-VQA: Benchmarking Vision-Language Models for Logistics Hazard Identification in Industrial ScenariosHanjing Zhou, Mingze Yin, Ying Lian, Jun Ma et al. · arXiv · Sep 9, 2026
Large Multimodal Models (LMMs) large-scale deployment in industrial warehouse settings specifically necessitates that models exhibit human-expert-level hazard-oriented perception, understanding, and reasoning capabilities. However, the scar…
- CARRE: Counterfactual Action Retrieval and Reason Evaluation for Explainable Churn PrescriptionMinJoo Kim, SanJin Park, SeungHwan Cho · arXiv · Sep 9, 2026
Churn models typically identify high-risk customers but do not specify which feasible retention action should be considered or why that action is appropriate. We present CARRE (Counterfactual Action Retrieval and Reason Evaluation), a three…
- What Does MMLU Actually Measure? A Psychometric Audit of Difficulty Structure in Aggregate Benchmark ScoresDana Paquin, Riddhiman Jain · arXiv · Sep 8, 2026
Although MMLU is widely adopted as a benchmark for calibrating general AI capabilities, we psychometrically demonstrate that its aggregate score primarily evaluates a model's factual retrieval capacity rather than its reasoning ability. By …
- ReCite: Agentic Reasoning for Faithful CitationYuyang Huang, Bobo Li, Jiajia Song, Yuzhe Ding et al. · arXiv · Sep 8, 2026
Accurate citations are the foundation of academic writing, tracing intellectual origins and substantiating core claims. However, manually navigating the growing volume of scientific literature is increasingly difficult, prompting reliance o…
- Answer-Distribution Trajectories: A Stochastic-Dynamics View of LLM ReasoningMar Gonzàlez I Català, Haitz Sáez de Ocáriz Borde, Davide Murari, Carola-Bibiane Schönlieb et al. · arXiv · Sep 8, 2026
Chain-of-thought reasoning provides a structured computation between a model's input and final answer. Yet it is often evaluated through endpoint accuracy, which ignores the path taken to reach that answer. An emerging line of work addresse…
- Step-level screening-driven reasoning enhancement for large language models in prefabricated substation fault diagnosisGengsheng Zhang, Shuai Zhang, Jin Zhu, Chunhou Zheng · Electric Power Systems Rese... · Sep 7, 2026
- Fine-tuning small reasoning models for quantum field theoryNathaniel Woodward, Zhiqi Gao, Yurii Kvasiuk, Kendrick M. Smith et al. · Machine Learning Science an... · Sep 7, 2026
Abstract Despite the growing application of large language models (LLMs) to theoretical physics, there has been little academic exploration of how domain-specific physics reasoning ability develops during training. To investigate this, we p…
- Thinking Without Words: A Survey of Latent Chain-of-Thought Reasoning in Large Language ModelsLihui Liu · OpenAlex · Sep 6, 2026
- WearableQA: A Benchmark for Health Reasoning over Real-World Wearable DataJi Soo Lee, Xilun Chen, Pierce Chuang, Ashish Shenoy et al. · arXiv · Sep 4, 2026
Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over a real user's longitudinal wearable record. We introduce We…
- A Verifier-Guided Explainable Reasoning Framework with Gold-Anchored QLoRA, Task-Aware Mixture-of-Experts, and Group-Relative RLVRThi Kim Trang Vo, Nam Tien Le, Thi Kim Nguyet Vo, Minh Khang Tran et al. · arXiv · Sep 4, 2026
Large language models (LLMs) show strong reasoning ability, but their explanations can remain inconsistent, weakly grounded, or difficult to verify. We propose a verifier-guided explainable reasoning framework for transparent educational qu…
- Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?Daan R. Henselmans, Derck W. E. Prinzhorn, Arno Libert · arXiv · Sep 4, 2026
AI oversight methods rely on ground truth for validation, but what constitutes appropriate AI behavior is contested. This leaves evaluation of moral reasoning in LLMs and debate-based oversight implicitly avoiding realistic ambiguity. We in…