Latest Question Answering Research Papers
The newest Question Answering papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Question Answering so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Question Answering papers in your inbox — free →Recent papers
- Moral Consistency Variance: A Pilot Benchmark for Decision Stability under Moral Prompt Perturbations in Large Language ModelsQiao Liang · OpenAlex · Dec 31, 2026
Static moral question-answering benchmarks do not test whether a model's decision distribution remains stable when the same dilemma is rephrased without changing the underlying facts. This paper introduces Moral Consistency Variance (MCV), …
- Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMsKillian Steunou, Yannis Tevissen, Mounîm A. El Yacoubi · arXiv · Sep 9, 2026
Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their strong performance o…
- Which Medical Questions Deserve Rationales? Perturbation-Sensitive Selection for Robust QAYuexin Wu, Dayou Yu, Vasile Rus · arXiv · Sep 9, 2026
Medical question-answering datasets often contain answer labels, whereas high-quality rationales remain scarce, noisy, or costly to validate. This changes the acquisition question: rather than asking which questions should be labeled, we as…
- When Models Defer to Wrong Answers: A Robustness Audit of Source-Attributed Cues in Multiple-Choice QAManikandan Ravikiran, Siddharth Vohra · arXiv · Sep 8, 2026
Language models often receive a question together with a claim about what another source answered. We audit whether such claims destabilize answers in multiple-choice question answering. For each item, we hold one wrong option fixed across …
- Evolution of Multimodal Question Answering: From Modality-Adaptive Extraction to Unified Language RepresentationAbdullah Al Shafi · arXiv · Sep 8, 2026
The rapid growth of multimodal data has intensified the need for question answering (QA) systems capable of reasoning across heterogeneous sources such as text, tables, and images. In this paper, we present a comprehensive methodological co…
- From Coordinates to Candidate Regions: Temporal Change Localization via Region Selection in Remote Sensing Multimodal LLMsJuwan Chung, Sungjune Park, Yeongyun Kim, Yong Man Ro · arXiv · Sep 8, 2026
Remote sensing multimodal large language models (RS-MLLMs) have advanced scene understanding and visual question answering over satellite imagery, yet localizing specific objects or changed regions remains challenging. Existing approaches r…
- SentryLine: Evidence-Grounded Question Answering over Evolving Documents in Oncology CareTampu Ravi Kumar, Gaurav Najpande, Muhammad Ali Khan, Kaneez Zahra Rubab Khakwani et al. · arXiv · Sep 8, 2026
Oncology care operates at constant pressure of absorbing rapidly evolving evidence base in biomedicine. The American Society of Clinical Oncology (ASCO) addresses this through living guidelines, but the format introduces a new burden: any r…
- IGT @ FinMMEval 2026 Task 2: Question-Type Prompting with Targeted Extraction for Multilingual Financial QAYuwen Chiu · arXiv · Sep 8, 2026
We present the IGT system for PolyFiQA Task 2 of the FinMMEval Lab at CLEF 2026, a multilingual financial question answering task over English SEC filings and multilingual news articles (English, Chinese, Japanese, Spanish, Greek) for four …
- WearableQA: A Benchmark for Health Reasoning over Real-World Wearable DataJi Soo Lee, Xilun Chen, Pierce Chuang, Ashish Shenoy et al. · arXiv · Sep 4, 2026
Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over a real user's longitudinal wearable record. We introduce We…
- A Verifier-Guided Explainable Reasoning Framework with Gold-Anchored QLoRA, Task-Aware Mixture-of-Experts, and Group-Relative RLVRThi Kim Trang Vo, Nam Tien Le, Thi Kim Nguyet Vo, Minh Khang Tran et al. · arXiv · Sep 4, 2026
Large language models (LLMs) show strong reasoning ability, but their explanations can remain inconsistent, weakly grounded, or difficult to verify. We propose a verifier-guided explainable reasoning framework for transparent educational qu…
- A Structured Debate-Mixture-of-Agents Framework for Complex Clinical Diagnostic Decision SupportChang Xia, Leilei Ouyang, Huimin Wang, Yong Zhao et al. · arXiv · Sep 4, 2026
Large language models (LLMs) show potential for medical tasks, but their single-turn question-answer format does not reflect how clinical diagnosis is performed in practice. As a result, they remain limited in complex diagnostic settings. W…
- BIT.UA at BioASQ 14B: Modular Retrieval with pg_textsearch and Qdrant, and Agent-Based Answer GenerationAndré Ribeiro, Rúben Garrido, Alexander Christiansen, Richard A. A. Jonker et al. · arXiv · Sep 4, 2026
This paper describes the participation of the BIT.UA team from the University of Aveiro in the 14th edition of the BioASQ Task B challenge on biomedical question answering. Building on our previous submissions, we introduced a substantially…
- MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical DomainSourav Malakar, Harshit Nigam, Akash Ghosh, Sriparna Saha et al. · arXiv · Sep 4, 2026
Time-series data in clinical settings is crucial for capturing dynamic changes in a patient's health over time, enabling timely diagnosis, personalized treatment, and early detection of critical events. However, the development of clinicall…
- PetQA: Benchmarking Veterinary Knowledge and Clinical ReasoningTaegyun Kim, Youngwook Ham, Jungwook Rhim, Ju-Hyun An et al. · arXiv · Sep 4, 2026
We introduce PetQA, a Korean long-form question-answering (QA) benchmark for evaluating veterinary knowledge and clinical reasoning in large language models (LLMs) and large vision-language models (LVLMs). PetQA contains 10,076 text-only an…
- Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model PipelinesSiddharth Vohra, Runmin Jiang, Xiaomo Li, Min Xu · arXiv · Sep 4, 2026
Grounded language-model pipelines can be divided into three stages: selecting an object, retrieving passages for it, and using that evidence to answer. If the selected object must reach the reader, losing it breaks the handoff. Benchmark re…
- RuleMem: Active Rule Memory for Long-Term Conversational AgentsXingyuan Zeng, Zuohan Wu, Quanming Yao, Yue Wang et al. · arXiv · Sep 3, 2026
Question answering agents in long-term conversations must reason over massive, temporally dispersed dialogue histories. However, existing memory mechanisms primarily treat past information as \textit{passively} stored facts, leading to sema…
- Enhancing Financial Question Answering: A Novel Benchmark Dataset of Banks' financial statementsArianna Miola, Bruno Spaccavento, Lorenzo Silotto, Marco Bianchetti et al. · arXiv · Sep 3, 2026
The comparative analysis of banks' financial statements poses significant challenges for automated question answering systems due to their complexity, substantial length, technical language, and inhomogeneity of both textual and numerical c…
- When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational AgentsWen-Yu Chang, Yun-Nung Chen · arXiv · Sep 3, 2026
Large language models (LLMs) are increas- ingly deployed as long-horizon conversational agents, motivating growing interest in mem- ory systems. However, existing benchmarks primarily evaluate memory through QA-style probing rather than in-…
- When Retrieval Helps: Selective Retrieval for Single-Turn Mental-Health QAHyunseo Oh, Chong-Kwon Kim, Yoonhyuk Choi · arXiv · Sep 3, 2026
Retrieval-augmented generation (RAG) can improve the specificity and grounding of large language model responses, but its effect is not uniformly beneficial in single-turn mental-health question answering, where user queries often combine e…
- TabScope: Question-Adaptive Scope Selection for Table Question AnsweringYuxiang Wang, Junhao Gan, Jianzhong Qi · arXiv · Sep 3, 2026
Large Language Models (LLMs) have shown strong performance on table question answering, yet their accuracy often degrades as table size increases. We find that this degradation is not uniform across question types. Localization-sensitive qu…
- MedQA-MM: Shortcuts Behind Medical Visual ReasoningBenlu Wang, Yifan Zhang, Jiaqing Yu, Chin Siang Ong et al. · arXiv · Sep 3, 2026
A benchmark score credits final answers, but not the route by which an item can be answered. In medical multimodal multiple-choice questions (MCQs), this distinction matters because a correct answer can be supported by the intended image fi…
- Untangling the Mechanisms of Misleading Context in Medical Question AnsweringRobin Linzmayer, Noémie Elhadad · arXiv · Sep 2, 2026
Large language models now answer medical questions with expert-level performance. However, the context these systems act on can be misleading, and misleading context can corrupt a model's medical judgment. To understand how misleading conte…
- ViSAR: Training-Free Adaptive-$k$ Retrieval for Visual Document Question AnsweringAdrien Mialland, Marc Plantevit, Julien Gallois, Céline Robardet · arXiv · Sep 2, 2026
Document Visual Question Answering (DocVQA) often leverages Retrieval-Augmented Generation (RAG), where late-interaction encoders are commonly used to identify document pages relevant to a user query, before answer generation by a Large Vis…
- APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question AnsweringJie Ding, Rui Sun, Xinyuan Zhang, Zeyu Zhang et al. · arXiv · Sep 2, 2026
Deep research agents augment large language models with external tools to answer complex, long-horizon questions through multi-turn reasoning. Learning from prior experience is crucial for continual improvement, yet existing methods either …
- Grounded, Compute-Efficient LLM Policy Agents for Energy-Poverty Equity in Physically-Constrained Peer-to-Peer Energy MarketsKunal Jadhav, Siddhesh More · arXiv · Sep 1, 2026
Energy poverty is nearly absent from NLP-for-social-good, and the little existing work is either static retrieval/QA or relies on carbon-intensive cloud LLMs, a self-defeating "computational irony" for a humanitarian setting. We present EqG…
- InSight: A Benchmark for Agentic Claim Verification in Interactive VisualizationsMaeve Hutchinson, Syed Mahbubul Huq, Mohammad Albinhassan, Radu Jianu et al. · arXiv · Sep 1, 2026
Vision Language Models have demonstrated remarkable proficiency in interpreting static visual artifacts, but modern data analysis is inherently dynamic, requiring the active interrogation of interactive environments. Existing benchmarks are…
- How Correct Is Your Answer? A Semantic Correctness Framework for Open QA EvaluationElitsa Yotkova, Violeta Kastreva, Petar Velkov, Hristo Boyanov et al. · arXiv · Sep 1, 2026
Reliable evaluation of open-ended question answering remains a bottleneck for measuring answer correctness of modern LLMs. Unlike multiple-choice tasks, free-form answers may be correct in many surface forms and may fail in qualitatively di…
- Investigating Linear Probe Robustness to Linguistic Register, Medical Specialty, and Corpus Shifts in Medical QANishant Mishra, Ameen Abu-Hanna, Iacer Calixto · arXiv · Sep 1, 2026
Linear classifiers trained on hidden states of a large language model (LLM), linear probes, can flag factual errors from a single forward pass. Geometrically, that implies that true and false statements separate along a stable direction in …
- FinLifeBench: Exhaustive Life-Event History and Financial-State Reconstruction from Longitudinal Banking DialogueHangyeul Lee, Juyoung Oh, Jaeyong Ko, Sunmin Kim et al. · arXiv · Sep 1, 2026
Repeated banking interactions require assistants to maintain complete, current, and traceable customer records as life changes emerge incidentally in routine requests. Existing benchmarks emphasize question answering, bounded episodes, or t…
- EDRAC: Benchmarking Arabic Dialect Reading ComprehensionNoor Abo Mokh, Kirill Chirkunov, Teresa Lynn, Nizar Habash et al. · arXiv · Sep 1, 2026
Dialectal Arabic (DA) remains under-resourced compared to Modern Standard Arabic (MSA), particularly for machine reading comprehension (MRC) and question answering (QA). Existing Arabic QA benchmarks primarily focus on formal written MSA or…