Latest Interpretability Research Papers
The newest Interpretability papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Interpretability so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Interpretability papers in your inbox — free →Recent papers
- Post-Interpretive Criticism: Volume III - The Canon of WitnessesVale, Dorian, Museum of One · Zenodo (CERN European Organ... · Apr 28, 2028
Title: The Canon of Witnesses Type: Publication > Journal Volume Journal Title: The Journal of Post-Interpretive Criticism Volume: 3 (upcoming) ISSN: 2819-7232 Author: Dorian Vale Publisher: Museum of One DOI: 10.5281/zenodo.17421408 Museum…
- Post-Interpretive Criticism: Volume III - The Canon of WitnessesVale, Dorian, Museum of One · Zenodo (CERN European Organ... · Apr 28, 2028
Title: The Canon of Witnesses Type: Publication > Journal Volume Journal Title: The Journal of Post-Interpretive Criticism Volume: 3 (upcoming) ISSN: 2819-7232 Author: Dorian Vale Publisher: Museum of One DOI: 10.5281/zenodo.17421408 Museum…
- Post-Interpretive Criticism: Volume III - The Canon of WitnessesVale, Dorian, Museum of One · Zenodo (CERN European Organ... · Apr 28, 2028
Title: The Canon of Witnesses Type: Publication > Journal Volume Journal Title: The Journal of Post-Interpretive Criticism Volume: 3 (upcoming) ISSN: 2819-7232 Author: Dorian Vale Publisher: Museum of One DOI: 10.5281/zenodo.17421408 Museum…
- On Zaya (2025) – Marcel van LuitVale, Dorian · OpenAlex · Apr 1, 2028
Abstract The Medium Betrayed the Miracle is a Post-Interpretive Criticism essay examining the encounter between image, memory, and medium through a detailed reading of Zaya’s digitally rendered allegorical tableau. What first appears to be …
- On Zaya (2025) – Marcel van LuitDorian Vale · Zenodo (CERN European Organ... · Apr 1, 2028
Abstract The Medium Betrayed the Miracle is a Post-Interpretive Criticism essay examining the encounter between image, memory, and medium through a detailed reading of Zaya’s digitally rendered allegorical tableau. What first appears to be …
- A Qualitative Exploration of the Attitudes and Perceptions of Female Immigrants of Chinese Descent Regarding Access to Mental Health ServicesMing Wai Myra Kan · USF Scholarship Repository ... · Aug 3, 2027
This qualitative study used Interpretative Phenomenological Analysis to explore the lived experiences and meaning-making of Chinese immigrants’ attitudes and perceptions of mental disorders and mental health-seeking behaviors in the US. Thi…
- Deciphering the threat: informational and analytical foundations for police training on encrypted criminal networks. The cases of Ukraine, England, and GermanyBohdan-Petro Koshovyi, Andrii Manko, P. P. Latkovskyi, Aleksey Demchenko · Zenodo (CERN European Organ... · Jan 1, 2027
This article interprets, from a hermeneutic and documentary perspective, the informational and analytical foundations that recent academic literature identifies as necessary for the training of police forces in dealing with encrypted crimin…
- Deciphering the threat: informational and analytical foundations for police training on encrypted criminal networks. The cases of Ukraine, England, and GermanyBohdan-Petro Koshovyi, Andrii Manko, P. P. Latkovskyi, Aleksey Demchenko · Zenodo (CERN European Organ... · Jan 1, 2027
This article interprets, from a hermeneutic and documentary perspective, the informational and analytical foundations that recent academic literature identifies as necessary for the training of police forces in dealing with encrypted crimin…
- HF-CAM: A Hard Prototype Anchored Concept Activation Mapping for Interpretable Deep LearningYuan Liu · Open MIND · Dec 31, 2026
The black-box nature of deep neural networks poses a significant challenge to their deployment in high-stakes decision-making scenarios, conflicting with the principles of transparency and accountability in Responsible AI. To address this, …
- Borrowed Shame: Hao Hongmei as a Narrative Resource for Sun Shaoping's Moral ImageKhoo-Hong Chang · Knowledge Commons (Lakehead... · Dec 31, 2026
This article reconsiders one of the most frequently moralized episodes in Lu Yao's Ordinary World: Hao Hongmei's theft of handkerchiefs at the end of high school and Sun Shaoping's intervention on her behalf. Existing readings often treat t…
- Explainability Assistant: A Conversational XAI Interface for Interpreting Energy Consumption ModelsRodion Krjutškov, Eduard Barbu, Nikos Sakkas, Sofia Yfanti · arXiv · Sep 10, 2026
Energy consumption forecasting relies on increasingly complex machine learning (ML) models, such as Genetic Programming-based symbolic regressors, whose predictions can be difficult for facility managers and building operators to interpret.…
- MUtE: A Dual Framework for Concept Erasure and Counterfactual InterventionsAntoine Saillenfest · arXiv · Sep 10, 2026
Erasing concept-specific information from representations has been proven useful for mitigating bias or interpreting model decisions. The joint objective is to transform the original representations such that the target concept becomes unpr…
- TailProp: content-adaptive light- and heavy-tailed propagation for visionJiahao Kong, Zihan Li · arXiv · Sep 10, 2026
Science-inspired vision models show that explicit propagation dynamics can provide structured and interpretable alternatives to conventional token mixing. Existing formulations, however, typically construct and adapt visual propagation with…
- scDEFT: A deep learning framework for drug-effect prediction and counterfactual reasoningMurthy Devarakonda · arXiv · Sep 9, 2026
Longitudinal single cell atlases now capture matched pre treatment and post treatment states from responders and non responders, presenting an opportunity to mechanistically explain why two patients on the same drug diverge. We introduce sc…
- A Multi-Stage Rule-Chaining Framework for Compositional and Interpretable Cognitive ReasoningDeblina Kar · arXiv · Sep 9, 2026
The Abstraction and Reasoning Corpus (ARC) benchmarks cognitive generalization, the ability to infer and apply abstract rules from limited examples. This paper presents a multi-stage rule-chaining framework that performs compositional reaso…
- A Dominant Diffuse Phase in the Sparse Autoencoder Phase DiagramAlexis D. Plascencia · arXiv · Sep 9, 2026
Sparse autoencoders (SAEs) are increasingly used to recover interpretable features from neural-network activations, yet systematic feature co-occurrence can cause distinct features to be absorbed or merged. The MAIS-O43 open problem propose…
- Artificial Intelligence Algorithms for the Detection of Pathologies Related to Lung Cancer through Image Analysis using Convolutional Neural Networks and Data Augmentation: a systematic mapping of the literaturePablo Ramirez Amador · arXiv · Sep 9, 2026
Lung cancer is one of the leading causes of death worldwide, and its early diagnosis is crucial to improving patients prognosis and quality of life. However, the process of interpreting medical images for the detection of lung cancer is com…
- SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?Yuqiao Tan, Shizhu He, Jun Zhao, Kang Liu · arXiv · Sep 8, 2026
While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autonomous development demands a missing pillar: post-hoc monitoring and auditing to understand what models learn and ensure s…
- A Verifier-Guided Explainable Reasoning Framework with Gold-Anchored QLoRA, Task-Aware Mixture-of-Experts, and Group-Relative RLVRThi Kim Trang Vo, Nam Tien Le, Thi Kim Nguyet Vo, Minh Khang Tran et al. · arXiv · Sep 4, 2026
Large language models (LLMs) show strong reasoning ability, but their explanations can remain inconsistent, weakly grounded, or difficult to verify. We propose a verifier-guided explainable reasoning framework for transparent educational qu…
- Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought ReasoningKevin Du, Alexander Hoyle, Laura Ruis, Acyr Locatelli · arXiv · Sep 3, 2026
Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its answer. A growing body of work treats them as such, using LLM judges to diagnose errors, evaluate faithfulness, and provide step-…
- The Implications of Linguistic Illegibility for LLM SecurityJames Mickens · arXiv · Sep 2, 2026
LLMs are trained to generate natural language. However, various strands of evidence indicate that an LLM's externalized linguistic outputs and mechanistically-extracted linguistic features can be an unreliable lens for understanding interna…
- LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style UpdatesDmitrii Andriianov, Andrey Veprikov, Aleksandr Beznosikov · arXiv · Sep 2, 2026
Low-rank adaptation (LoRA) is the standard way to fine-tune large models, yet when its two factors are trained independently, the update ignores the geometry of the low-rank weight change it induces. We introduce LoRA-TSD, an optimizer that…
- Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization EvaluationHimil Vasava, Ming Jiang · arXiv · Sep 1, 2026
LLM-based evaluators of natural language generation (NLG) quality are widely deployed as scoring tools and as automated training signals, yet the internal procedure by which they assign a rating remains poorly understood. We investigate thi…
- Advancing Interaction-Sensitive Feature Selection: Novel Relief-Based Algorithms, Expanded Comparisons, and Recommendations for Biomedical Data MiningKia Kazemi-Nia, Harsh Bandhey, Philip J. Freda, Ryan J. Urbanowicz · arXiv · Aug 28, 2026
As a precursor to high-dimensional biomedical data modeling, reliable feature selection can reduce computational expense, improve modeling performance, and yield simpler, more interpretable models. However, most filter-based feature selecti…
- Acquire, Repair, Preserve: A Diagnosis-Guided Post-Training Recipe for Small-Model Dialogue Game AgentsNan Li · arXiv · Aug 28, 2026
Interactive dialogue games test a capability that static benchmarks largely leave implicit: a model must carry state across turns, interpret feedback, and choose valid actions under changing constraints. We study this setting in the LM Play…
- QuantumBoostNet: A Hybrid Classical-Quantum Architecture for Enhanced Accuracy in Cardiac Ultrasound View IdentificationMihai Udrescu-Milosav, Stefan-Alexandru Jura, Mihai Udrescu, Gerhard-Paul Diller · arXiv · Aug 27, 2026
Accurate identification of the correct view or angle in cardiac ultrasound (echocardiogram) is a critical component of cardiologic imaging. This step is essential for precise anatomical interpretation, reliable measurement, and the reductio…
- Circuit Condensation: Post-Training that Concentrates a Behavior's Causal CircuitSai Adith Senthil Kumar · arXiv · Aug 27, 2026
One approach to mechanistic interpretability explains behavior through circuits: the components and connections that carry it. Frozen discovery often returns hundreds of edges, making them hard to inspect, compare, or verify exhaustively. W…
- Importance Scoring of Transformer Attention Heads in Learning Tabular DataAhmad Jad Allah, Kazi F. Akhter, Md. Kamrozzaman Bhuiyan, Manar D. Samad · arXiv · Aug 27, 2026
Computationally demanding and opaque deep learning models can be better understood and optimized by analyzing how they transform data. While deep transformers have been widely studied in computer vision and natural language processing, thei…
- Weakly Supervised Seafloor Segmentation for Seagrass Habitat Mapping in Side-Scan Sonar ImageryHayat Rajani, Nuno Gracias, Rafael Garcia · arXiv · Aug 25, 2026
Seagrass meadows are crucial blue-carbon habitats, and mapping their extent is a prerequisite for coastal management and carbon inventory. Optical satellite sensors cover large areas but cannot reach deep or turbid water, whereas side-scan …
- Truthful Calibration Measures for Sequential PredictionAnagha Gokul, Jason Hartline, Lunjia Hu, Jonathan Ullman et al. · arXiv · Aug 21, 2026
Calibration requires probabilistic reports to be conditionally unbiased and reliably interpretable as probabilities. A calibration measure assigns numerical error to miscalibrated reports. Haghtalab et al. (2024) proposed an approximately t…