Latest Interpretability Research Papers
The newest Interpretability papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Interpretability so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Interpretability papers in your inbox — free →Recent papers
- On Zaya (2025) – Marcel van LuitVale, Dorian · Zenodo (CERN European Organ... · Apr 1, 2028
Abstract The Medium Betrayed the Miracle is a Post-Interpretive Criticism essay examining the encounter between image, memory, and medium through a detailed reading of Zaya’s digitally rendered allegorical tableau. What first appears to be …
- On Zaya (2025) – Marcel van LuitVale, Dorian · Zenodo (CERN European Organ... · Apr 1, 2028
Abstract The Medium Betrayed the Miracle is a Post-Interpretive Criticism essay examining the encounter between image, memory, and medium through a detailed reading of Zaya’s digitally rendered allegorical tableau. What first appears to be …
- Advanced optimization methods for interpretable machine learning models and their applications in chemical engineeringJiayang Ren · cIRcle (University of Briti... · Jan 1, 2028
The full abstract for this thesis is available in the body of the thesis, and will be available when the embargo expires....
- Climate-resilient electric vehicle charging infrastructure for sustainable cities: An interpretable causal-ensemble framework for preventive maintenance and low-carbon mobilityCande Lian, Wentao Zeng, Jiabin Wu, Yiming Bie et al. · arXiv · Jul 23, 2026
Reliable electric vehicle (EV) charging infrastructure is a cornerstone of sustainable, low-carbon cities, yet urban climate stress such as extreme heat, heavy precipitation, and humidity increasingly raises equipment fault risk and undermi…
- Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought ModelsRenuka Oladri, Niveda Jawahar, Abdirisak Mohamed · arXiv · Jul 23, 2026
Chain-of-thought reasoning models such as DeepSeek-R1-Distill-Qwen-7B exhibit a bimodal convergence pattern: generations either terminate within a token budget (converged) or exhaust it without reaching a conclusion (non-converged). We char…
- PG-KINN: A Physics-Informed Petrov-Galerkin Kolmogorov-Arnold Network for Solving Forward and Inverse PDEsAmirhossein Sadr, Nima Soltani, Vahideh Moghtadaiee, Aida Pakniyat et al. · arXiv · Jul 22, 2026
Physics-informed learning of partial differential equations (PDEs) has been dominated by multilayer perceptrons (MLPs), whose spectral bias and dense parameterization limit both accuracy and interpretability. Kolmogorov Arnold Networks (KAN…
- Interpretable Fuzzy Rule-Based Regression Extension for Ex-Fuzzy LibraryCayan Deniz Kucuktopana, Javier Fumanal-Idocin, Richard Pitts, Javier Andreu-Perez · FUZZ-IEEE · Jul 22, 2026
Machine learning models achieve high predictive accuracy in regression tasks, but their deployment in safety-critical and regulated domains requires interpretability. While fuzzy rule-based systems offer transparent, linguistically explicit…
- PhaseAware: Interpretable Human-in-the-Loop Rehabilitation Scoring with Boundary MonitoringYankai Zheng, Yuhe Liu, Yuxin Ma, Tianci Xue et al. · arXiv · Jul 22, 2026
Rehabilitation scoring systems are most useful when their outputs can be reviewed and interpreted within clinical workflows. This study presents PhaseAware, a compact framework for continuous rehabilitation quality assessment that combines …
- User-Centric Modeling of Transactional Sequences with Explainable State Space ModelsIvan Palagin · arXiv · Jul 22, 2026
We propose a hybrid approach for user-centric modeling of transactional event sequences that combines contrastive representation learning (CoLES) with State Space Models (SSMs). While contrastive methods yield high-quality compressed user r…
- CircuitKIT : Circuit Discovery, Evaluation, and Application Toolkit for Mechanistic InterpretabilityPratinav Seth, Hem Gosalia, Aditya Kasliwal, Vinay Kumar Sankarapu · arXiv · Jul 21, 2026
Circuit analysis can support not only model explanation but also downstream interventions such as pruning, editing, steering, and selective fine-tuning. However, conducting such analyses currently requires stitching together separate implem…
- Toward Auditable Fraud Detection: Combining Graph Features, Model Explanations, and Agentic Case InvestigationRahil Sharma · arXiv · Jul 21, 2026
Fraud detection systems must scale with rising transaction volume while remaining explainable and reviewable. We study a layered pipeline on the PaySim dataset that combines a gradient-boosted classifier, graph-derived structural features, …
- BadWAM: When World-Action Models Dream Right but Act WrongQi Li, Xingyi Yang, Xinchao Wang · arXiv · Jul 16, 2026
World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often view…
- An Introduction to Sparse Identification of Nonlinear Dynamics for Engineering ApplicationsYao Cheng Li, Ana Larrañaga, Steven L. Brunton, Urban Fasel · arXiv · Jul 16, 2026
Many engineering problems involve phenomena whose governing equations are poorly characterized or only partially known. Surrogate modeling techniques such as neural networks can capture the behavior of these systems, but they typically dema…
- Robustness of Deep Learning Models for PV Power Forecasting under NWP Forecast Errors: A Spatiotemporal and Physically Interpretable AnalysisDandan Chen, Yan Zhao, Xuepeng Chen · arXiv · Jul 14, 2026
Engineering use of AI forecasting models requires not only high nominal accuracy but also predictable behavior under uncertain inputs. In photovoltaic (PV) forecasting, this requirement is especially challenging because numerical weather pr…
- Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge BiasZixiang Xu, Sixian Li, Huaxing Liu, Xiang Wang et al. · arXiv · Jul 13, 2026
Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and propose prompt-level mitigations. We argue that the same biases admit a representation-level account …
- All Explanations are Wrong, But Many Are Useful: Exploring the Rashomon Explanation Set with Large Language ModelsPan Li · arXiv · Jul 10, 2026
Explaining machine-learning models is increasingly important for decision-making and consumer trust, yet it is widely believed to come at a cost: existing Explainable AI (XAI) methods suffer from a persistent accuracy-explainability trade-o…
- Steering Neural Network Training through Interpretable Constraints Based on Partial DependenceYann Claes, Pierre Geurts, Vân Anh Huynh-Thu · arXiv · Jul 9, 2026
Over the last few years, there has been an increased interest in making machine learning models more interpretable. Although a great deal of effort goes into developing techniques for interpreting the interactions learned by a given model, …
- When Structured Sparse Autoencoders Learn Consistent Concepts Across ModalitiesWeiduo Liao, Yunqiao Yang, Ying Wei · arXiv · Jul 9, 2026
Sparse autoencoders (SAEs) have emerged as a promising technique for mechanistic interpretability by learning a set of sparse latent features in large models, each of which encodes a distinct concept. However, in vision-language models (VLM…
- Higher-Order Geometric Updates for Levenberg-Marquardt Method via Riemann Normal CoordinatesJianing Liu, Dong H. Zhang · arXiv · Jul 8, 2026
Nonlinear least-squares optimization is central to regression, physics-informed neural networks, and other machine-learning tasks. Such problems have a natural geometric interpretation, model predictions form a manifold in data space, while…
- Dithered Gaussian Mechanism for Randomness-Efficient Differential PrivacyNikita P. Kalinin, Rasmus Pagh · arXiv · Jul 7, 2026
We present the dithered Gaussian mechanism, a novel alternative to the discrete Gaussian mechanism for differential privacy that discretizes the private output rather than the noise distribution itself. By interpreting this discretization a…
- TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific ScenariosHong Lyu, Mingru Yang, Qianhua He, Yanxiong Li et al. · arXiv · Jul 7, 2026
There are some datasets of varying scales for audio classification (AC) applied to different tasks. However, annotated data is limited for most scenarios, such as domestic environments. To address this challenge, we propose an $\textbf{A}$u…
- Controllable Sim Agents with Behavior LatentsJuanwu Lu, Junyu Zhu, Ziran Wang · arXiv · Jul 2, 2026
Realistic traffic simulation requires agents that imitate logged behavior and can also be steered along interpretable axes. Such controllability enables engineers to isolate variables, reproduce specific edge cases, and test autonomous syst…
- Fast Multi-dimensional Refusal Subspaces via RFM-AGOPThomas Winninger · arXiv · Jul 2, 2026
Steering and monitoring activations in Large Language Models (LLMs) are increasingly used for both safety and interpretability. Early work assumed behaviours are encoded along single linear directions, but recent findings suggest complex be…
- Muon as a Residual ConnectionHao Huang · arXiv · Jul 1, 2026
Muon has recently emerged as one of the most effective optimizers for training large neural networks, yet its empirical success has been explained from several different perspectives. In this paper, we propose a simple mechanistic interpret…
- Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning ModelsSubramanyam Sahoo, Aman Chadha, Vinija Jain, Divya Chaudhary · arXiv · Jun 29, 2026
Conservative offline training is widely advocated as a safe foundation for subsequent online adaptation: if a policy stays close to well-supported behaviour, the argument goes, it is less likely to exploit imperfections in a learned reward …
- C$^{2}$R: Cross-sample Consistency Regularization Mitigates Feature Splitting and Absorption in Sparse AutoencodersHaoran Jin, Xiting Wang, Shijie Ren, Hong Xie et al. · arXiv · Jun 29, 2026
Sparse Autoencoders (SAEs) are widely used to interpret large language models by decomposing activations into sparse, human-understandable features, but scaling to large dictionaries exposes fundamental challenges. Systematic studies reveal…
- Democratic ICAI: Debating Our Way to Steering Principles from PreferencesKevin Kingslin, Anish Natekar, Ashutosh Ranjan, Vivek Srivastava et al. · arXiv · Jun 26, 2026
Preference-based alignment often struggles to capture the reasoning that underlies human judgments. Many evaluations rely on multiple interacting criteria, yet pairwise labels reveal only the final choice rather than the considerations that…
- Bridging Ab Initio Symmetries and Global Nuclear Masses with Interpretable Neural NetworksPhong Dang, Evander Espinoza, Xiaoliang Wan, Michela Negro et al. · arXiv · Jun 26, 2026
Ab initio modeling has established Wigner's SU(4) and Elliott's SU(3) as dominant symmetries of the nuclear force in light and intermediate-mass nuclei. We ask whether they also govern nuclear binding across the entire chart. Our aim is not…
- COCOLogic-V2: Identifying Logical Inconsistencies via Truly Hard-NegativesDavid Steinmann, Antonia Wüst, Kristian Kersting, Wolfgang Stammer · arXiv · Jun 26, 2026
While interpretable models such as concept bottleneck models (CBMs) and program synthesis methods enable verification of model decisions, their evaluation is typically limited to simple tasks, leaving complex reasoning on real-world images …
- Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty PredictionChenguang Wang, Ming Li, Xinyue Zeng, Zhuochun Li et al. · arXiv · Jun 26, 2026
Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test construction. Existing methods often depend on costly human calibration or item-level textual representation…