Latest Efficient ML Research Papers
The newest Efficient ML papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Efficient ML so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Efficient ML papers in your inbox — free →Recent papers
- Model-Aware Schedules Improve Generation via Fiberwise Optimal TransportLuyi Jia, Boyan Zhang, Yilun Liu, Steffen Rulands · arXiv · Sep 10, 2026
Diffusion and flow-matching schedules control the signal and noise coefficients that mix data and noise along affine probability paths. Minimizing a kinetic action defined on coefficient paths, motivated by optimal transport, helps explain …
- A Unified Per-Token Gating Family for On-Policy Distillation: FKL/RKL Mixing with Multi-Channel and Bias CoefficientsSuwan Wu, Yumeng Lin, Pengcheng Yuan, Xiaolong Jiang · arXiv · Sep 10, 2026
Per-token gating of forward/reverse KL losses has become a standard technique for on-policy knowledge distillation (OPD), but existing methods such as EOPD (Jin et al., 2026) and ToDi (Jung et al., 2025) each fix a single gating signal and …
- LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language GenerationDongfang Zhao · arXiv · Sep 10, 2026
Large language model serving costs scale directly with output sequence length, yet standard preference alignment often inflates response verbosity without improving utility. We study whether the parameterization of post-training updates aff…
- Why Does Post-Training Quantization Work?Yuxiang Chen, Michael Beyer, Jun Zhu, Jianfei Chen · arXiv · Sep 10, 2026
Post-training quantization compresses large language models (LLMs) by storing their weights at reduced precision, and each quantized weight introduces an error into the hidden states. Naively, these errors should accumulate with depth and c…
- Negative Self-Distillation: Learning to Reason by Avoiding FlawsRongcan Pei, Zhepei Wei, Shuyao Xu, Xinyu Zhu et al. · arXiv · Sep 10, 2026
On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However,…
- Silver Rate Is (Almost) Optimal for Gradient Descent AccelerationYuhan Ye, Kaizhao Liu · arXiv · Sep 8, 2026
We study how far gradient descent (GD) can be accelerated by predetermined nonnegative stepsizes in smooth convex optimization. Writing $p_{\mathrm{sil}}=\log_2(1+\sqrt{2})$, we prove an $Ω\left(n^{-p_{\mathrm{sil}}-O(\sqrt{\log\log n/\log …
- PlayTrain: An Efficient Reinforcement Learning Framework for LLM-Generated Adaptable JavaScript GamesRyan Truong, Lance Ying, Samuel J. Gershman, Kazuki Irie · arXiv · Sep 8, 2026
While many video-game environments (VGEs) have played crucial roles in advancing reinforcement learning (RL), developing novel VGEs or modifying existing ones to support new features, has been a laborious process requiring extensive hand-co…
- Operator Inference for Elliptic Eigenvalue ProblemsHaoqian Li, Jiguang Sun, Zhiwen Zhang · Communications in Computati... · Sep 8, 2026
Efficient computation of elliptic eigenvalue problems is critical in science and engineering. In this paper, we propose a novel operator learning framework that directly maps arbitrary domains to their associated eigenvalues and eigenfuncti…
- Structure-oriented deep learning for semantic segmentation of bridge point cloudsYu Chen, Chao Lin, Tatsuro Yamane, Shiori Kubo et al. · Automation in Construction · Sep 7, 2026
Bridges built during the post-war construction boom are reaching the end of their design life, and their inspection still relies mainly on manual visual assessment. The specific question addressed is whether bridge components can be segment…
- Distill Globally, Adapt Locally: Reasoning Distillation and Product-Type Test-Time Training for Scalable Trade-Up RecommendationSiliang Liu, Mohammad Ghasemi, Sapan Patel, Amin Banitalebi-Dehkordi · arXiv · Sep 4, 2026
Trade-up recommendation identifies higher-quality alternatives that preserve a customer's purchase intent while offering upgraded benefits. Large language models (LLMs) can reason about such distinctions, but applying them directly to hundr…
- Lightweight Vision Transformer Compression for On-Device Plant Disease Detection in Resource-Constrained Agricultural Field ConditionsMahadev Sunil Kumar, Bhavika Gondi, Desaisetty Venkata Satya Sai Swapnith, Gangireddy Rahul Jogi et al. · arXiv · Sep 4, 2026
Chilli (Capsicum annuum) is one of India's most economically significant crops, yet its productivity is persistently threatened by diseases that are difficult to identify without expert intervention. While Vision Transformers (ViTs) have ac…
- Proton Irradiation Characterization of an Open-Source ML Accelerator on a Zynq UltraScale+ MPSoCSaad Memon, Rafal Graczyk, Jan Swakoń, Leszek Grzanka et al. · arXiv · Sep 4, 2026
As spaceborne computing systems increasingly rely on neural network (NN) accelerators, the opacity of commercial, black-box architectures severely restricts the development of verifiable radiation mitigation strategies. Open-source, registe…
- Hessian-based molecular conformation augmentation for a scalable and efficient strategy of machine learning interatomic potentialsBumju Kwak, Jeonghee Jo · arXiv · Sep 4, 2026
While machine-learning interatomic potentials (MLIPs) have successfully learned potential energy surfaces (PES) and atomic forces, many practical applications, such as vibrational analysis and transition state search, rely heavily on the PE…
- Quantum-enhanced hybrid-model compression using knowledge distillationLuigi Barbato, Massimo Esposito, Francesco Gargiulo · Quantum Machine Intelligence · Sep 4, 2026
Abstract Quantum computing has emerged as a promising paradigm for addressing computational tasks intractable for classical systems, leveraging quantum mechanical principles such as superposition and entanglement to efficiently explore high…
- Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVRBoyan Li, Bingsen Chen, Chenghao Yang, Ping Nie et al. · arXiv · Sep 3, 2026
Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have emerged as two dominant methods for post-training reasoning LLMs. Prior work uses OPD's dense token-level supervision to complement the sparse RL re…
- Hardware-Aware FP4 FlashAttention-4Robert Hu · arXiv · Sep 3, 2026
Blackwell's 4-bit floating-point (FP4) tensor cores do not automatically make attention faster because softmax conversion and on-chip dependencies dominate once its matrix products shrink. We address this with \emph{Direct-P} for noncausal …
- Conditioning Degenerate Diffusion ModelsUğur Aydın, Tamer Başar · arXiv · Sep 3, 2026
Current conditioned generative models heavily rely on score functions for guidance during training. When the generative model is a diffusion process with a singular diffusion coefficient and the underlying (conditional) densities either do …
- Subspace Inference Enables Efficient Active Reward Learning from PreferencesYutai Zhou, Erdem Bıyık · arXiv · Sep 3, 2026
Reinforcement learning from human feedback (RLHF) has emerged as a powerful yet sample-inefficient approach for learning reward models from human preferences, making active learning a critical component in synthesizing informative preferenc…
- Three-dimensional magnetotelluric Bayesian inversion based on Stein variational gradient descentZehan Liao, Hao Yang, Xin Zhang, Ji Gao et al. · Geophysical Journal Interna... · Sep 3, 2026
Summary To address the challenge of assessing the reliability of three-dimensional (3D) magnetotelluric (MT) inversion results, we have developed a variational inference (VI) inversion framework (VI-MT) based on the Stein Variational Gradie…
- Joint Exploration of Neural Networks and Systolic Hardware for Improved AI Accelerator PerformanceAnnina Gutermann, Alexey Serdyuk, Foivos Paraskevas, Hella Toto Kiesa et al. · Journal of Signal Processin... · Sep 3, 2026
Abstract When executing common Neural Networks (NNs) on custom AI accelerators, the high performance suggested by advertised Giga or Tera Operations per Second (GOPS/TOPS) is typically not achieved, as low hardware utilization often leads t…
- Machine‐Learning Enhanced Parametric Reynolds‐Averaged Navier‐Stokes Equations at the Full‐ and Reduced‐Order LevelsDavide Oberto, Maria Strazzullo, Stefano Berrone · International Journal for N... · Sep 3, 2026
ABSTRACT In this contribution, we focus on the Reynolds‐averaged Navier‐Stokes (RANS) models and their exploitation to build reliable reduced‐order models to further accelerate predictions for real‐time applications and many‐query scenarios…
- Digital twin-driven edge–cloud collaborative remote fault diagnosis and intelligent predictive maintenance for critical hydropower equipmentDong Yang, Zhile Jiang, Kai Zheng, Yue Wang · Discover Artificial Intelli... · Sep 3, 2026
Abstract To address the problems of lagging remote monitoring, complex fault mechanisms, and inefficient maintenance decisions for key equipment in hydropower stations, this paper proposes a remote diagnosis and intelligent maintenance meth…
- Improved Gradient Descent Lower Bounds Beyond NesterovYuhan Ye, Kaizhao Liu · arXiv · Sep 2, 2026
We study how far gradient descent (GD) can be accelerated by predetermined stepsizes in smooth convex optimization. Going beyond the classical $Ω(n^{-2})$ first-order oracle lower bound of Nemirovsky and Yudin, we prove an $Ω(n^{-1.6342})$ …
- Cliff: Learning Process Rewards from the First MistakePeixuan Han, Runhui Wang, Ketan Ramaneti, Jie Hao et al. · arXiv · Sep 2, 2026
Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for large language model (LLM) post-training, but its reliance on coarse outcome rewards leads to limited guidance on intermediate reasoning processes.…
- LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style UpdatesDmitrii Andriianov, Andrey Veprikov, Aleksandr Beznosikov · arXiv · Sep 2, 2026
Low-rank adaptation (LoRA) is the standard way to fine-tune large models, yet when its two factors are trained independently, the update ignores the geometry of the low-rank weight change it induces. We introduce LoRA-TSD, an optimizer that…
- The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent GloballyJundong Hu, Shekar Ramachandran · arXiv · Sep 1, 2026
Post-training quantization (PTQ) is widely used to reduce the cost of serving large language models (LLMs), but its accuracy cost is uneven and is often tuned per model. We study where quantization damage occurs and how to allocate a small …
- A Mathematical Theory of Reusable Neural Bases for Network CompressionBinshuai Wang · arXiv · Sep 1, 2026
As large AI models become increasingly prevalent across a wide range of applications, memory cost has become a critical bottleneck in both training and inference. To mitigate this issue, we introduce the Linear Reusable Neural Bases Archite…
- Sierpiński--Knopp Wasserstein Distance for Persistence Diagrams and Applications to 2-Wasserstein ApproximationSebastien Tchitchek, Julien Tierny · arXiv · Sep 1, 2026
This paper introduces the Sierpiński-Knopp (SK) Wasserstein distance, a fast metric between persistence diagrams. The SK-Wasserstein distance, denoted $d_{\mathrm{SK}}$, maps diagram points and their diagonal projections to the unit interva…
- LatentPress: Context Compression Beyond Text and VisionZhengze Zhou, Hejian Sang · arXiv · Sep 1, 2026
Compressed context is usually carried as human-readable text or as rendered images that must be decoded, even when its consumer is a language model. We introduce LatentPress, which writes conversational histories and long documents into a t…
- Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy SearchZhiliang Chen, Sebastian Ament, David Eriksson, Maximilian Balandat et al. · arXiv · Sep 1, 2026
Optimal hyperparameter scaling laws describe how the best hyperparameters for large language model (LLM) training change with model and data scale, enabling practitioners to predict optimal configurations at production scales without expens…