Latest Mixture of Experts Research Papers
The newest Mixture of Experts papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Mixture of Experts so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Mixture of Experts papers in your inbox — free →Recent papers
- Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated DataAtindra Jha, Margaret Li, Jure Leskovec, Percy Liang et al. · arXiv · Sep 10, 2026
As the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activated Transformers, but the effects of data repetition remains …
- T1: Terminal Agent Reinforcement Learning for Long-Horizon TasksJunyao Yang, Yucheng Shi, Zhongzhi Li, Ruhan Wang et al. · arXiv · Sep 10, 2026
Agent usage is shifting toward long-horizon tasks such as coding and scientific discovery, among which terminal tasks are especially important. We introduce T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, o…
- Domain-Adaptive Mixture-of-Experts for Cross-Dataset Lithium-Ion Battery State-of-Health Prediction via Adaptive Strategy SelectionTeng Liu, Wei Li, Zhiqiang Li · Batteries · Sep 10, 2026
Accurate cross-dataset state-of-health prediction for lithium-ion batteries remains challenging due to distribution shifts arising from diverse cathode chemistries, operating temperatures, and charge–discharge protocols across heterogeneous…
- TopoMoE: a topology-fidelity discrete mixture-of-experts framework for multi-modal knowledge graph completionHongwei Chen, Chang Liu, Peng Shao · The Journal of Supercomputing · Sep 6, 2026
- Regime-Aware Reinforcement Learning: A Mixture-of-Experts Framework for Dynamic Asset AllocationYirui Luo, John M. Mulvey · The Journal of Financial Da... · Sep 5, 2026
Recent advances in reinforcement learning (RL) have spurred growing interest in its application to multi-period financial planning. Existing literature broadly follows three paradigms: hybrid RL, scalable RL, and end-to-end RL. This paper d…
- A Verifier-Guided Explainable Reasoning Framework with Gold-Anchored QLoRA, Task-Aware Mixture-of-Experts, and Group-Relative RLVRThi Kim Trang Vo, Nam Tien Le, Thi Kim Nguyet Vo, Minh Khang Tran et al. · arXiv · Sep 4, 2026
Large language models (LLMs) show strong reasoning ability, but their explanations can remain inconsistent, weakly grounded, or difficult to verify. We propose a verifier-guided explainable reasoning framework for transparent educational qu…
- MS2-DAMoE: Multiscale multi-source domain alignment with adaptive mixture-of-experts for industrial time-series predictionXue Xu, Kaifang Li, Yuanjian Fu, Chaomin Luo et al. · Journal of Process Control · Sep 4, 2026
- Image-aware mixture of experts automates the detection of ovarian follicle rupture in a high-throughput ex vivo ovulation assayMahdi Babaei, Jiyang Zhang, Shuo Xiao, Yu Gan · Biology of Reproduction · Sep 3, 2026
An AI-based dynamic, image-aware Mixture-of-Experts (MoE) framework automates ex vivo ovulation detection and offers a highly accurate, scalable solution for studying the biology of ovulation, drug development, and reproductive toxicity tes…
- LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style UpdatesDmitrii Andriianov, Andrey Veprikov, Aleksandr Beznosikov · arXiv · Sep 2, 2026
Low-rank adaptation (LoRA) is the standard way to fine-tune large models, yet when its two factors are trained independently, the update ignores the geometry of the low-rank weight change it induces. We introduce LoRA-TSD, an optimizer that…
- Archetype-based mixture of experts for interpretable and generalisable data-driven turbulence modelling - 2nd International Symposium on AI and Fluid Mechanics.A Montoya Santamaría, P Cinnella · OpenAlex · Sep 2, 2026
- Mixture of Mesh Experts With Random Walk Transformer GatingAmir Belder, Ayellet Tal · Computer Graphics Forum · Sep 2, 2026
Abstract In recent years, various methods have been proposed for mesh analysis, each offering distinct advantages and often excelling on different object classes. We present a novel Mixture of Experts (MoE) framework designed to harness the…
- Layer-Scoped Expert-Budget Expansion Discovers Succinct Convergence in Sparse Mixture-of-Experts Reasoning: Reducing Reasoning Tokens at Matched Accuracy on MMLU-ProVincenzo Agrillo · Zenodo (CERN European Organ... · Sep 2, 2026
This paper introduces layer-scoped expert-budget expansion, a training-free, runtime-only routing modification for sparse Mixture-of-Experts (MoE) language models. By expanding the expert selection budget ($N \ge K$) exclusively within late…
- Layer-Scoped Expert-Budget Expansion Discovers Succinct Convergence in Sparse Mixture-of-Experts Reasoning: Reducing Reasoning Tokens at Matched Accuracy on MMLU-ProVincenzo Agrillo · Zenodo (CERN European Organ... · Sep 2, 2026
This paper introduces layer-scoped expert-budget expansion, a training-free, runtime-only routing modification for sparse Mixture-of-Experts (MoE) language models. By expanding the expert selection budget ($N \ge K$) exclusively within late…
- Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark ConclusionsAhmed El Kady, Aravind Narayanan, Rehana Noorani, Yani Ioannou et al. · arXiv · Aug 31, 2026
Efficient evaluation changes the protocol used to support claims about model behavior, yet it is rarely tested whether those claims remain stable after the evaluation itself is made cheaper. We stress-test conclusion robustness in responsib…
- Plasticity-Aware Mixture of Experts (PA-MoE) for Learning Under QoE Shifts in Adaptive Video StreamingAhmadreza Montazerolghaem, Sharifian · OpenAlex · Aug 31, 2026
- Budgeted Reward Allocation: A Position on Unifying Sparse, Verifiable, and Hierarchical Reward Signals in LLM Post-TrainingMarko Tahvanainen · Zenodo (CERN European Organ... · Aug 29, 2026
Reward hacking, verbosity, and hallucination are usually treated as separate failure modes in the post-training of large language models, addressed by separate mechanisms: reward shaping for verbosity, verifiable reward models (RLVR) for ha…
- MM-Spectrum: Multimodal Multi-spectral Molecular Structural Elucidation with a Stable MoE FrameworkHai-tao Yu, Nan Min, Zheng Fang, Hongyu Zhan et al. · arXiv · Aug 27, 2026
Inferring molecular structures from multimodal spectroscopic measurements requires integrating complementary yet highly heterogeneous signals. However, the common paradigm of directly concatenating multispectral sequences can exhibit anomal…
- MoTE: Mixture of Task Experts for Multi-Task Video UnderstandingMuhammad Asad Ali, Umar Khan, Nadia Robertini, Didier Stricker · arXiv · Aug 25, 2026
Procedural video-language models must solve heterogeneous tasks from the same visual evidence, including action recognition, forecasting, and procedure prediction. Dense transformer decoders share the same feed-forward networks across tasks…
- RIS-MoE: robust and secure image steganography via latent-space optimization with mixture-of-experts denoisingGenfan Yang, Rongchang Duan, Hong Zhang, Jie Gan et al. · Cybersecurity · Aug 25, 2026
Abstract Diffusion-based generative image steganography enables covert communication by synthesizing stego images without relying on cover images. However, existing latent-space methods still struggle to balance robustness, steganographic s…
- DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-ExpertsVlad Hondru, Florinel Alin Croitoru, Iuliana Georgescu, A. Sophia Koepke et al. · arXiv · Aug 24, 2026
Audio-visual deepfake detection is an actively studied topic, where one of the main challenges is to develop detectors able to generalize across deepfake generation methods. We conjecture that overfitting can be mitigated by extracting mult…
- A hybrid data-driven approach for state of charge and state of health estimation of lithium-ion batteries: using transformer combining multi-gate mixture-of-experts as a battery modelJiabo Li, Xingyu Liu, Di Tian, Yuanjin Zhang et al. · Ionics · Aug 24, 2026
- Adaptive frequency-band mixture of experts for long-horizon multivariate time-series forecastingJiawei Zhang, Hairui Wang, Ya Li, Guifu Zhu · Computers & Electrical Engi... · Aug 22, 2026
- Multi-task mixture-of-experts vision transformer for cross-bridge anomaly detectionQiuyue Pan, Yuequan Bao, Feiyuan Long · Mechanical Systems and Sign... · Aug 22, 2026
- Towards continual stance detection via Dual Prototype Aware Mixture of Experts NetworkYuzhe Ding, Yuxiang Peng, Kang He, Li Zheng et al. · Information Processing & Ma... · Aug 20, 2026
- Implicit integration of a Mixture of Experts neural network surrogate elasto-viscoplastic constitutive model of HT9 steel in a finite element framework: Verification, comparative assessment, and validationK. M. Zaheen Nasir, Giang Huynh, Benjamin W. Spencer, Laurent Capolungo et al. · Finite Elements in Analysis... · Aug 20, 2026
- Deep learning-based PET/CT mixture-of-experts model for relapse risk stratification in relapsed/refractory classical hodgkin lymphoma: a multicenter studyChong Jiang, Zekun Jiang, Xinyu Zhang, Zitong Zhang et al. · European Journal of Nuclear... · Aug 19, 2026
- On the Fragility of Self-Improving Agents: Variance, Task Order, and UnderspecificationQinyuan Ye, Yu Li, Yada Pruksachatkun, Jiaxin Zhang et al. · arXiv · Aug 18, 2026
Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank--have shown great promise in recent literature. However, the reliability aspects of these methods…
- DeaMoE: Efficient MoE Structure for Fast Small-Batch DecodingZewen Jin, Shen Fu, Zeping Duan, Shannon Wang et al. · arXiv · Aug 14, 2026
Mixture-of-Experts (MoE) models have been widely adopted in real-time interactive applications such as coding assistants, real-time audio-video interaction systems. To meet the extremely low response latency requirements of these scenarios,…
- Regime-Gated Residual Mixture-of-Experts for Cross-Sectional Volatility ForecastingJunyi Ye, Gargi Vijay Borde · arXiv · Aug 12, 2026
Financial volatility is regime dependent, yet incorporating regime information into neural networks can also destabilize training. This paper asks where such information should enter a neural cross-sectional volatility forecasting model. We…
- SAGE-UNet: A mixture-of-experts spatially-aware gated enhancement U-Net for flood water segmentation from Sentinel-1 SAR and post-event impact assessment—learning from flood events in Germany and SpainArmin Moghimi, Alireza Bahrami Mahtaj, Ali Jamali, Mario Welzel et al. · Figshare · Aug 12, 2026
The increasing frequency and intensity of floods pose substantial risks to residential areas, infrastructure, and economic systems, highlighting the need for reliable flood mapping to support emergency response and damage assessment. Even t…