Latest Scaling Laws Research Papers
The newest Scaling Laws papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Scaling Laws so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Scaling Laws papers in your inbox — free →Recent papers
- LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style UpdatesDmitrii Andriianov, Andrey Veprikov, Aleksandr Beznosikov · arXiv · Sep 2, 2026
Low-rank adaptation (LoRA) is the standard way to fine-tune large models, yet when its two factors are trained independently, the update ignores the geometry of the low-rank weight change it induces. We introduce LoRA-TSD, an optimizer that…
- Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy SearchZhiliang Chen, Sebastian Ament, David Eriksson, Maximilian Balandat et al. · arXiv · Sep 1, 2026
Optimal hyperparameter scaling laws describe how the best hyperparameters for large language model (LLM) training change with model and data scale, enabling practitioners to predict optimal configurations at production scales without expens…
- On the Fragility of Self-Improving Agents: Variance, Task Order, and UnderspecificationQinyuan Ye, Yu Li, Yada Pruksachatkun, Jiaxin Zhang et al. · arXiv · Aug 18, 2026
Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank--have shown great promise in recent literature. However, the reliability aspects of these methods…
- Scaling-Law Analysis of SignSGD: From Feature-Space Linear Regression to LLM Pre-trainingBinghui Li, Jianan Wang, Jinbo Wang, Lean Wang et al. · Sci4DL 2026 · Mar 2, 2026
Despite their widespread use in deep learning, the mechanisms underlying the effectiveness of adaptive gradient methods in large-scale training remain poorly understood. In this work, we provide a scaling-law analysis of SignSGD, a minimal …
- Configuration-to-Performance Scaling Law with Neural AnsatzHuaqing Zhang, Kaiyue Wen, Tengyu Ma · Sci4DL 2026 · Mar 2, 2026
Researchers build scaling laws to forecast the training performance of expensive large-scale runs with larger model size $N$ and data size $D$. These laws assume that other training hyperparameters are optimally chosen, which can require si…
- Configuration-to-Performance Scaling Law with Neural AnsatzHuaqing Zhang, Kaiyue Wen, Tengyu Ma · ICLR 2026 Workshop DATA-FM · Mar 2, 2026
Researchers build scaling laws to forecast the training performance of expensive large-scale runs with larger model size $N$ and data size $D$. These laws assume that other training hyperparameters are optimally chosen, which can require si…
- DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous DrivingYingyan Li, Shuyao Shang, Weisong Liu, Bing Zhan et al. · arXiv.org · Oct 14, 2025
Scaling Vision-Language-Action (VLA) models on large-scale data offers a promising path to achieving a more generalized driving intelligence. However, VLA models are limited by a ``supervision deficit'': the vast model capacity is supervise…
- Relative-Based Scaling Law for Neural Language ModelsBaoqing Yue, Jinyuan Zhou, Zixi Wei, Jingtao Zhan et al. · ICLR 2026 Conference Withdrawn Submission · Sep 20, 2025
Scaling laws aim to accurately predict model performance across different scales. Existing scaling-law studies almost exclusively rely on cross-entropy as the evaluation metric. However, cross-entropy provides only a partial view of perform…
- Scaling Law for Code: A More Data-Hungry RegimeXianzhen Luo, Wenzhen Zheng, Qingfu Zhu, Rongyi Zhang et al. · ICLR 2026 Conference Withdrawn Submission · Sep 19, 2025
The training of large language models (LLMs) for code generation incurs substantial computational costs, yet the resource allocation strategies are often guided by scaling laws derived from natural language (NL). Given the distinct statisti…
- LayerMix Law: Scaling Law for Large Language Models on Quality-Weighted Mixture Data with RepetitionFengze Liu, Weidong Zhou, BINBINLIU, Ping Guo et al. · ICLR 2026 Conference Withdrawn Submission · Sep 19, 2025
Upweighting high-quality data in large language model (LLM) pretraining typically improves performance. However, the limited availability of high-quality data—particularly in overtrained regimes—means that stronger upweighting often increas…
- P-Law: Predicting Quantitative Scaling Law with Entropy Guidance in Large Recommendation ModelsTingjia Shen, Hao Wang, Chuhan Wu, Jin Yao Chin et al. · NeurIPS 2025 poster · Sep 18, 2025
With the growing size of data and models in Large Recommendation Models, the time required for debugging has become increasingly prohibitive, underscoring the urgent need for effective guidance in parameter configuration. The Scaling Law (S…
- Predictable Scale (Part II) --- Farseer: A Refined Scaling Law in LLMsHouyi Li, Wenzhen Zheng, Qiufeng Wang, Zhenyu Ding et al. · NeurIPS 2025 spotlight · Sep 18, 2025
Training Large Language Models (LLMs) is prohibitively expensive, creating a critical scaling gap where insights from small-scale experiments often fail to transfer to resource-intensive production systems, thereby hindering efficient innov…
- Kinetics: Rethinking Test-Time Scaling LawRanajoy Sadhukhan, Zhuoming Chen, Haizhong Zheng, Beidi Chen · NeurIPS 2025 poster · Sep 18, 2025
We rethink test-time scaling laws from a practical efficiency perspective, revealing that the effectiveness of smaller models is significantly overestimated. Prior work, grounded in compute-optimality, overlooks critical memory access bottl…
- Parallel Scaling Law for Language ModelsMouxiang Chen, Binyuan Hui, Zeyu Cui, Jiaxi Yang et al. · NeurIPS 2025 poster · Sep 18, 2025
It is commonly believed that scaling language models should commit a significant space or time cost, by increasing the parameters (parameter scaling) or output tokens (inference-time scaling). We introduce another and more inference-efficie…
- Beyond Scaling Law: A Data-Efficient Distillation Framework for ReasoningXiaojun Wu, Xiaoguang Jiang, Huiyang Li, Jucai Zhai et al. · arXiv.org · Aug 13, 2025
Large language models (LLMs) demonstrate remarkable reasoning capabilities in tasks such as algorithmic coding and mathematical problem-solving. Recent methods have improved reasoning through expanded corpus and multistage training combinin…
- Revisiting the theory of van Driest: a general scaling law for the skin-friction coefficient of high-speed turbulent boundary layersZhiye Zhao, Lin Fu (Associate Professor) · Journal of Fluid Mechanics · May 29, 2025
Abstract The skin-friction coefficient is a dimensionless quantity defined by the wall shear stress exerted on an object moving in a fluid, and it decreases as the Reynolds number increases for wall-bounded turbulent flows over a flat plate…
- Scaling Law for Quantization-Aware TrainingMengzhao Chen, Chaoyi Zhang, Jing Liu, Yutao Zeng et al. · arXiv.org · May 20, 2025
Large language models (LLMs) demand substantial computational and memory resources, creating deployment challenges. Quantization-aware training (QAT) addresses these challenges by reducing model precision while maintaining performance. Howe…
- Parallel Scaling Law for Language ModelsMouxiang Chen, Binyuan Hui, Zeyu Cui, Jiaxin Yang et al. · Neural Information Processing Systems · May 15, 2025
It is commonly believed that scaling language models should commit a significant space or time cost, by increasing the parameters (parameter scaling) or output tokens (inference-time scaling). We introduce the third and more inference-effic…
- Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem SolvingXinji Mai, Haotian Xu, Zhong-Zhi Li, W. Xing et al. · Semantic Scholar · May 12, 2025
Large Language Models (LLMs) often struggle with mathematical reasoning tasks requiring precise, verifiable computation. While Reinforcement Learning (RL) from outcome-based rewards enhances text-based reasoning, understanding how agents au…
- A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling LawQianjun Pan, Wenkai Ji, Yuyang Ding, Junsong Li et al. · arXiv.org · May 5, 2025
This survey explores recent advancements in reasoning large language models (LLMs) designed to mimic"slow thinking"- a reasoning process inspired by human cognition, as described in Kahneman's Thinking, Fast and Slow. These models, like Ope…
- a1: Steep Test-time Scaling Law via Environment Augmented GenerationLingrui Mei, Shenghua Liu, Yiwei Wang, Baolong Bi et al. · Annual Meeting of the Association for Computational Linguistics · Apr 20, 2025
Large Language Models (LLMs) have made remarkable breakthroughs in reasoning, yet continue to struggle with hallucinations, logical errors, and inability to self-correct during complex multi-step tasks. Current approaches like chain-of-thou…
- Unsourced Random Access in MIMO Quasi-Static Rayleigh Fading Channels: Finite Blocklength and Scaling Law AnalysesJunyuan Gao, Yongpeng Wu, Giuseppe Caire, Wei Yang et al. · IEEE Transactions on Information Theory · Mar 21, 2025
This paper considers the unsourced random access (URA) problem with a random and unknown number of active users in multiple-input multiple-output (MIMO) quasi-static Rayleigh fading channels. We derive non-asymptotic achievability bounds on…
- L2M: Mutual Information Scaling Law for Long-Context Language ModelingZhuo Chen, Oriol Mayn'e i Comas, Zhuotao Jin, Di Luo et al. · Neural Information Processing Systems · Mar 6, 2025
We present a universal theoretical framework for understanding long-context language modeling based on a bipartite mutual information scaling law that we rigorously verify in natural language. We demonstrate that bipartite mutual informatio…
- Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model PretrainingHouyi Li, Wenzheng Zheng, Qiufeng Wang, Hanshan Zhang et al. · Semantic Scholar · Mar 6, 2025
The impressive capabilities of Large Language Models (LLMs) across diverse tasks are now well established, yet their effective deployment necessitates careful hyperparameter optimization. Although existing methods have explored the influenc…
- Unlocking Scaling Law in Industrial Recommendation Systems with a Three-step Paradigm based Large User ModelBencheng Yan, Shilei Liu, Zhiyuan Zeng, Zihao Wang et al. · Web Search and Data Mining · Feb 12, 2025
Recent advancements in autoregressive Large Language Models (LLMs) have achieved remarkable progress, largely driven by their scalability—commonly formalized as the scaling law. Inspired by these successes, there has been growing interest i…
- Scaling Law for Intrinsic Fracture Energy of Diverse Stretchable NetworksChase M. Hartquist, Shu Wang, Qiaodong Cui, W. Matusik et al. · Physical Review X · Jan 8, 2025
Networks of interconnected materials permeate throughout nature, biology, and technology due to exceptional mechanical performance. Despite the importance of failure resistance in network design and utility, no existing physical model effec…
- Predictable Scale: Part I - Optimal Hyperparameter Scaling Law in Large Language Model PretrainingHouyi Li, Wenzheng Zheng, Jingcheng Hu, Qiufeng Wang et al. · arXiv.org · Jan 1, 2025
- (Mis)Fitting Scaling Laws: A Survey of Scaling Law Fitting Techniques in Deep LearningMargaret Li, Sneha Kudugunta, Luke S. Zettlemoyer · International Conference on Learning Representations · Jan 1, 2025
- From Scaling Law to Sub-Scaling Law: Understanding the Diminishing Returns of Larger ModelsZhengyu Chen, Siqi Wang, Teng Xiao, Yudong Wang et al. · ICLR 2025 Conference Withdrawn Submission · Sep 28, 2024
Traditional scaling laws suggest that performance metrics of language models improve predictably with increases in model or dataset size. However, recent works display sub-scaling growth for large language models, where performance improvem…
- D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language ModelsHaoran Que, Jiaheng Liu, Ge Zhang, Chenchen Zhang et al. · NeurIPS 2024 poster · Sep 25, 2024
Continual Pre-Training (CPT) on Large Language Models (LLMs) has been widely used to expand the model’s fundamental understanding of specific downstream domains (e.g., math and code). For the CPT on domain-specific LLMs, one important quest…