Latest Foundation Models Research Papers
The newest Foundation Models papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Foundation Models so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Foundation Models papers in your inbox — free →Recent papers
- CausalArena: Benchmarking Causal Discovery in the Foundation Model EraZi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye · arXiv · Sep 10, 2026
Causal discovery aims to uncover causal structures from data and is fundamental to scientific reasoning and intervention-based decision making. Its evaluation relies heavily on structural causal models (SCMs), which specify a causal graph t…
- Evaluating Time-Series Foundation Models and Multimodal Dietary Context for CGM ForecastingBowen Zhang, Hsiu-Wen Cheng, Hongyu Yang, Evie L. Shen et al. · arXiv · Sep 10, 2026
Continuous glucose monitoring (CGM) provides high-frequency measurements of glucose dynamics and enables short-term glucose forecasting for diabetes management. Although time-series foundation models have shown strong general forecasting ab…
- SIRF: A Spec-Internalized Risk Foundation Model for Industrial Content Risk ControlSuwan Wu, Yumeng Lin, Pengcheng Yuan, Xiaolong Jiang · arXiv · Sep 10, 2026
For industrial content risk control, the real deployment constraint is not average accuracy but how much risk can be auto-handled under high precision and second-level latency. We present SIRF (Spec-Internalized Risk Foundation Model), whic…
- Why Does Post-Training Quantization Work?Yuxiang Chen, Michael Beyer, Jun Zhu, Jianfei Chen · arXiv · Sep 10, 2026
Post-training quantization compresses large language models (LLMs) by storing their weights at reduced precision, and each quantized weight introduces an error into the hidden states. Naively, these errors should accumulate with depth and c…
- Geospatial Foundation Models Capture Health-Relevant Dimensions of Place Beyond Conventional Social Risk IndicesNathaniel Hendrix, Carl Y. Zhang, Chris Heitzig, Andrew Bazemore et al. · arXiv · Sep 10, 2026
Area-based social risk indices summarize residents' socioeconomic conditions but incompletely capture physical features of place that may affect health. We evaluated whether numerical representations of physical place produced by four geosp…
- A Later Test Set Is Not a New Domain: Pretraining Familiarity Survives a Contamination-Free Hold-OutMahdi Naser Moghadasi, Faezeh Ghaderi · arXiv · Sep 9, 2026
Time-series foundation models are evaluated almost exclusively on public archives that predate them, so a strong score cannot be separated from having seen the test set during pretraining. The obvious remedy is a hold-out that postdates the…
- Training-Free Task Vectors for LLM Behavioral ControlGabriel J. Perin, Lucas Boscaini, André Araujo, Nina S. T. Hirata · arXiv · Sep 8, 2026
Task vectors enable post-training model editing by identifying semantically meaningful directions in weight space, typically computed as the difference between a fine-tuned model and its pretrained initialization. However, this reliance on …
- UniMate: One Unified Model to Animate Diverse SkeletonsLinzhan Mou, Jiahui Lei, Zhiyang Dou, Chenyue Cai et al. · arXiv · Sep 4, 2026
Recent advances in automatic rigging now deliver animation-ready 3D assets at scale, yet generating the motion to drive them remains a bottleneck. Existing learned animators are topology-constrained: they rely on category-specific templates…
- Graph Machine: Towards Better Pretraining via EdgesLintai Hou · arXiv · Sep 2, 2026
We introduce the Graph Machine (GM), an architecture that maintains an $O(n)$-sized state and accesses it through sparse, dynamic routing. Unlike methods with fixed-size states or sparse but static routing, GM preserves $O(n)$ complexity in…
- UE5M3 FP4 Block Scaling for Stable Language Model PretrainingRobert Hu, Carlo Luschi, Paul Balanca · arXiv · Sep 2, 2026
Stable 4-bit floating-point (FP4) pretraining is difficult because the E2M1 payload represents only a narrow range of magnitudes. NVIDIA's Transformer Engine \nv{} recipe addresses this with current-tensor scaling, a randomized Hadamard tra…
- Do Tabular Foundation Models Know Physics? Contamination, Units, and the Deterministic LimitWassim Tenachi, Yashar Hezaveh, Laurence Perreault Levasseur, Pierre-Luc Bacon · arXiv · Sep 2, 2026
Tabular foundation models (TFMs) learn to fill in tables the way language models fill in text, and tables are arguably the format in which most physical measurement arrives. Did they learn any physics in the process? They are Bayesian by co…
- LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style UpdatesDmitrii Andriianov, Andrey Veprikov, Aleksandr Beznosikov · arXiv · Sep 2, 2026
Low-rank adaptation (LoRA) is the standard way to fine-tune large models, yet when its two factors are trained independently, the update ignores the geometry of the low-rank weight change it induces. We introduce LoRA-TSD, an optimizer that…
- Facet-0: A Robotic Foundation Model for Contact-Rich Precise ManipulationHaoyuan Deng, Haichao Liu, Wenkai Guo, Yuan Ling et al. · arXiv · Sep 1, 2026
Real-world robotic assembly at sub-millimeter tolerances demands spatial precision, compliant interaction, and robustness to contact failures. We present Facet-0, a robotic foundation model that predicts and values the contact consequences …
- One Adapter, Many Tasks: Task-Conditioned Feature Transformations for Continual LearningYunxiang Fu, Meng Lou, Yizhou Yu · arXiv · Aug 31, 2026
Class-incremental learning (CIL) requires a model to incrementally learn tasks that contain new classes without accessing earlier training data while preserving the ability to recognize all seen classes. Recently, pretrained-model-based app…
- TSPFN: A Temporal Tabular Foundation Model for Physiological Time Series ClassificationJérémie Stym-Popper, Clément Rambour, Federica Granese, Nicolas Thome et al. · arXiv · Aug 31, 2026
Designing models that generalize effectively in low- to medium-data regimes remains a primary challenge in medical machine learning, particularly for physiological time-series classification. While tabular foundation models such as TabPFN o…
- Curvature-Conditioned Multiscale Momentum with Sphere Constraints for LLM PretrainingShuchen Zhu, Yuxin Fang, Mingze Wang, Kun Yuan · arXiv · Aug 28, 2026
Pretraining accounts for a large fraction of the total computational cost in LLM training. However, noise-dominant gradients and the highly ill-conditioned loss landscape bring severe challenges. Although modern adaptive optimizers such as …
- Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090Kairong Luo, Jiarui Cui, Yaorui Yin, Shengqi Chen et al. · arXiv · Aug 27, 2026
Language model pretraining has become almost synonymous with prohibitive cost, placing it out of reach for much of the academic and open-source communities. Although strong open-source efforts already exist, including open-weight models and…
- HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity RecognitionZihan Ding, Liyu Zhang, Xiaomin Ouyang · arXiv · Aug 27, 2026
Human Activity Recognition (HAR) using inertial measurement units (IMUs) enables a wide range of applications, yet the field still lacks a unified model that can generalize across diverse subjects, devices, and activities. Training such a m…
- Effective Learning Rate Governs Loss Dynamics in Language Model PretrainingZihan Liu, Ruiheng Zheng, Shaobo Zhang, Changxin Tian et al. · arXiv · Aug 25, 2026
We uncover ELR collapse in language model pretraining: learning rate (LR) and parameter norm govern loss dynamics primarily through their ratio, the effective learning rate (ELR). When ELR is matched across runs, their loss trajectories col…
- Parameter-Efficient Self-Supervised Adaptation for EEG-FM under Fixed Computational BudgetsMeghal Dani, Stefanie Liebe · arXiv · Aug 25, 2026
EEG foundation models pretrained via self-supervised learning promise transferable representations, but their generalization remains limited, especially across diverse clinical datasets. Full fine-tuning is impractical for resource-constrai…
- Tydra: An Efficient Hybrid Model for Tabular DataMieszko Komisarczyk, Saurabh Mathur, Maurice Kraus, Sriraam Natarajan et al. · arXiv · Aug 21, 2026
Transformer-based tabular foundation models such as TabPFN achieve strong predictive performance but incur quadratic computational cost with context length. On the other hand, subquadratic SSM-based alternatives such as Hydra trade away acc…
- Explainable Transformer Models for Clinical Prediction Tasks on Structured Electronic Health RecordsJun Ni Du, Lukas Adamek, Maxim Kryukov, Flavio Dormont et al. · arXiv · Aug 20, 2026
Predictive models over structured electronic health records (EHRs) remain central to machine learning for healthcare, but few have jointly emphasized quantitative laboratory information and interpretability with respect to input medical eve…
- Pretraining Reusable Inference Across Views with Synthetic Task PriorsJielong Lu, Zhihao Wu, Jiajun Yu, Zhaoliang Chen et al. · arXiv · Aug 19, 2026
Modern pretrained encoders make representations from heterogeneous views increasingly reusable, but the procedure that determines view utility and combines evidence is still relearned for each downstream task. Consequently, knowledge about …
- On the Fragility of Self-Improving Agents: Variance, Task Order, and UnderspecificationQinyuan Ye, Yu Li, Yada Pruksachatkun, Jiaxin Zhang et al. · arXiv · Aug 18, 2026
Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank--have shown great promise in recent literature. However, the reliability aspects of these methods…
- RecirculationMichael C. Mozer, Shoaib Ahmed Siddiqui, Danny Sawyer, Sunny Sanyal et al. · arXiv · Aug 18, 2026
We describe an inference-time architectural enhancement for off-the-shelf foundation models that markedly reduces perplexity and boosts accuracy across generation and reasoning tasks. Our approach incurs essentially no additional latency du…
- Understanding the Surprising Generalization Properties of Tabular Foundation ModelsNour Shaheen, Junwei Ma, Alex Labach, Frank Hutter et al. · arXiv · Aug 18, 2026
Tabular Foundation Models (TFMs) increasingly rely on in-context learning, where a model receives labelled examples at inference time and predicts labels for new inputs without updating its weights. Existing TFMs are typically trained on ei…
- Approximate Muon with low-rank adaptersBen Anson, Conor Houghton, Edward Milsom · arXiv · Aug 14, 2026
The Muon optimizer shows clear benefits versus alternatives when pretraining neural networks. However, it is used less frequently for parameter-efficient fine-tuning (PEFT). One potential reason is that the most common PEFT method, LoRA, do…
- CytoBERT: A Foundation Model for Cytometry DataSyed Abdul Haseeb Qadri, Bjarne C. Hiller, Felix Blanke, Vanja Sophie Cangalovic et al. · arXiv · Aug 14, 2026
Cytometry measures the complex characteristics of single cells (e.g., counts and protein expression of immune cells) and is widely used across immunological research and clinical settings. However, cytometry data is highly heterogeneous and…
- LittleLearner: Language Models Under Pedagogically Controlled Knowledge ExposureFanfei Li, Jana Zeller, Manuel Prada-Corral, Thaddäus Wiedemer et al. · arXiv · Aug 13, 2026
Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we int…
- Synthetic Persona Pretraining: Alignment from Token ZeroJulian Minder, Viktor Moskvoretskii, Raghav Singhal, Difan Jiao et al. · arXiv · Aug 13, 2026
As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical. Today, alignment, and the assistant identity itself, are typically introduced only after pretra…