Latest Transfer Learning Research Papers
The newest Transfer Learning papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks Transfer Learning so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest Transfer Learning papers in your inbox — free →Recent papers
- Domain-Specific Hallucination Detection in Large Language ModelsVarun Teja Chundru, Debasmita Biswas · arXiv · Sep 10, 2026
Large language models generate fluent text that can contain unfaithful claims -- a phenomenon known as hallucination. We present a multi-signal detection pipeline combining fine-tuned DeBERTa-v3 classification, Monte Carlo (MC) Dropout unce…
- Training-Free Task Vectors for LLM Behavioral ControlGabriel J. Perin, Lucas Boscaini, André Araujo, Nina S. T. Hirata · arXiv · Sep 8, 2026
Task vectors enable post-training model editing by identifying semantically meaningful directions in weight space, typically computed as the difference between a fine-tuned model and its pretrained initialization. However, this reliance on …
- UniMate: One Unified Model to Animate Diverse SkeletonsLinzhan Mou, Jiahui Lei, Zhiyang Dou, Chenyue Cai et al. · arXiv · Sep 4, 2026
Recent advances in automatic rigging now deliver animation-ready 3D assets at scale, yet generating the motion to drive them remains a bottleneck. Existing learned animators are topology-constrained: they rely on category-specific templates…
- LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style UpdatesDmitrii Andriianov, Andrey Veprikov, Aleksandr Beznosikov · arXiv · Sep 2, 2026
Low-rank adaptation (LoRA) is the standard way to fine-tune large models, yet when its two factors are trained independently, the update ignores the geometry of the low-rank weight change it induces. We introduce LoRA-TSD, an optimizer that…
- Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMsJingtan Wang, Arun Verma, Xiaoqiang Lin, Zhengyuan Liu et al. · arXiv · Sep 1, 2026
How to divide a fixed annotation budget between supervised fine-tuning (SFT) and reinforcement learning (RL) during LLM post-training remains an open problem. Existing work characterizes only broad trends (e.g., SFT dominates in low-data re…
- TSPFN: A Temporal Tabular Foundation Model for Physiological Time Series ClassificationJérémie Stym-Popper, Clément Rambour, Federica Granese, Nicolas Thome et al. · arXiv · Aug 31, 2026
Designing models that generalize effectively in low- to medium-data regimes remains a primary challenge in medical machine learning, particularly for physiological time-series classification. While tabular foundation models such as TabPFN o…
- Language-Informed Flow Matching for Trend-Guided Structure-Based 3D Molecular GenerationTianyu Gao, Zhikai Su, Jiashu Li, Wenjun Gao et al. · arXiv · Aug 31, 2026
Structure-based drug design (SBDD) requires ligands that satisfy both 3D target affinity and 1D chemical validity. Existing controllable generation methods often rely on task-specific fine-tuning or externally imposed sampling-time guidance…
- DARTS: Decoder-Aware Representation Tuning via Surgery for Model MergingAaryan Ajay Sharma, Sai Nishanth Padala, Seganrasan Subramanian · arXiv · Aug 28, 2026
Model merging combines multiple task-specific fine-tuned LLMs into a single multi-task model without additional training. However, merged models are known to suffer from representation bias: systematic drift between the merged model's hidde…
- Parameter-Efficient Self-Supervised Adaptation for EEG-FM under Fixed Computational BudgetsMeghal Dani, Stefanie Liebe · arXiv · Aug 25, 2026
EEG foundation models pretrained via self-supervised learning promise transferable representations, but their generalization remains limited, especially across diverse clinical datasets. Full fine-tuning is impractical for resource-constrai…
- Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-PlanningVarun Giridhar, Anant Khandelwal, Jeremy A. Collins, Ignat Georgiev et al. · arXiv · Aug 21, 2026
Behaviour Cloning (BC) has driven remarkable progress in robot manipulation, yet it is fundamentally limited by its inability to self-improve: a policy that fails cannot learn from that failure without additional human demonstrations. Reinf…
- Explainable Transformer Models for Clinical Prediction Tasks on Structured Electronic Health RecordsJun Ni Du, Lukas Adamek, Maxim Kryukov, Flavio Dormont et al. · arXiv · Aug 20, 2026
Predictive models over structured electronic health records (EHRs) remain central to machine learning for healthcare, but few have jointly emphasized quantitative laboratory information and interpretability with respect to input medical eve…
- Transfer Learning in Nonparametric Regression with Deep ReLU NetworksJunpeng Ren, Carlos Misael Madrid Padilla, Yanzhen Chen, Oscar Hernan Madrid Padilla · arXiv · Aug 20, 2026
This paper develops a general transfer learning framework for nonparametric regression with data consisting of multiple groups. Under the assumption that groups share a common structure along with group-specific deviations in additive form,…
- On the Fragility of Self-Improving Agents: Variance, Task Order, and UnderspecificationQinyuan Ye, Yu Li, Yada Pruksachatkun, Jiaxin Zhang et al. · arXiv · Aug 18, 2026
Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank--have shown great promise in recent literature. However, the reliability aspects of these methods…
- Approximate Muon with low-rank adaptersBen Anson, Conor Houghton, Edward Milsom · arXiv · Aug 14, 2026
The Muon optimizer shows clear benefits versus alternatives when pretraining neural networks. However, it is used less frequently for parameter-efficient fine-tuning (PEFT). One potential reason is that the most common PEFT method, LoRA, do…
- MetaboLLM: a metabolomics-specialized large language model for biochemical knowledge integration and predictive metabolite graph constructionDohyun Ku, Min Gu Kwak, Francisco J. Pasquel, Jing Li · arXiv · Aug 6, 2026
Metabolomics knowledge is distributed across heterogeneous resources and remains difficult to translate into predictive representations. We developed MetaboLLM, a metabolomics-specialized large language model adapted through continual pretr…
- A Six-Dimensional Taxonomy of Post-Training Adaptation Techniques with Applications in AI GovernanceFardin Afdideh, Fernando Seoane, Farhad Abtahi · arXiv · Aug 6, 2026
Post-training adaptation has become central to modern machine learning practice and includes techniques such as retraining, fine-tuning, parameter-efficient adaptation, alignment, retrieval augmentation, model editing, unlearning, calibrati…
- MarsCast: Transfer Learning of AI Weather Foundation Models to Planetary AtmospheresM. L. Carroll, J. Li, S. D. Guzewich, G. Villanueva et al. · arXiv · Aug 5, 2026
We investigate the transferability of Earth weather foundation models to planetary atmospheres by adapting the GraphCast graph neural weather forecasting model to Mars. While GraphCast achieves state-of-the-art performance for terrestrial f…
- Enhancing VLM Reward Models Through Structure-Aware Fine-TuningPyrros Koussios, Chenhao Li, Xin Chen, Andreas Krause · arXiv · Aug 4, 2026
Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL). Recent work uses large foundation Vision-Language Models (VLMs) as reward models, computing text-observation similarity to bypass manual reward …
- Bi-semantic Chemical Embedder for Joint Representation Learning of SMILES and Natural LanguageDavid Ming Segura, Jeremy Goumaz, Joshua W. Sin, Bojana Ranković et al. · arXiv · Aug 4, 2026
Transformer models have revolutionized natural language processing (NLP), and text-based molecular representations like SMILES have successfully extended these architectures to chemistry. However, domain-adaptive pre-training often causes m…
- The Parts Are Greater Than the Sum: Automated Task Sequencing for Efficient Training of Multi-Policy LLMsJiajia Tang, Sizhe Yuen, Francisco Gomez Medina, Yali Du et al. · arXiv · Jul 31, 2026
Parameter-Efficient Fine-Tuning (PEFT) commonly adapts large language models using a single shared Low-Rank Adapter (LoRA). This shared optimization space often suffers from interference when adapting heterogeneous task sequences, leading t…
- Ordered-to-disordered transfer learning with graph neural networks for formation-energy and HOMO-LUMO gap prediction in high-entropy perovskite oxidesPanupol Untarabut, Narjes Jomaa, Sylvian Cadars, Olivier Masson et al. · arXiv · Jul 31, 2026
High-entropy perovskite oxides (HEPOs) represent a chemically complex class of materials with promising functional properties, yet their vast compositional space and, chemical/structural disorder pose significant challenge for accurate prop…
- Leveraging Transfer Learning with Class-Specific Decoders for Laparoscopic SegmentationPriya Tomar, Aditya Parikh, Christian Bauckhage, Rafet Sifa · arXiv · Jul 31, 2026
Effective multi-organ segmentation in surgical data requires learning the intricate anatomical features and alleviating the challenge of class imbalance, which results from relatively lower proportions of small and limitedly exposed structu…
- Cybersecurity Detection Classification with Reasoning-enabled Language ModelsAmol Khanna, Manu Nandan, Cristian Viorel Popa, Joan Pujol-Roig et al. · arXiv · Jul 30, 2026
A major issue in Security Operations Centers (SOCs) is alert fatigue, as the number of detections reported is more than staff can triage in a given day. Prior work prompts or fine-tunes large language models (LLMs) to emit a triage label di…
- Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?Perry Dong, Ron Polonsky, Dorsa Sadigh, Chelsea Fin · arXiv · Jul 29, 2026
Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) this raises a natural question: given a pretrained policy, should the Q-function be pretrai…
- On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust RealignmentYongjian Guo, Wanlun Ma, Lingyu Shen, Xi Xiao et al. · arXiv · Jul 29, 2026
Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that retain professio…
- Re-thinking Mammography Transfer Learning: The Dataset-Informed Transfer Learning (DITL) Framework for Breast Cancer Screening and Lesion DiagnosisAdarsh Bhandary Panambur, Siming Bayer, Andreas Maier · arXiv · Jul 28, 2026
Enhancing classification performance in mammography remains a persistent challenge across both small curated datasets and large-scale clinical cohorts. Conventional transfer learning approaches often neglect dataset-specific characteristics…
- \k{appa}-LoRA: Condition Numbers Reveal Which LoRA Matrices Worth UpdatingJianghui Wang, Silong Yong, Francesco Orabona, Marco Canini et al. · arXiv · Jul 24, 2026
Low-Rank Adaptation (LoRA) has become a widely adopted technique for efficient neural network fine-tuning, decomposing model updates into low-rank matrices. However, LoRA remains computationally costly because it updates all matrices unifor…
- Universal BCI Personalization: One API for Frozen EEG Trunks and Foundation ModelsSergey Musienko · arXiv · Jul 24, 2026
Frozen EEG encoders proliferate; per-model fine-tune defaults do not scale. We present Nimbus Personalizer: one contract encode to Bayesian head to BrainState (optional affine mid-tier) that sits on heterogeneous frozen trunks without a new…
- Online Variance Reduction for Domain Adaptation on Streaming DataAndrea Napoli · arXiv · Jul 22, 2026
This paper studies the problem of stochastic variance reduction (SVR) for the maximum mean discrepancy (MMD) and correlation alignment (CORAL) loss functions. Although various offline SVR algorithms for these losses have been proposed, thes…
- Variance-reduced Domain Adaptation using Paired SamplingAndrea Napoli · arXiv · Jul 22, 2026
Correlation alignment and the maximum mean discrepancy are two widely used distribution-matching frameworks for unsupervised domain adaptation (UDA). However, high variance in these losses has been shown to undermine their effectiveness in …