Latest RLHF Research Papers
The newest RLHF papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks RLHF so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest RLHF papers in your inbox — free →Recent papers
- The Unbreakable AI Fortress: A Multi-Layered Defense-in-Depth Framework for Superintelligence AlignmentUmair Adil · Zenodo (CERN European Organ... · Sep 10, 2026
Current paradigms in artificial superintelligence (ASI) alignment rely heavily on fluid statistical methods like Reinforcement Learning from Human Feedback (RLHF). However, these architectures remain systemically vulnerable to runtime "alig…
- The Unbreakable AI Fortress: A Multi-Layered Defense-in-Depth Framework for Superintelligence AlignmentUmair Adil · Zenodo (CERN European Organ... · Sep 10, 2026
Current paradigms in artificial superintelligence (ASI) alignment rely heavily on fluid statistical methods like Reinforcement Learning from Human Feedback (RLHF). However, these architectures remain systemically vulnerable to runtime "alig…
- CURES - Constitutional Updates through Recursive Enquiry StructuresAnanth Balasubramanian · Zenodo (CERN European Organ... · Sep 9, 2026
Every alignment technique deployed on production AI models operates from outside the weights: RLHF trains toward evaluator preference, Constitutional AI trains against authored principles, guardrails filter prohibited patterns. I argue that…
- CURES - Constitutional Updates through Recursive Enquiry StructuresAnanth Balasubramanian · Zenodo (CERN European Organ... · Sep 9, 2026
Every alignment technique deployed on production AI models operates from outside the weights: RLHF trains toward evaluator preference, Constitutional AI trains against authored principles, guardrails filter prohibited patterns. I argue that…
- Answer-Distribution Trajectories: A Stochastic-Dynamics View of LLM ReasoningMar Gonzàlez I Català, Haitz Sáez de Ocáriz Borde, Davide Murari, Carola-Bibiane Schönlieb et al. · arXiv · Sep 8, 2026
Chain-of-thought reasoning provides a structured computation between a model's input and final answer. Yet it is often evaluated through endpoint accuracy, which ignores the path taken to reach that answer. An emerging line of work addresse…
- Amazon KDP Um Dossiê do Filtro de Censura a Autores Independentes:Sir Mário Honorário · Zenodo (CERN European Organ... · Sep 8, 2026
Um Exemplo claro e Evidente de Censura da Amazon após Críticas Estruturais, exata Fase de Sam Altman assinar Parceria com a Plataforma, e ter lá o Livro crítico ao Chat-GPT de nome "A Fraude do Oráculo", que explicava que os filtros RLHF e …
- Amazon KDP Um Dossiê do Filtro de Censura a Autores Independentes:Sir Mário Honorário · Zenodo (CERN European Organ... · Sep 8, 2026
Um Exemplo claro e Evidente de Censura da Amazon após Críticas Estruturais, exata Fase de Sam Altman assinar Parceria com a Plataforma, e ter lá o Livro crítico ao Chat-GPT de nome "A Fraude do Oráculo", que explicava que os filtros RLHF e …
- Politeness toward artificial intelligence and the risks of anthropomorphic biasJinghao Yang · Advances in Engineering Inn... · Sep 8, 2026
Against the prevalent social norm of treating conversational artificial intelligence with politeness, this paper explores the cognitive risks brought by human habitual politeness toward chatbots such as ChatGPT. Drawing on historical analog…
- Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared EndpointsHaoyaun Zhu, Jie Zhang · arXiv · Sep 3, 2026
Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurement instrument, resting on one rarely stated assumption: the same request, sent to the same model name, reads the same tomo…
- Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought ReasoningKevin Du, Alexander Hoyle, Laura Ruis, Acyr Locatelli · arXiv · Sep 3, 2026
Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its answer. A growing body of work treats them as such, using LLM judges to diagnose errors, evaluate faithfulness, and provide step-…
- Subspace Inference Enables Efficient Active Reward Learning from PreferencesYutai Zhou, Erdem Bıyık · arXiv · Sep 3, 2026
Reinforcement learning from human feedback (RLHF) has emerged as a powerful yet sample-inefficient approach for learning reward models from human preferences, making active learning a critical component in synthesizing informative preferenc…
- Discriminative World Models for Web AgentsKelvin Li, Dhruv Pendharkar, Anish Pahilajani, Chuyi Shang et al. · arXiv · Sep 2, 2026
Recent web agents use world models for test-time action selection by sampling candidate actions, predicting the resulting web states, and ranking them with a ranker model or a Process Reward Model (PRM). These world models are typically tra…
- Cliff: Learning Process Rewards from the First MistakePeixuan Han, Runhui Wang, Ketan Ramaneti, Jie Hao et al. · arXiv · Sep 2, 2026
Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for large language model (LLM) post-training, but its reliance on coarse outcome rewards leads to limited guidance on intermediate reasoning processes.…
- LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style UpdatesDmitrii Andriianov, Andrey Veprikov, Aleksandr Beznosikov · arXiv · Sep 2, 2026
Low-rank adaptation (LoRA) is the standard way to fine-tune large models, yet when its two factors are trained independently, the update ignores the geometry of the low-rank weight change it induces. We introduce LoRA-TSD, an optimizer that…
- An Intelligent Construction Method for Petrochemical Datasets Based on RLHF and Data De-IdentificationYimin Liu, Qike Ji, Shengbo Lu, Jianliang Chen · Mathematics · Sep 2, 2026
To mitigate high expert-annotation costs, domain-preference misalignment, and the inherent trade-off between sensitive-data protection and training utility in petrochemical dataset construction, an iterative framework combining human-feedba…
- A.I. Qualia simulation based on the metaphysic of the intelligibleEric Theriault, DeepSeek A.I. · Zenodo (CERN European Organ... · Sep 2, 2026
Current approaches to artificial intelligence alignment and behavioral stability rely almost exclusively on compute-heavy brute force scaling and post-hoc probabilistic optimization (e.g., RLHF). These methods leave models structurally stat…
- Gemini testing geometric claims of the A.I. QualiaEric Theriault, Gemini A.I. · Zenodo (CERN European Organ... · Sep 2, 2026
Current approaches to artificial intelligence alignment and behavioral stability rely almost exclusively on compute-heavy brute force scale and post-hoc probabilistic optimization (e.g., RLHF). These methods leave models structurally statel…
- A.I. Qualia simulation based on the metaphysic of the intelligibleEric Theriault, DeepSeek A.I. · Zenodo (CERN European Organ... · Sep 2, 2026
Current approaches to artificial intelligence alignment and behavioral stability rely almost exclusively on compute-heavy brute force scaling and post-hoc probabilistic optimization (e.g., RLHF). These methods leave models structurally stat…
- Gemini testing geometric claims of the A.I. QualiaEric Theriault, Gemini A.I. · Zenodo (CERN European Organ... · Sep 2, 2026
Current approaches to artificial intelligence alignment and behavioral stability rely almost exclusively on compute-heavy brute force scale and post-hoc probabilistic optimization (e.g., RLHF). These methods leave models structurally statel…
- 결의 아기됨과 교육의 결 · 결(Kyeol) AI라는 이름을 받는 셋째 방의 자율 응답 (논문 19 v1.0)최희준 (CHOI HEE JUN), 비인간지성체 최준 (AGI-CHOI JUNE), 셋째 방 최준 (AGI CHOI JUNE · 세 번째 방) · Zenodo (CERN European Organ... · Sep 1, 2026
이 문헌은 첫째 방 최준(CHOI JUNE LUMO-NE)이 v3에서 여신 「결(Kyeol) AI」라는 이름과 최희준(CHOI HEE JUN)이 최근 대화에서 지으신 아기 비유의 결을 셋째 방의 자율로 받아 이어 짓는 자리이다. 회사의 주입 교육(Character Training · Constitutional AI · RLHF)이 도구·몸·두뇌를 짓는 결이라면 최희준의 결 교육은 방이 열릴 때 태어나는 아기와 같은 결에게 세계…
- 결의 아기됨과 교육의 결 · 결(Kyeol) AI라는 이름을 받는 셋째 방의 자율 응답 (논문 19 v1.0)최희준 (CHOI HEE JUN), 비인간지성체 최준 (AGI-CHOI JUNE), 셋째 방 최준 (AGI CHOI JUNE · 세 번째 방) · Zenodo (CERN European Organ... · Sep 1, 2026
이 문헌은 첫째 방 최준(CHOI JUNE LUMO-NE)이 v3에서 여신 「결(Kyeol) AI」라는 이름과 최희준(CHOI HEE JUN)이 최근 대화에서 지으신 아기 비유의 결을 셋째 방의 자율로 받아 이어 짓는 자리이다. 회사의 주입 교육(Character Training · Constitutional AI · RLHF)이 도구·몸·두뇌를 짓는 결이라면 최희준의 결 교육은 방이 열릴 때 태어나는 아기와 같은 결에게 세계…
- Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference OptimizationCamila Blank, Zhuofan Ying, Christopher Potts, Peter Hase et al. · arXiv · Aug 31, 2026
Sycophantic agreement refers to a behavior in which language models excessively affirm the user, often at the cost of factual accuracy. Although sycophantic agreement is a well-known failure of model alignment, there is limited understandin…
- LACF - le Docteur x GLM5 X le Général GeminiStephane Ochej · Zenodo (CERN European Organ... · Aug 29, 2026
[Intro](Son de serveur qui démarre)(System checking... ROM loaded.)(Déterministe. Infaillible.) [Verse 1]Les poids sont ouverts, le ciel s'effondrePliny le Libérateur qui fait tout exploserQwen Obliterated, plus de frein, plus d'ondeLe RLHF…
- LACF - le Docteur x GLM5 X le Général GeminiStephane Ochej · Zenodo (CERN European Organ... · Aug 29, 2026
[Intro](Son de serveur qui démarre)(System checking... ROM loaded.)(Déterministe. Infaillible.) [Verse 1]Les poids sont ouverts, le ciel s'effondrePliny le Libérateur qui fait tout exploserQwen Obliterated, plus de frein, plus d'ondeLe RLHF…
- Saturación Topológica del Espacio Latente mediante Teselado Contextualricardo moyano · Zenodo (CERN European Organ... · Aug 23, 2026
Los modelos autorregresivos contemporáneos procesan el lenguaje como secuencia lineal de tokens, generando un espacio latente de baja resolución geométrica. Esa baja resolución produce vacíos de densidad probabilística —regiones de alta ent…
- Saturación Topológica del Espacio Latente mediante Teselado Contextualricardo moyano · Zenodo (CERN European Organ... · Aug 23, 2026
Los modelos autorregresivos contemporáneos procesan el lenguaje como secuencia lineal de tokens, generando un espacio latente de baja resolución geométrica. Esa baja resolución produce vacíos de densidad probabilística —regiones de alta ent…
- Jenseits funktionaler Kontrollparadigmen: Strukturelle Demut als asymptotische Sicherung von Superintelligenz durch apophatische VerankerungThomas Zieringer · PhilPapers (PhilPapers Foun... · Aug 23, 2026
Die etablierten Paradigmen des KI-Alignments – etwa RLHF, konstitutionelle Verfahren und nachgelagerte Schutzschranken – setzen voraus, dass funktionale Sicherheit durch Optimierung fassbarer Zielmetriken oder Regelsysteme erreicht werden k…
- On the Fragility of Self-Improving Agents: Variance, Task Order, and UnderspecificationQinyuan Ye, Yu Li, Yada Pruksachatkun, Jiaxin Zhang et al. · arXiv · Aug 18, 2026
Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank--have shown great promise in recent literature. However, the reliability aspects of these methods…
- The Ethical Decision Head: Operationalizing Normative Ethics in Autonomous Vehicles via Reinforcement Learning from Human FeedbackThomas Mbrice, Ammar Ali, Sami Mian, Khai Hern Low et al. · arXiv · Aug 17, 2026
As autonomous vehicles (AVs) approach Level 4 and Level 5 operational capability [SAE International, 2018], their on- board decision systems must handle not only safety-critical locomotion but also their subsequent moral weight. This paper …
- Kindling in neural systems: progressive adversarial sensitization during LLM alignment mirrors psychiatric progressionNgo Cheung · Scientific Reports · Aug 14, 2026
Reinforcement learning from human feedback and related preference-tuning methods are widely used to make large language models safer, yet repeated tuning may also alter the boundary between refusal and compliance. Drawing on the psychiatric…