Latest World Models Research Papers
The newest World Models papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks World Models so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest World Models papers in your inbox — free →Recent papers
- Semigroup-JEPA: Latent Dynamics Consistency for Zero-Shot Physics GeneralizationAndy Zeyi Liu, Haoran Sun, Lucas Baker, Randall Balestriero et al. · arXiv · Sep 9, 2026
Joint-Embedding Predictive Architecture (JEPA) world models learn a compact latent representation of the world that supports prediction and planning, but their capability to learn physics and generate physically realistic dynamics remains h…
- Discriminative World Models for Web AgentsKelvin Li, Dhruv Pendharkar, Anish Pahilajani, Chuyi Shang et al. · arXiv · Sep 2, 2026
Recent web agents use world models for test-time action selection by sampling candidate actions, predicting the resulting web states, and ranking them with a ranker model or a Process Reward Model (PRM). These world models are typically tra…
- Dutch Books for Language ModelsIsaiah Andrews, Suproteem Sarkar · arXiv · Sep 2, 2026
People increasingly use language models to support life decisions. Many such decisions involve a probabilistic forecast: How likely is a major life event, a natural disaster, or an economic outcome? Users of language models may implicitly t…
- LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style UpdatesDmitrii Andriianov, Andrey Veprikov, Aleksandr Beznosikov · arXiv · Sep 2, 2026
Low-rank adaptation (LoRA) is the standard way to fine-tune large models, yet when its two factors are trained independently, the update ignores the geometry of the low-rank weight change it induces. We introduce LoRA-TSD, an optimizer that…
- An Enclosed Mode Is a Gauge Choice: Topology Relative to Reach in Certified Code World ModelsJavier Aguilar Martín · arXiv · Aug 28, 2026
A code world model accepted by a sampling gate can be exactly right on everything the gate can see and arbitrarily wrong beyond it. We characterize what a certified model can know, and what its errors can cost, when the omission is an annul…
- Explainable Transformer Models for Clinical Prediction Tasks on Structured Electronic Health RecordsJun Ni Du, Lukas Adamek, Maxim Kryukov, Flavio Dormont et al. · arXiv · Aug 20, 2026
Predictive models over structured electronic health records (EHRs) remain central to machine learning for healthcare, but few have jointly emphasized quantitative laboratory information and interpretability with respect to input medical eve…
- On the Fragility of Self-Improving Agents: Variance, Task Order, and UnderspecificationQinyuan Ye, Yu Li, Yada Pruksachatkun, Jiaxin Zhang et al. · arXiv · Aug 18, 2026
Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank--have shown great promise in recent literature. However, the reliability aspects of these methods…
- Towards Zero-Shot Task Transfer with Neurosymbolic World ModelsIsidoro Tamassia, Lennert De Smet, Giuseppe Marra · arXiv · Aug 18, 2026
State-of-the-art model-based reinforcement learning methods learn neural world models that allow policy improvement by planning in a latent space, without assumptions on the structure of the underlying environment. While expressive, these m…
- An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World ModelsJavier Aguilar Martín · arXiv · Aug 18, 2026
In the Code World Model paradigm an LLM synthesizes an executable world model that a classical planner searches, and the model is accepted when it reproduces sampled transitions. We ask what that acceptance certifies in continuous control. …
- CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?Jonathan Sadeghi, Jenny Seidenschwarz, Jesse Allardice, Sirish Srinivasan et al. · arXiv · Aug 17, 2026
Video world models approximate the stochastic distribution of physical outcomes through generative sampling, but existing benchmarks score individual generations or compare distributions coarsely over a whole dataset, leaving the fine-grain…
- Intervention-Aware Clinical World Model for Post-Op Outcome Forecasting in CardiologyYunsung Chung, Yingshuo Liu, Abboud F. Hassan, Han Feng et al. · arXiv · Aug 13, 2026
Many clinical prediction models treat post-intervention outcomes as a one-step mapping from baseline measurements to a future endpoint. However, recovery after a procedure often unfolds as an irregular trajectory: clinical observations, med…
- ScreenShot: A Foundation Model for Few-Shot Combination Drug ScreeningAntoine de Mathelin, Christopher Tosh, Wesley Tansey · arXiv · Aug 12, 2026
Treating patients with combinations of drugs reduces the risk of resistance to any individual drug. Finding effective combinations is difficult because the large search space makes combinatorial screens prohibitively expensive, time consumi…
- Better Slots, Better Worlds: Representation Quality & Robustness in Object-Centric World ModelsShukrullo Nazirjonov, Sai Prasanna, Anna Manasyan, Georg Martius · arXiv · Aug 12, 2026
Learning world models from offline trajectories enables agents to accomplish different tasks through planning. Object-centric (OC) representations, which decompose a scene into a set of slots that bind to its objects, have been proposed as …
- Beyond Myopic World Models: Long-Horizon End-to-End Training for Direct Future PredictionXinyi Li, Zaishuo Xia, Chenjie Hao, Yubei Chen · arXiv · Aug 7, 2026
World models are expected to support imagination over extended temporal horizons, yet most are still trained through local few-step prediction objectives and deployed by recursively rolling out their own predictions. This creates a fundamen…
- Addressable Memory for Video World ModelsXindi Wu, Sven Elflein, James Lucas, Olga Russakovsky et al. · arXiv · Aug 7, 2026
We study visual persistence in interactive video world models. These models rely on a Key-Value (KV) cache as a growing visual memory to carry forward previously generated frames. However, we find that models can no longer reliably address …
- Reinformed Dreamer: An Asymmetric World Model Efficiently Trained through Latent GuidanceGaspard Lambrechts, Adrien Bolland, Daniel Ebi, Damien Ernst · arXiv · Jul 28, 2026
Much like humans benefit from guidance while learning, reinforcement learning algorithms may benefit from additional supervision beyond rewards. Leveraging additional information during training to learn better representations and behaviors…
- The balance between compactness and forecast accuracy of data-driven latent-space reduced-order models in controlled wake flowsAlberto Solera-Rico, Patricia García-Caspueñas, Carlos Sanmiguel Vila, Stefano Discetti · arXiv · Jul 27, 2026
Model-based active flow control requires predictive models that are accurate, stable, and fast enough for real-time optimisation. In controlled wake flows, this is often achieved through Reduced-Order Models (ROMs) that first compress high-…
- On the Identifiability of Controlled World ModelsXiangteng Zhang, Yang Guan, Bo Zhang, Ya-Qin Zhang et al. · arXiv · Jul 24, 2026
Learning world models that infer environment dynamics from high-dimensional observations and predict outcomes under candidate actions is central to planning and control. Joint-Embedding Predictive Architectures (JEPAs) provide a compelling …
- Concept-Guided Spatial Regularization for World Models in Atari PongYukuan Lu, Zaishuo Xia, Weyl Lu, Yubei Chen · arXiv · Jul 16, 2026
World models are usually evaluated as components of model-based reinforcement learning (MBRL) systems, while the world models themselves are rarely studied in isolation. We examine five representative visual world-model agents in Atari Pong…
- DriftWorld: Fast World Modeling through DriftingSusie Lu, Haonan Chen, Weirui Ye, Yilun Du · arXiv · Jul 16, 2026
Predictive world models enable robots to plan by imagining the outcomes of their actions, but their value for control hinges on generating many rollouts quickly. This creates a bottleneck for diffusion-based world models: multistep sampling…
- World Action Verifier: Self-Improving World Models via Forward-Inverse AsymmetryYuejiang Liu, Fan Feng, Lingjing Kong, Weifeng Lu et al. · ICLR 2026 Workshop World Models · Mar 2, 2026
General-purpose world models promise scalable policy evaluation, optimization, and planning, yet achieving the required level of robustness remains challenging. Unlike policy learning which primarily focuses on optimal actions, a world mode…
- Consistent Video World Model With Geometry-Aware Rotary Position EmbeddingChendong Xiang, Jiajun Liu, Jintao Zhang, Xiao Yang et al. · ICLR 2026 Workshop World Models · Mar 2, 2026
Predictive world models that simulate future observations under explicit camera control are fundamental to interactive AI. Despite rapid advances, current systems lack spatial persistence: they fail to maintain stable scene structures over …
- Computer-Using World ModelYiming Guan, Rui Yu, John Zhang, Lu Wang et al. · ICLR 2026 Workshop World Models · Mar 2, 2026
Agents operating in complex software environments benefit from reasoning about the consequences of their actions, as even a single incorrect user interface (UI) operation can derail long, artifact-preserving workflows. This challenge is par…
- Action Shapley: A training data selection metric for Training World Models for Reinforcement LearningRajat Ghosh, Debojyoti Dutta · ICLR 2026 Workshop World Models · Mar 2, 2026
World models are central to model-based reinforcement learning, enabling agents to predict environment dynamics and reason about future outcomes. In real-world settings, however, training high-fidelity world models is often constrained by l…
- Ctrl-World: A Controllable Generative World Model for Robot ManipulationYanjiang Guo, Lucy Xiaoyang Shi, Jianyu Chen, Chelsea Finn · ICLR 2026 Workshop World Models · Mar 2, 2026
Generalist robot policies can now perform a wide range of manipulation skills, but evaluating and improving their ability with unfamiliar objects and instructions remains a significant challenge. Rigorous evaluation requires a large number …
- Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action ModelsZhilong Zhang, Haoxiang Ren, Yihao Sun, Yifei Sheng et al. · ICLR 2026 Workshop World Models · Mar 2, 2026
Vision-Language-Action (VLA) models show strong generalization for robotic control, but finetuning them with reinforcement learning (RL) is constrained by the high cost and safety risks of real-world interaction. Training VLA models in inte…
- Ctrl-World: A Controllable Generative World Model for Robot ManipulationYanjiang Guo, Lucy Xiaoyang Shi, Jianyu Chen, Chelsea Finn · ICLR 2026 Poster · Jan 26, 2026
Generalist robot policies can now perform a wide range of manipulation skills, but evaluating and improving their ability with unfamiliar objects and instructions remains a significant challenge. Rigorous evaluation requires a large number …
- Motus: A Unified Latent Action World ModelHongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyu Wang et al. · arXiv.org · Dec 15, 2025
While a general embodied agent must function as a unified system, current methods are built on isolated models for understanding, world modeling, and control. This fragmentation prevents unifying multimodal generative capabilities and hinde…
- RELIC: Interactive Video World Model with Long-Horizon MemoryYicong Hong, Yiqun Mei, Chongjian Ge, Yiran Xu et al. · arXiv.org · Dec 3, 2025
A truly interactive world model requires three key ingredients: real-time long-horizon streaming, consistent spatial memory, and precise user control. However, most existing approaches address only one of these aspects in isolation, as achi…
- RynnVLA-002: A Unified Vision-Language-Action and World ModelJun Cen, Siteng Huang, Yuqian Yuan, Kehan Li et al. · arXiv.org · Nov 21, 2025
We introduce RynnVLA-002, a unified Vision-Language-Action (VLA) and world model. The world model leverages action and visual inputs to predict future image states, learning the underlying physics of the environment to refine action generat…