A comprehensive list of papers for the definition of World Models and using World Models for General Video Generation, Embodied AI, and Autonomous Driving, including papers, codes, and related websites.
set. de 25out. de 25dez. de 25fev. de 26abr. de 26mai. de 26jul. de 26
ArtefatosPyPIpip install awesome-world-models
README
Awesome World Models for Robotics
This repository provides a curated list of papers for World Models for General Video Generation, Embodied AI, and Autonomous Driving. Template from Awesome-LLM-Robotics and Awesome-World-Model
Contributions are welcome! Please feel free to submit pull requests or reach out via email to add papers!
If you find this repository useful, please consider citing and giving this list a star ⭐. Feel free to share it with others!
Wayve, Introducing GAIA-1: A Cutting-Edge Generative AI Model for Autonomy. [Paper] [Blog]
Yann LeCun, A Path Towards Autonomous Machine Intelligence. [Paper]
Survey
"From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence", arxiv 2026.07. [Paper]
"From World Models to World Action Models: A Concise Tutorial for Robotics", arxiv 2026.07. [Paper]
"Autonomous Video Generation with Counterfactual Controllability for Self-Evolving World Models", arxiv 2026.06. [Paper]
"Medical world models: representing medical states, modelling clinical dynamics and guiding intervention policies", arxiv 2026.06. [Paper]
"Bridging the Agent-World Gap: Text World Models for LLM-based Agents", arxiv 2026.06. [Paper]
"Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends", arxiv 2026.05. [Paper]
"Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses" arXiv 2026.05. [Paper] [Website]
"Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform" arXiv 2026.05. [Paper]
"World Action Models: The Next Frontier in Embodied AI" arXiv 2026.05. [Paper]
"Latent State Design for World Models under Sufficiency Constraints" arXiv 2026.05. [Paper]
"World Model for Robot Learning: A Comprehensive Survey" arXiv 2026.05. [Paper] [Website] [Code]
"Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling" arXiv 2026.04. [Paper] [Code]
"Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond" arXiv 2026.04. [Paper]
"Infrastructure-Centric World Models: Bridging Temporal Depth and Spatial Breadth for Roadside Perception" arXiv 2026.04. [Paper]
"Human Cognition in Machines: A Unified Perspective of World Models" arXiv 2026.04. [Paper]
"Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms" arXiv 2026.03. [Paper]
"From Digital Twins to World Models:Opportunities, Challenges, and Applications for Mobile Edge General Intelligence" arXiv 2026.03. [Paper]
"The Trinity of Consistency as a Defining Principle for General World Models" arXiv 2026.02. [Paper] [Code]
"A Mechanistic View on Video Generation as World Models: State and Dynamics", arXiv 2026.01. [Paper]
"From Generative Engines to Actionable Simulators: The Imperative of Physical Grounding in World Models", arXiv 2026.01. [Paper]
"Modeling the Mental World for Embodied AI: A Comprehensive Review", arXiv 2026.01. [Paper]
"Digital Twin AI: Opportunities and Challenges from Large Language Models to World Models", arXiv 2026.01. [Paper]
"Beyond World Models: Rethinking Understanding in AI Models", AAAI 2026. [Paper]
"Simulating the Visual World with Artificial Intelligence: A Roadmap", arXiv 2025.11. [Paper] [Wesite] [Code]
"A Step Toward World Models: A Survey on Robotic Manipulation", arXiv 2025.11. [Paper]
"World Models Should Prioritize the Unification of Physical and Social Dynamics", NIPS 2025. [Paper] [Website]
"From Masks to Worlds: A Hitchhiker's Guide to World Models", arXiv 2025.10. [Paper] [Website]
"A Comprehensive Survey on World Models for Embodied AI", arXiv 2025.10. [Paper] [Website]
"The Safety Challenge of World Models for Embodied AI Agents: A Review", arXiv 2025.10. [Paper]
"Embodied AI: From LLMs to World Models", IEEE CASM. [Paper]
"3D and 4D World Modeling: A Survey", arXiv 2025.09. [Paper]
"Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges", arXiv 2025.08. [Paper]
"A Survey: Learning Embodied Intelligence from Physical Simulators and World Models", arXiv 2025.07. [Paper] [Code]
"Embodied AI Agents: Modeling the World", arXiv 2025.06. [Paper]
"From 2D to 3D Cognition: A Brief Survey of General World Models", arXiv 2025.06. [Paper]
"A Survey on World Models Grounded in Acoustic Physical Information", arXiv 2025.06. [Paper]
"Exploring the Evolution of Physics Cognition in Video Generation: A Survey", arXiv 2025.03. [Paper] [Code]
"World Models in Artificial Intelligence: Sensing, Learning, and Reasoning Like a Child", arXiv 2025.03. [Paper]
"Simulating the Real World: A Unified Survey of Multimodal Generative Models", arXiv 2025.03. [Paper] [Code]
"Four Principles for Physically Interpretable World Models", arXiv 2025.03. [Paper]
"The Role of World Models in Shaping Autonomous Driving: A Comprehensive Survey", arXiv 2025.02. [Paper] [Code]
"A Survey of World Models for Autonomous Driving", TPAMI. [Paper]
"Understanding World or Predicting Future? A Comprehensive Survey of World Models", arXiv 2024.11. [Paper]
"World Models: The Safety Perspective", ISSRE WDMD. [Paper]
"Exploring the Interplay Between Video Generation and World Models in Autonomous Driving: A Survey", arXiv 2024.11. [Paper]
"From Efficient Multimodal Models to World Models: A Survey", arXiv 2024.07. [Paper]
"Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI", arXiv 2024.07. [Paper] [Code]
"Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond", arXiv 2024.05. [Paper] [Code]
"World Models for Autonomous Driving: An Initial Survey", TIV. [Paper]
"A survey on multimodal large language models for autonomous driving", WACVW 2024. [Paper] [Code]
Datasets & Benchmarks & Evaluation
WorldRoamBench: "WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models", arxiv 2026.06. [Paper]
"A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models", arxiv 2026.06. [Paper]
PhysEditWorld: "PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models", arxiv 2026.06. [Paper]
"Current World Models Lack a Persistent State Core", arxiv 2026.06. [Paper]
EgoCS-400K: "EgoCS-400K: An Egocentric Gameplay Dataset for World Models", arxiv 2026.06. [Paper]
ARB4WM: "ARB4WM: An Adversarial Robustness Benchmark for World Models in Continuous Control", arxiv 2026.06. [Paper]
"Diffusion Transformer World-Action Model for AV Scene Prediction", arxiv 2026.06. [Paper]
"World Model Self-Distillation: Training World Models to Solve General Tasks", arxiv 2026.06. [Paper]
WorldOlympiad: "WorldOlympiad: Can Your World Model Survive a Triathlon?", arxiv 2026.06. [Paper]
"Can Image Models Imagine Time? ImageTime: A Novel Benchmark for Probing Visual World Modeling Through Spatiotemporal Consistency", arxiv 2026.06. [Paper]
"Towards World Models in Biomedical Research", arxiv 2026.06. [Paper]
RoboTrustBench: "RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation", arxiv 2026.06. [Paper] [Website]
MBench: "MBench: A Comprehensive Benchmark on Memory Capability for Video World Models", arxiv 2026.05. [Paper] [Website] [Code]
MiraBench: "MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models", arxiv 2026.05. [Paper]
What-If World: "What-If World: A Causal Benchmark for General World Models in Embodied Scenarios", arxiv 2026.05. [Paper]
WBench: "WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation", arxiv 2026.05. [Paper] [Website]
stable-worldmodel: "stable-worldmodel: A Platform for Reproducible World Modeling Research and Evaluation", arxiv 2026.05. [Paper]
HalluWorld: "HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models", arxiv 2026.05. [Paper] [Website]
WorldArena 2.0: "WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform", arxiv 2026.05. [Paper] [Website]
DeTrack: "DeTrack: A Benchmark and Altitude-Aware Dual World Model for Drone-embodied Tracking", arxiv 2026.05. [Paper]
"Quantitative Video World Model Evaluation for Geometric-Consistency", arxiv 2026.05. [Paper] [Website]
"Is Your Driving World Model an All-Around Player?", arxiv 2026.05. [Paper] [Website] [Code]
PhyGround: "PhyGround: Benchmarking Physical Reasoning in Generative World Models", arxiv 2026.05. [Paper] [Website]
iWorld-Bench: "iWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation Framework", ICML 2026. [Paper]
WorldMark: "WorldMark: A Unified Benchmark Suite for Interactive Video World Models", arxiv 2026.04. [Paper]
RoboWM-Bench: "RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation", arxiv 2026.04. [Paper] [Code]
MotionScape: "MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models", arxiv 2026.04. [Paper] [Code]
EgoVerse: "EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World", arxiv 2026.04. [Paper] [Website]
Omni-WorldBench: "Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models", arxiv 2026.03. [Paper]
"Out of Sight, Out of Mind? Evaluating State Evolution in Video World Models", arxiv 2026.03. [Paper] [Website]
MicroVerse: "MicroVerse: A Preliminary Exploration Toward a Micro-World Simulation", ICLR 2026. [Paper] [Code]
WorldArena: "WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models" arXiv 2026.02. [Paper] [Code]
MIND: "MIND: Benchmarking Memory Consistency and Action Control in World Models" arXiv 2026.02. [Paper] [Code]
WoW-bench: "World of Workflows: a Benchmark for Bringing World Models to Enterprise Systems" arXiv 2026.01. [Paper]
WorldBench: "WorldBench: Disambiguating Physics for Diagnostic Evaluation of World Models" arXiv 2026.01. [Paper] [Website]
PhysicsMind: "PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models" arXiv 2026.01. [Paper]
RBench: "Rethinking Video Generation Model for the Embodied World" arXiv 2026.01. [Paper] [Website] [Code]
Wow, wo, val!: "Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Test" arXiv 2026.01. [Paper]
DrivingGen: "DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving" arXiv 2026.01. [Paper] [Website]
"A Unified Definition of Hallucination, Or: It's the World Model, Stupid" arXiv 2025.12. [Paper]
"Active Intelligence in Video Avatars via Closed-loop World Modeling" arXiv 2025.12. [Paper] [Code]
MobileWorldBench: "MobileWorldBench: Towards Semantic World Modeling For Mobile Agents" arXiv 2025.12. [Paper] [Code]
WorldLens: "WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World" arXiv 2025.12. [Paper] [Website]
"Evaluating Gemini Robotics Policies in a Veo World Simulator" arXiv 2025.12. [Paper]
On Memory: "On Memory: A comparison of memory mechanisms in world models" World Modeling Workshop 2026. [Paper]
SmallWorlds: "SmallWorlds: Assessing Dynamics Understanding of World Models in Isolated Environments", arXiv 2025.11. [Paper]
4DWorldBench: "4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models", arXiv 2025.11. [Paper]
Target-Bench: "Target-Bench: Can World Models Achieve Mapless Path Planning with Semantic Targets?", arXiv 2025.11. [Paper]
PragWorld: "PragWorld: A Benchmark Evaluating LLMs' Local World Model under Minimal Linguistic Alterations and Conversational Dynamics", AAAI 2026. [Paper]
"Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark", arXiv 2025.11. [Paper] [Code]
"Scalable Policy Evaluation with Video World Models", arXiv 2025.11. [Paper]
"Expert Evaluation of LLM World Models: A High-Tc Superconductivity Case Study", ICML 2025 workshop on Assessing World Models and the Explorations in AI Today. [Paper]
LikePhys: "LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference", ICLR 2026. [Paper] [Website]
World-in-World: "World-in-World: World Models in a Closed-Loop World", arXiv 2025.10. [Paper] [Website]
VideoVerse: "VideoVerse: How Far is Your T2V Generator from a World Model?", arXiv 2025.10. [Paper]
OmniWorld: "OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling", arXiv 2025.09. [Paper] [Website]
"Beyond Simulation: Benchmarking World Models for Planning and Causality in Autonomous Driving", ICRA 2025. [Paper]
WM-ABench: "Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation", ACL 2025(Findings). [Paper] [Website]
UNIVERSE: "Adapting Vision-Language Models for Evaluating World Models", arxiv 2025.06. [Paper]
WorldPrediction: "WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning", arxiv 2025.06. [Paper]
"Toward Memory-Aided World Models: Benchmarking via Spatial Consistency", arxiv 2025.05. [Paper] [Datasets] [Code]
SimWorld: "SimWorld: A Unified Benchmark for Simulator-Conditioned Scene
Generation via World Model", arxiv 2025.05. [Paper] [Code]
EWMBench: "EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models", arxiv 2025.05. [Paper] [Code]
"Toward Stable World Models: Measuring and Addressing World Instability in Generative Environments", arxiv 2025.03. [Paper]
WorldModelBench: "WorldModelBench: Judging Video Generation Models As World Models", CVPR 2025. [Paper] [Website]
Text2World: "Text2World: Benchmarking Large Language Models for Symbolic World Model Generation", arxiv 2025.02. [Paper] [Website]
ACT-Bench: "ACT-Bench: Towards Action Controllable World Models for Autonomous Driving", arxiv 2024.12. [Paper]
WorldSimBench: "WorldSimBench: Towards Video Generation Models as World Simulators", arxiv 2024.10. [Paper] [Website]
EVA: "EVA: An Embodied World Model for Future Video Anticipation", ICML 2025. [Paper] [Website]
AeroVerse: "AeroVerse: UAV-Agent Benchmark Suite for Simulating, Pre-training, Finetuning, and Evaluating Aerospace Embodied World Models", arxiv 2024.08. [Paper]
CityBench: "CityBench: Evaluating the Capabilities of Large Language Model as World Model", arXiv 2024.06. [Paper] [Code]
"Imagine the Unseen World: A Benchmark for Systematic Generalization in Visual World Models", NIPS 2023. [Paper]
General World Models
SAGE: "SAGE: Subgoal-Conditioned Action Generation for Latent World Model Planning", arxiv 2026.07. [Paper]
"Mobile Network Control with a World Model", arxiv 2026.07. [Paper]
"Concept-Guided Spatial Regularization for World Models in Atari Pong", arxiv 2026.07. [Paper]
DriftWorld: "DriftWorld: Fast World Modeling through Drifting", arxiv 2026.07. [Paper]
RENEW: "RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences", arxiv 2026.07. [Paper]
"From Pixels to States: Rethinking Interactive World Models as Game Engines", arxiv 2026.07. [Paper]
"Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models", arxiv 2026.07. [Paper]
Cycle-World: "Cycle-World: Mitigating Error Accumulation in Long-term Video World Models via Reverse-Prediction Cycle Consistency", arxiv 2026.07. [Paper]
ABot-3DWorld 0: "ABot-3DWorld 0: A Universal World Model to Explore Any 3D Space", arxiv 2026.07. [Paper]
"A Control Theory of Predictability in Latent World Models", arxiv 2026.07. [Paper]
RetailSMV: "RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail", arxiv 2026.07. [Paper]
AdaJEPA: "AdaJEPA: An Adaptive Latent World Model", arxiv 2026.06. [Paper]
MemLearner: "MemLearner: Learning to Query Context memory for Video World Models", arxiv 2026.06. [Paper]
Delta-JEPA: "Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding", arxiv 2026.06. [Paper]
DreamForge-World 0.1: "DreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World Model", arxiv 2026.06. [Paper]
"Prototype Latent World Model Replay for Class-Incremental Learning", arxiv 2026.06. [Paper]
"Flow Matching in Feature Space for Stochastic World Modeling", arxiv 2026.06. [Paper]
"Hallucination in World Models is Predictable and Preventable", arxiv 2026.06. [Paper]
EO-WM: "EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting", arxiv 2026.06. [Paper]
"A Generalization Theory for JEPA-Based World Models", arxiv 2026.06. [Paper]
LithoDreamer: "LithoDreamer: A Physics-Informed World Model for Multi-Stage Computational Lithography", arxiv 2026.06. [Paper]
Causal-rCM: "Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models", arxiv 2026.06. [Paper]
"Compression and Retrieval: Implicit Memory Retrieval for Video World Models", arxiv 2026.06. [Paper]
"Reference-Free Assessment of Physical Consistency in World Model-based Video Generation", arxiv 2026.06. [Paper]
"Sensorimotor World Models: Perception for Action via Inverse Dynamics", arxiv 2026.06. [Paper]
Holo-World: "Holo-World: Unified Camera, Object and Weather Control for Video World Model", arxiv 2026.06. [Paper]
SurgVista: "SurgVista: Long-Horizon Surgical World Modeling with Plausible Instrument-Tissue Dynamics", arxiv 2026.06. [Paper]
DreamReg: "DreamReg: Belief-Driven World Model for 2D-3D Ultrasound Registration", arxiv 2026.06. [Paper]
MaineCoon: "MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model", arxiv 2026.06. [Paper]
ActWorld: "ActWorld: From Explorable to Interactive World Model via Action-Aware Memory", arxiv 2026.06. [Paper]
DreamX-World 1.0: "DreamX-World 1.0: A General-Purpose Interactive World Model", arxiv 2026.06. [Paper]
BadWorld: "BadWorld: Adversarial Attacks on World Models", arxiv 2026.06. [Paper]
"Slots, Transitions, Loops: Learning Composable World Models for ARC", arxiv 2026.06. [Paper]
BiWM: "BiWM: Advancing Open-Source Interactive Video World Models with Bidirectional Autoregression", arxiv 2026.06. [Paper]
"Latent Spatial Memory for Video World Models", arxiv 2026.06. [Paper]
Echo-Memory: "Echo-Memory: A Controlled Study of Memory in Action World Models", arxiv 2026.06. [Paper]
Prisma-World: "Prisma-World: Camera-Controllable Multi-Agent Video World Model", arxiv 2026.06. [Paper]
FF-JEPA: "FF-JEPA: Long-Horizon Planning in World Models with Latent Planners", arxiv 2026.06. [Paper]
AnchorWorld: "AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization", arxiv 2026.06. [Paper]
"A 3D Isovist World Model -- Revealing a City's Unseen Geometry and Its Emergent Cross-City Signature", arxiv 2026.06. [Paper]
"World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning", arxiv 2026.06. [Paper] [Code]
MetaWorld: "MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data", arxiv 2026.06. [Paper]
"From Zero to Hero: Training-Free Custom Concept Spawning in World Models", arxiv 2026.06. [Paper]
"Geometry-Aware Implicit Memory for Video World Models", arxiv 2026.06. [Paper] [Website]
"Policy and World Modeling Co-Training for Language Agents", arxiv 2026.06. [Paper]
COMAP: "COMAP: Co-Evolving World Models and Agent Policies for LLM Agents", arxiv 2026.06. [Paper] [Code]
"Behavior-Invariant Task Representation Learning with Transformer-based World Models for Offline Meta-Reinforcement Learning", arxiv 2026.06. [Paper]
YoCausal: "YoCausal: How Far is Video Generation from World Model? A Causality Perspective", arxiv 2026.05. [Paper] [Website]
minWM: "minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models", arxiv 2026.05. [Paper] [Website]
"Thoughts-as-Planning: Latent World Models for Chain-of-Thoughts Optimization via Reinforcement Planning", arxiv 2026.05. [Paper] [Code]
Gamma-World: "Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players", arxiv 2026.05. [Paper] [Website]
Chreode: "Chreode: A Cell World Model for One-Step Temporal Dynamics and Perturbation Prediction", arxiv 2026.05. [Paper]
"When Does LeJEPA Learn a World Model?", arxiv 2026.05. [Paper]
"Back to Parsimonious Latents: Learning Task-Centric World Models from Visual Foundations", arxiv 2026.05. [Paper]
UWM-JEPA: "UWM-JEPA: Predictive World Models That Imagine in Belief Space", arxiv 2026.05. [Paper] [Website]
WorldCraft: "WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models", arxiv 2026.05. [Paper] [Website]
"Drift-Resistant Navigation World Model with Anchored Epipolar Guidance", arxiv 2026.05. [Paper]
"World Models as Group Actions", arxiv 2026.05. [Paper]
"Distilling Game Code World Model Generation into Lightweight Large Language Models", arxiv 2026.05. [Paper]
"Unified 3D Scene Understanding Through Physical World Modeling", ICLR 2026. [Paper]
"A World Model of Radiologist Reading for Medical Image Representation Learning", arxiv 2026.05. [Paper]
SCOPE: "SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models", arxiv 2026.05. [Paper] [Website] [Code]
WMAttack: "WMAttack: Automated Attack Search for Adversarial Evaluation of World-Model Agents", arxiv 2026.05. [Paper]
WorldKV: "WorldKV: Efficient World Memory with World Retrieval and Compression", arxiv 2026.05. [Paper] [Website]
"Beyond Euclidean Proximity: Repairing Latent World Models with Horizon-Matched Trajectory Reachability Metrics", arxiv 2026.05. [Paper]
ChronoMedicalWorld: "ChronoMedicalWorld: A Medical World Model for Learning Patient Trajectories from Longitudinal Care Data", arxiv 2026.05. [Paper]
"World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks", arxiv 2026.05. [Paper]
AffectVerse: "AffectVerse: Emotional World Models for Multimodal Affective Computing", arxiv 2026.05. [Paper]
FlyMirage: "FlyMirage: A Fully Automated Generation Pipeline for Diverse and Scalable UAV Flight Data via Generative World Model", arxiv 2026.05. [Paper]
PhyWorld: "PhyWorld: Physics-Faithful World Model for Video Generation", arxiv 2026.05. [Paper]
"Transformers Linearly Represent Highly Structured World Models", arxiv 2026.05. [Paper]
"Composition of Memory Experts for Diffusion World Models", ICLR 2026. [Paper]
PROWL: "PROWL: Prioritized Regret-Driven Optimization for World Model Learning", arxiv 2026.05. [Paper]
Incantation: "Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models", arxiv 2026.05. [Paper]
PH-Dreamer: "PH-Dreamer: A Physics-Driven World Model via Port-Hamiltonian Generative Dynamics", arxiv 2026.05. [Paper]
PanoWorld: "PanoWorld: A Generative Spatial World Model for Consistent Whole-House Panorama Synthesis", arxiv 2026.05. [Paper] [Website]
ECG-WM: "ECG-WM: A Physiology-Informed ECG World Model for Clinical Intervention Simulation", arxiv 2026.05. [Paper]
"Self-supervised Hierarchical Visual Reasoning with World Model", ICML 2026. [Paper] [Website]
"Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models", arxiv 2026.05. [Paper]
Map2World: "Map2World: Segment Map Conditioned Text to 3D World Generation", arxiv 2026.05. [Paper] [Website]
GWMs: "Graph World Models: Concepts, Taxonomy, and Future Directions", arxiv 2026.04. [Paper]
MultiWorld: "MultiWorld: Scalable Multi-Agent Multi-View Video World Models", arxiv 2026.04. [Paper] [Website]
"Learning Ad Hoc Network Dynamics via Graph-Structured World Models", arxiv 2026.04. [Paper]
"Do LLMs Build Spatial World Models? Evidence from Grid-World Maze Tasks", arxiv 2026.04. [Paper]
"Zero-shot World Models Are Developmentally Efficient Learners", arxiv 2026.04. [Paper]
WorldMAP: "WorldMAP: Bootstrapping Vision-Language Navigation Trajectory Prediction with Generative World Models", arxiv 2026.04. [Paper]
"CausalVAE as a Plug-in for World Models: Towards Reliable Counterfactual Dynamics", arxiv 2026.04. [Paper]
INSPATIO-WORLD: "INSPATIO-WORLD: A Real-Time 4D World Simulator via Spatiotemporal Autoregressive Modeling", arxiv 2026.04. [Paper]
Telecom World Models: "Telecom World Models: Unifying Digital Twins, Foundation Models, and Predictive Planning for 6G", arxiv 2026.04. [Paper]
InCoder-32B-Thinking: "InCoder-32B-Thinking: Industrial Code World Model for Thinking", arxiv 2026.04. [Paper]
Learn2Fold: "Learn2Fold: Structured Origami Generation with World Model Planning", arxiv 2026.03. [Paper]
WorldFlow3D: "WorldFlow3D: Flowing Through 3D Distributions for Unbounded World Generation", arxiv 2026.03. [Paper] [Website]
LOME: "LOME: Learning Human-Object Manipulation with Action-Conditioned Egocentric World Model", arxiv 2026.03. [Paper]
VGGRPO: "VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward", arxiv 2026.03. [Paper] [Website]
PiJEPA: "Policy-Guided World Model Planning for Language-Conditioned Visual Navigation", arxiv 2026.03. [Paper]
"Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models", arxiv 2026.03. [Paper]
Lingshu-Cell: "Lingshu-Cell: A generative cellular world model for transcriptome modeling toward virtual cells", arxiv 2026.03. [Paper]
AI-Supervisor: "AI-Supervisor: Autonomous AI Research Supervision via a Persistent Research World Model", arxiv 2026.03. [Paper]
WildWorld: "WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG", arxiv 2026.03. [Paper] [Website] [Code]
"Model Predictive Control with Differentiable World Models for Offline Reinforcement Learning", arxiv 2026.03. [Paper]
WorldCache: "WorldCache: Content-Aware Caching for Accelerated Video World Models", arxiv 2026.03. [Paper] [Website]
"From Part to Whole: 3D Generative World Model with an Adaptive Structural Hierarchy", ICME 2026. [Paper]
EgoForge: "EgoForge: Goal-Directed Egocentric World Simulator", arxiv 2026.03. [Paper]
"Structured Latent Dynamics in Wireless CSI via Homomorphic World Models", IEEE ICC. [Paper]
WorldAgents: "WorldAgents: Can Foundation Image Models be Agents for 3D World Models?", arxiv 2026.03. [Paper] [Website]
R2-Dreamer: "R2-Dreamer: Redundancy-Reduced World Models without Decoders or Augmentation", ICLR 2026. [Paper] [Code]
StereoWorld: "Stereo World Model: Camera-Guided Stereo Video Generation", arxiv 2026.03. [Paper] [Website]
MosaicMem: "MosaicMem: Hybrid Spatial Memory for Controllable Video World Models", arxiv 2026.03. [Paper] [Website]
WorldCam: "WorldCam: Interactive Autoregressive 3D Gaming Worlds with Camera Pose as a Unifying Geometric Representation", arxiv 2026.03. [Paper] [Website]
SWM: "Grounding World Simulation Models in a Real-World Metropolis", arxiv 2026.03. [Paper] [Website]
NavThinker: "NavThinker: Action-Conditioned World Models for Coupled Prediction and Planning in Social Navigation", arxiv 2026.03. [Paper] [Website]
EyeWorld: "EyeWorld: A Generative World Model of Ocular State and Dynamics", arxiv 2026.03. [Paper]
CtrlAttack: "CtrlAttack: A Unified Attack on World-Model Control in Diffusion Models", arxiv 2026.03. [Paper]
SAW: "SAW: Toward a Surgical Action World Model via Controllable and Scalable Video Generation", arxiv 2026.03. [Paper]
VGGT-World: "VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model", arxiv 2026.03. [Paper]
ARROW: "ARROW: Augmented Replay for RObust World models", arxiv 2026.03. [Paper]
RAE-NWM: "RAE-NWM: Navigation World Model in Dense Visual Representation Space", arxiv 2026.03. [Paper] [Code]
SPIRAL: "SPIRAL: A Closed-Loop Framework for Self-Improving Action World Models via Reflective Planning Agents", arxiv 2026.03. [Paper]
MWM: "MWM: Mobile World Models for Action-Conditioned Consistent Prediction", arxiv 2026.03. [Paper] [Website] [Code]
Brain-WM: "Brain-WM: Brain Glioblastoma World Model", arxiv 2026.03. [Paper] [Code]
DreamSAC: "DreamSAC: Learning Hamiltonian World Models via Symmetry Exploration", arxiv 2026.03. [Paper]
LiveWorld: "LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models", arxiv 2026.03. [Paper] [Website]
"What if? Emulative Simulation with World Models for Situated Reasoning", arxiv 2026.03. [Paper]
WorldCache: "WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching", arxiv 2026.03. [Paper] [Website]
"Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model", CVPR 2026. [Paper]
"Beyond Pixel Histories: World Models with Persistent 3D State", arxiv 2026.03. [Paper] [Website]
"Contextual Latent World Models for Offline Meta Reinforcement Learning", arxiv 2026.03. [Paper]
"Next Embedding Prediction Makes World Models Stronger", arxiv 2026.03. [Paper]
COMBAT: "COMBAT: Conditional World Models for Behavioral Agent Training", arxiv 2026.03. [Paper]
DreamWorld: "DreamWorld: Unified World Modeling in Video Generation", arxiv 2026.03. [Paper] [Code]
MetaOthello: "MetaOthello: A Controlled Study of Multiple World Models in Transformers", arxiv 2026.02. [Paper]
GeoWorld: "GeoWorld: Geometric World Models", CVPR 2026. [Paper] [Website]
UCM: "UCM: Unifying Camera Control and Memory with Time-aware Positional Encoding Warping for World Models", arxiv 2026.02. [Paper] [Website]
"Code World Models for Parameter Control in Evolutionary Algorithms", arxiv 2026.02. [Paper]
Solaris: "Solaris: Building a Multiplayer Video World Model in Minecraft", arxiv 2026.02. [Paper] [Website]
"MRI Contrast Enhancement Kinetics World Model", CVPR 2026. [Paper]
"Neural Fields as World Models", arXiv 2026.02. [Paper]
"Learning Invariant Visual Representations for Planning with Joint-Embedding Predictive World Models", arXiv 2026.02. [Paper]
Generated Reality: "Generated Reality: Human-centric World Simulation using Interactive Video Generation with Hand and Camera Control", arXiv 2026.02. [Paper]
VLM-DEWM: "VLM-DEWM: Dynamic External World Model for Verifiable and Resilient Vision-Language Planning in Manufacturing", arXiv 2026.02. [Paper]
"World-Model-Augmented Web Agents with Action Correction", arXiv 2026.02. [Paper]
"Cold-Start Personalization via Training-Free Priors from Structured World Models", arXiv 2026.02. [Paper] [Code]
"World Models for Policy Refinement in StarCraft II", arXiv 2026.02. [Paper]
WebWorld: "WebWorld: A Large-Scale World Model for Web Agent Training", arXiv 2026.02. [Paper]
WIMLE: "WIMLE: Uncertainty-Aware World Models with IMLE for Sample-Efficient Continuous Control", ICLR 2026. [Paper]
Causal-JEPA: "Causal-JEPA: Learning World Models through Object-Level Latent Interventions", arXiv 2026.02. [Paper] [Website] [Code]
Olaf-World: "Olaf-World: Orienting Latent Actions for Video World Modeling", arXiv 2026.02. [Paper] [Website] [Code]
Agent World Model: "Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning", arXiv 2026.02. [Paper] [Code]
WorldCompass: "WorldCompass: Reinforcement Learning for Long-Horizon World Models", arXiv 2026.02. [Paper] [Website]
"Horizon Imagination: Efficient On-Policy Training in Diffusion World Models", ICLR 2026. [Paper] [Code]
"Geometry-Aware Rotary Position Embedding for Consistent Video World Model", arXiv 2026.02. [Paper]
"Debugging code world models", arXiv 2026.02. [Paper]
"Interpreting Physics in Video World Models", arXiv 2026.02. [Paper]
"Neural Sabermetrics with World Model: Play-by-play Predictive Modeling with Large Language Model", arXiv 2026.02. [Paper]
"From Kepler to Newton: Inductive Biases Guide Learned World Models in Transformers", arXiv 2026.02. [Paper]
"Self-Improving World Modelling with Latent Actions", arXiv 2026.02. [Paper]
"Reinforcement World Model Learning for LLM-based Agents", arXiv 2026.02. [Paper]
LIVE: "LIVE: Long-horizon Interactive Video World Modeling", arXiv 2026.02. [Paper] [Website]
EHRWorld: "EHRWorld: A Patient-Centric Medical World Model for Long-Horizon Clinical Trajectories", arXiv 2026.02. [Paper]
"Joint Learning of Hierarchical Neural Options and Abstract World Model", arXiv 2026.02. [Paper]
"Test-Time Mixture of World Models for Embodied Agents in Dynamic Environments", ICLR 2026. [Paper] [Code]
"The Patient is not a Moving Document: A World Model Training Paradigm for Longitudinal EHR", arXiv 2026.01. [Paper]
PathWise: "PathWise: Planning through World Model for Automated Heuristic Design via Self-Evolving LLMs", arXiv 2026.01. [Paper]
"From Observations to Events: Event-Aware World Model for Reinforcement Learning", arXiv 2026.01. [Paper]
NuiWorld: "NuiWorld: Exploring a Scalable Framework for End-to-End Controllable World Generation", arXiv 2026.01. [Paper]
""Just in Time" World Modeling Supports Human Planning and Reasoning", arXiv 2026.01. [Paper]
Action Shapley: "Action Shapley: A Training Data Selection Metric for World Model in Reinforcement Learning", arXiv 2026.01. [Paper]
"Inference-time Physics Alignment of Video Generative Models with Latent World Models", CVPR 2026. [Paper] [Code]
Imagine-then-Plan: "Imagine-then-Plan: Agent Learning from Adaptive Lookahead with World Models", arXiv 2026.01. [Paper]
Puzzle it Out: "Puzzle it Out: Local-to-Global World Model for Offline Multi-Agent Reinforcement Learning", arXiv 2026.01. [Paper]
"Object-Centric World Models Meet Monte Carlo Tree Search", arXiv 2026.01. [Paper]
"Learning Latent Action World Models In The Wild", arXiv 2026.01. [Paper]
VerseCrafter: "VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control", arXiv 2026.01. [Paper] [Website]
"Choreographing a World of Dynamic Objects", arXiv 2026.01. [Paper] [Website]
MobileDreamer: "MobileDreamer: Generative Sketch World Model for GUI Agent", arXiv 2026.01. [Paper]
"Current Agents Fail to Leverage World Model as Tool for Foresight", arXiv 2026.01. [Paper]
"Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments", arXiv 2026.01. [Paper] [Website]
"Value-guided action planning with JEPA world models", ICLR 2026 World Modeling Workshop. [Paper]
NeoVerse: "NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos", arXiv 2026.01. [Paper] [Website]
TeleWorld: "TeleWorld: Towards Dynamic Multimodal Synthesis with a 4D World Model", arXiv 2026.01. [Paper]
"World model inspired sarcasm reasoning with large language model agents", arXiv 2025.12. [Paper]
LEWM: "Large Emotional World Model", arXiv 2025.12. [Paper]
WWM: "Web World Models", arXiv 2025.12. [Paper] [Website]
SurgWorld: "SurgWorld: Learning Surgical Robot Policies from Videos via World Modeling", arXiv 2025.12. [Paper]
Agent2World: "Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback", arXiv 2025.12. [Paper] [Website]
Yume-1.5: "Yume-1.5: A Text-Controlled Interactive World Generation Model", arXiv 2025.12. [Paper] [Website] [Code]
"Aerial World Model for Long-horizon Visual Generation and Navigation in 3D Space", arXiv 2025.12. [Paper]
"From Word to World: Can Large Language Models be Implicit Text-based World Models?", arXiv 2025.12. [Paper]
"Dexterous World Models", arXiv 2025.12. [Paper] [Website]
PhysFire-WM: "PhysFire-WM: A Physics-Informed World Model for Emulating Fire Spread Dynamics", `arxiv 2025.12. [Paper]
WorldPlay: "WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling", `arxiv 2025.12. [Paper] [Website]
"The Double Life of Code World Models: Provably Unmasking Malicious Behavior Through Execution Traces", `arxiv 2025.12. [Paper]
LongVie 2: "LongVie 2: Multimodal Controllable Ultra-Long Video World Model", `arxiv 2025.12. [Paper] [Website]
VFMF: "VFMF: World Modeling by Forecasting Vision Foundation Model Features", `arxiv 2025.12. [Paper] [Website] [Code]
VDAWorld: "VDAWorld: World Modelling via VLM-Directed Abstraction and Simulation", `arxiv 2025.12. [Paper] [Website]
"Closing the Train-Test Gap in World Models for Gradient-Based Planning", `arxiv 2025.12. [Paper]
WonderZoom: "WonderZoom: Multi-Scale 3D World Generation", `arxiv 2025.12. [Paper] [Website]
Astra: "Astra: General Interactive World Model with Autoregressive Denoising", `arxiv 2025.12. [Paper] [Website] [Code]
Visionary: "Visionary: The World Model Carrier Built on WebGPU-Powered Gaussian Splatting Platform", `arxiv 2025.12. [Paper] [Website]
CLARITY: "CLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space", `arxiv 2025.12. [Paper]
UnityVideo: "UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation", `arxiv 2025.12. [Paper] [Website] [Code]
"Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech", `arxiv 2025.12. [Paper]
"Probing the effectiveness of World Models for Spatial Reasoning through Test-time Scaling", `arxiv 2025.12. [Paper] [Code]
ProPhy: "ProPhy: Progressive Physical Alignment for Dynamic World Simulation", `arxiv 2025.12. [Paper]
BiTAgent: "BiTAgent: A Task-Aware Modular Framework for Bidirectional Coupling between Multimodal Large Language Models and World Models", `arxiv 2025.12. [Paper]
RELIC: "RELIC: Interactive Video World Model with Long-Horizon Memory", `arxiv 2025.12. [Paper] [Website]
"Better World Models Can Lead to Better Post-Training Performance", `arxiv 2025.12. [Paper]
SeeU: "SeeU: Seeing the Unseen World via 4D Dynamics-aware Generation", `arxiv 2025.12. [Paper] [Website]
DynamicVerse: "DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling", `arxiv 2025.12. [Paper]
IC-World: "IC-World: In-Context Generation for Shared World Modeling", `arxiv 2025.12. [Paper] [Code]
WorldPack: "WorldPack: Compressed Memory Improves Spatial Consistency in Video World Modeling", `arxiv 2025.12. [Paper]
GrndCtrl: "GrndCtrl: Grounding World Models via Self-Supervised Reward Alignment", `arxiv 2025.12. [Paper]
ChronosObserver: "ChronosObserver: Taming 4D World with Hyperspace Diffusion Sampling", `arxiv 2025.12. [Paper]
AVWM: "Audio-Visual World Models: Towards Multisensory Imagination in Sight and Sound", `arxiv 2025.12. [Paper]
VCWorld: "VCWorld: A Biological World Model for Virtual Cell Simulation", `arxiv 2025.12. [Paper] [Code]
VISTAv2: "VISTAv2: World Imagination for Indoor Vision-and-Language Navigation", `arxiv 2025.12. [Paper] [Website]
Captain Safari: "Captain Safari: A World Engine", arxiv 2025.11. [Paper] [Website]
WorldWander: "WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation", arxiv 2025.11. [Paper] [Code]
Inferix: "Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation", arxiv 2025.11. [Paper] [Code]
MagicWorld, MagicWorld: Towards Long-Horizon Stability for Interactive Video World Exploration. [Paper][Website] [Code]
"Counterfactual World Models via Digital Twin-conditioned Video Diffusion", arxiv 2025.11. [Paper]
WorldGen: "WorldGen: From Text to Traversable and Interactive 3D Worlds", arxiv 2025.11. [Paper] [Website]
X-WIN: "X-WIN: Building Chest Radiograph World Model via Predictive Sensing", arxiv 2025.11. [Paper]
"Object-Centric World Models for Causality-Aware Reinforcement Learning", AAAI 2026. [Paper]
"Latent-Space Autoregressive World Model for Efficient and Robust Image-Goal Navigation", arxiv 2025.11. [Paper]
Dynamic Sparsity: "Dynamic Sparsity: Challenging Common Sparsity Assumptions for Learning World Models in Robotic Reinforcement Learning Benchmarks", AAAI 2026. [Paper]
MrCoM: "MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-Scenarios", AAAI 2026. [Paper]
"Next-Latent Prediction Transformers Learn Compact World Models", arxiv 2025.11. [Paper]
DR. WELL: "DR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent Collaboration", NeurIPS 2025 Workshop: Bridging Language, Agent, and World Models for Reasoning and Planning (LAW). [Paper] [Website]
"How Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert Assessment", arxiv 2025.11. [Paper]
"From Pixels to Cooperation Multi Agent Reinforcement Learning based on Multimodal World Models", arxiv 2025.11. [Paper]
"Bootstrap Off-policy with World Model", NIPS 2025. [Paper]
"Clone Deterministic 3D Worlds with Geometrically-Regularized World Models", arxiv 2025.10. [Paper]
"Semantic Communications with World Models", arxiv 2025.10. [Paper]
TRELLISWorld: "TRELLISWorld: Training-Free World Generation from Object Generators", arxiv 2025.10. [Paper]
WorldGrow: "WorldGrow: Generating Infinite 3D World", arxiv 2025.10. [Paper] [Code]
PhysWorld: "PhysWorld: From Real Videos to World Models of Deformable Objects via Physics-Aware Demonstration Synthesis", arxiv 2025.10. [Paper]
"How Hard is it to Confuse a World Model?", arxiv 2025.10. [Paper]
"Social World Model-Augmented Mechanism Design Policy Learning", NIPS 2025. [Paper]
VAGEN: "VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents", NIPS 2025. [Paper] [Website]
Cosmos-Surg-dVRK: "Cosmos-Surg-dVRK: World Foundation Model-based Automated Online Evaluation of Surgical Robot Policy Learning", arxiv 2025.10. [Paper]
"Zero-shot World Models via Search in Memory", arxiv 2025.10. [Paper]
"Vector Quantization in the Brain: Grid-like Codes in World Models", NIPS 2025. [Paper]
Terra: "Terra: Explorable Native 3D World Model with Point Latents", arxiv 2025.10. [Paper] [Website]
Deep SPI: "Deep SPI: Safe Policy Improvement via World Models", arxiv 2025.10. [Paper]
"One Life to Learn: Inferring Symbolic World Models for Stochastic Environments from Unguided Exploration", arxiv 2025.10. [Paper] [Code]
R-WoM: "R-WoM: Retrieval-augmented World Model For Computer-use Agents", arxiv 2025.10. [Paper]
WorldMirror: "WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting", arxiv 2025.10. [Paper]
Unified World Models: "Unified World Models: Memory-Augmented Planning and Foresight for Visual Navigation", arxiv 2025.10. [Paper] [code]
"Code World Models for General Game Playing", arxiv 2025.10. [Paper]
MorphoSim: "MorphoSim: An Interactive, Controllable, and Editable Language-guided 4D World Simulator", arxiv 2025.10. [Paper] [code]
ChronoEdit: "ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation", arxiv 2025.10. [Paper] [Website]
SFP: "Spatiotemporal Forecasting as Planning: A Model-Based Reinforcement Learning Approach with Generative World Models", arxiv 2025.10. [Paper]
EvoWorld: "EvoWorld: Evolving Panoramic World Generation with Explicit 3D Memory", arxiv 2025.10. [Paper] [Code]
"World Model for AI Autonomous Navigation in Mechanical Thrombectomy", MICCAI 2025. Lecture Notes in Computer Science. [Paper]
DyMoDreamer: "DyMoDreamer: World Modeling with Dynamic Modulation", NeurIPS 2025. [Paper] [Code]
Dreamer4: "Training Agents Inside of Scalable World Models", arxiv 2025.09. [Paper] [Website]
"Reinforcement Learning with Inverse Rewards for World Model Post-training", arxiv 2025.09. [Paper]
"Context and Diversity Matter: The Emergence of In-Context Learning in World Models", arxiv 2025.09. [Paper]
FantasyWorld: "FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D Prediction", arxiv 2025.09. [Paper]
"Remote Sensing-Oriented World Model", arxiv 2025.09. [Paper]
"World Modeling with Probabilistic Structure Integration", arxiv 2025.09. [Paper]
"One Model for All Tasks: Leveraging Efficient World Models in Multi-Task Planning", arxiv 2025.09. [Paper] [Code]
LatticeWorld: "LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation", arxiv 2025.09. [Paper]
"Planning with Reasoning using Vision Language World Model", arxiv 2025.09. [Paper]
MobiWorld: "MobiWorld: World Models for Mobile Wireless Network", arxiv 2025.07. [Paper]
"Continual Reinforcement Learning by Planning with Online World Models", ICML 2025 Spotlight. [Paper]
AirScape: "AirScape: An Aerial Generative World Model with Motion Controllability", arxiv 2025.07. [Paper] [Website]
Geometry Forcing: "Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling", arxiv 2025.07. [Paper] [Website]
Martian World Models: "Martian World Models: Controllable Video Synthesis with Physically Accurate 3D Reconstructions", arxiv 2025.07. [Paper] [Website]
"What Has a Foundation Model Found? Using Inductive Bias to Probe for World Models", ICML 2025. [Paper]
"Critiques of World Models", arxiv 2025.07. [Paper]
"When do World Models Successfully Learn Dynamical Systems?", arxiv 2025.07. [Paper]
"Accurate and Efficient World Modeling with Masked Latent Transformers", arxiv 2025.07. [Paper]
Dyn-O: "Dyn-O: Building Structured World Models with Object-Centric Representations", arxiv 2025.07. [Paper]
NavMorph: "NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments", ICCV 2025. [Paper] [Code]
"A “Good” Regulator May Provide a World Model for Intelligent Systems", arxiv 2025.06. [Paper]
Xray2Xray: "Xray2Xray: World Model from Chest X-rays with Volumetric Context", arxiv 2025.06. [Paper]
MATWM: "Transformer World Model for Sample Efficient Multi-Agent Reinforcement Learning", arxiv 2025.06. [Paper]
"Measuring (a Sufficient) World Model in LLMs: A Variance Decomposition Framework", arxiv 2025.06. [Paper]
"Efficient Generation of Diverse Cooperative Agents with World Models", arxiv 2025.06. [Paper]
WorldLLM: "WorldLLM: Improving LLMs' world modeling using curiosity-driven theory-making", arxiv 2025.06. [Paper]
"LLMs as World Models: Data-Driven and Human-Centered Pre-Event Simulation for Disaster Impact Assessment", arxiv 2025.06. [Paper]
"Bootstrapping World Models from Dynamics Models in Multimodal Foundation Models", arxiv 2025.06. [Paper]
"Video World Models with Long-term Spatial Memory", arxiv 2025.06. [Paper] [Website]
DSG-World: "DSG-World: Learning a 3D Gaussian World Model from Dual State Videos", arxiv 2025.06. [Paper]
"Safe Planning and Policy Optimization via World Model Learning", arxiv 2025.06. [Paper]
FOLIAGE: "FOLIAGE: Towards Physical Intelligence World Models Via Unbounded Surface Evolution", arxiv 2025.06. [Paper]
"Linear Spatial World Models Emerge in Large Language Models", arxiv 2025.06. [Paper] [Code]
Simple, Good, Fast: "Simple, Good, Fast: Self-Supervised World Models Free of Baggage", ICLR 2025. [Paper] [Code]
Medical World Model: "Medical World Model: Generative Simulation of Tumor Evolution for Treatment Planning", arxiv 2025.06. [Paper]
"General agents need world models", ICML 2025. [Paper]
"Learning Abstract World Models with a Group-Structured Latent Space", arxiv 2025.06. [Paper]
DeepVerse: "DeepVerse: 4D Autoregressive Video Generation as a World Model", arxiv 2025.06. [Paper]
"World Models for Cognitive Agents: Transforming Edge Intelligence in Future Networks", arxiv 2025.06. [Paper]
Dyna-Think: "Dyna-Think: Synergizing Reasoning, Acting, and World Model Simulation in AI Agents", arxiv 2025.06. [Paper]
StateSpaceDiffuser: "StateSpaceDiffuser: Bringing Long Context to Diffusion World Models", arxiv 2025.05. [Paper]
"Learning World Models for Interactive Video Generation", arxiv 2025.05. [Paper]
"Revisiting Multi-Agent World Modeling from a Diffusion-Inspired Perspective", arxiv 2025.05. [Paper]
"Long-Context State-Space Video World Models", arxiv 2025.05. [Paper] [Website]
"Unlocking Smarter Device Control: Foresighted Planning with a World Model-Driven Code Execution Approach", arxiv 2025.05. [Paper]
"World Models as Reference Trajectories for Rapid Motor Adaptation", arxiv 2025.05. [Paper]
"Policy-Driven World Model Adaptation for Robust Offline Model-based Reinforcement Learning", arxiv 2025.05. [Paper]
"Building spatial world models from sparse transitional episodic memories", arxiv 2025.05. [Paper]
PoE-World: "PoE-World: Compositional World Modeling with Products of Programmatic Experts", arxiv 2025.05. [Paper] [Website]
"Explainable Reinforcement Learning Agents Using World Models", arxiv 2025.05. [Paper]
seq-JEPA: "seq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World Models", arxiv 2025.05. [Paper]
"Coupled Distributional Random Expert Distillation for World Model Online Imitation Learning", arxiv 2025.05. [Paper]
"Learning Local Causal World Models with State Space Models and Attention", arxiv 2025.05. [Paper]
WebEvolver: "WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model", arxiv 2025.04. [Paper]
WALL-E 2.0: "WALL-E 2.0: World Alignment by NeuroSymbolic Learning improves World Model-based LLM Agents", arxiv 2025.04. [Paper] [Code]
ViMo: "ViMo: A Generative Visual GUI World Model for App Agent", arxiv 2025.04. [Paper]
"Simulating Before Planning: Constructing Intrinsic User World Model for User-Tailored Dialogue Policy Planning", SIGIR 2025. [Paper]
CheXWorld: "CheXWorld: Exploring Image World Modeling for Radiograph Representation Learning", CVPR 2025. [Paper] [Code]
EchoWorld: "EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance", CVPR 2025. [Paper] [Code]
"Adapting a World Model for Trajectory Following in a 3D Game", ICLR 2025 Workshop on World Models. [Paper]
MineWorld: "MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft", arXiv 2025.04. [Paper] [Website]
MoSim: "Neural Motion Simulator Pushing the Limit of World Models in Reinforcement Learning", CVPR 2025. [Paper]
"Improving World Models using Deep Supervision with Linear Probes", ICLR 2025 Workshop on World Models. [Paper]
"Decentralized Collective World Model for Emergent Communication and Coordination", arXiv 2025.04. [Paper]
"Adapting World Models with Latent-State Dynamics Residuals", arXiv 2025.04. [Paper]
"Can Test-Time Scaling Improve World Foundation Model?", arXiv 2025.03. [Paper] [Code]
"Synthesizing world models for bilevel planning", arXiv 2025.03. [Paper]
"Long-context autoregressive video modeling with next-frame prediction", arXiv 2025.03. [Paper] [Code] [Website]
Aether: "Aether: Geometric-Aware Unified World Modeling", arXiv 2025.03. [Paper] [Website]
FUSDREAMER: "FUSDREAMER: Label-efficient Remote Sensing World Model for Multimodal Data Classification", arXiv 2025.03. [Paper] [Website]
"Inter-environmental world modeling for continuous and compositional dynamics", arXiv 2025.03. [Paper]
Disentangled World Models: "Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement Learning", arXiv 2025.03. [Paper]
"Revisiting the Othello World Model Hypothesis", ICLR World Models Workshop. [Paper]
"Learning Transformer-based World Models with Contrastive Predictive Coding", arXiv 2025.03. [Paper]
"Surgical Vision World Model", arXiv 2025.03. [Paper]
"World Models for Anomaly Detection during Model-Based Reinforcement Learning Inference", arXiv 2025.03. [Paper]
WMNav: "WMNav: Integrating Vision-Language Models into World Models for Object Goal Navigation", arXiv 2025.03. [Paper] [Website]
SENSEI: "SENSEI: Semantic Exploration Guided by Foundation Models to Learn Versatile World Models", arXiv 2025.03. [Paper] [Website]
"Learning Actionable World Models for Industrial Process Control", arXiv 2025.03. [Paper]
"Implementing Spiking World Model with Multi-Compartment Neurons for Model-based Reinforcement Learning", arXiv 2025.03. [Paper]
"Discrete Codebook World Models for Continuous Control", ICLR 2025. [Paper]
Multimodal Dreaming: "Multimodal Dreaming: A Global Workspace Approach to World Model-Based Reinforcement Learning", arXiv 2025.02. [Paper]
"Generalist World Model Pre-Training for Efficient Reinforcement Learning", arXiv 2025.02. [Paper]
"Learning To Explore With Predictive World Model Via Self-Supervised Learning", arXiv 2025.02. [Paper]
M^3: "M^3: A Modular World Model over Streams of Tokens", arXiv 2025.02. [Paper]
"When do neural networks learn world models?", arXiv 2025.02. [Paper]
"Pre-Trained Video Generative Models as World Simulators", arXiv 2025.02. [Paper]
DMWM: "DMWM: Dual-Mind World Model with Long-Term Imagination", arXiv 2025.02. [Paper]
EvoAgent: "EvoAgent: Agent Autonomous Evolution with Continual World Model for Long-Horizon Tasks", arXiv 2025.02. [Paper]
"Acquisition through My Eyes and Steps: A Joint Predictive Agent Model in Egocentric Worlds", arXiv 2025.02. [Paper]
"Generating Symbolic World Models via Test-time Scaling of Large Language Models", arXiv 2025.02. [Paper] [Website]
"Improving Transformer World Models for Data-Efficient RL", arXiv 2025.02. [Paper]
"Trajectory World Models for Heterogeneous Environments", arXiv 2025.02. [Paper]
"Enhancing Memory and Imagination Consistency in Diffusion-based World Models via Linear-Time Sequence Modeling", arXiv 2025.02. [Paper]
"Objects matter: object-centric world models improve reinforcement learning in visually complex environments", arXiv 2025.01. [Paper]
GLAM: "GLAM: Global-Local Variation Awareness in Mamba-based World Model", arXiv 2025.01. [Paper]
GAWM: "GAWM: Global-Aware World Model for Multi-Agent Reinforcement Learning", arXiv 2025.01. [Paper]
"Generative Emergent Communication: Large Language Model is a Collective World Model", arXiv 2025.01. [Paper]
"Towards Unraveling and Improving Generalization in World Models", arXiv 2025.01. [Paper]
"Towards Physically Interpretable World Models: Meaningful Weakly Supervised Representations for Visual Trajectory Prediction", arXiv 2024.12. [Paper]
"Transformers Use Causal World Models in Maze-Solving Tasks", arXiv 2024.12. [Paper]
"Causal World Representation in the GPT Model", NIPS 2024 Workshop. [Paper]
Owl-1: "Owl-1: Omni World Model for Consistent Long Video Generation", arXiv 2024.12. [Paper]
"Navigation World Models", arXiv 2024.12. [Paper] [Website]
"Evaluating World Models with LLM for Decision Making", arXiv 2024.11. [Paper]
LLMPhy: "LLMPhy: Complex Physical Reasoning Using Large Language Models and World Models", arXiv 2024.11. [Paper]
WebDreamer: "Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents", arXiv 2024.11. [Paper] [Code]
"Scaling Laws for Pre-training Agents and World Models", arXiv 2024.11. [Paper]
DINO-WM: "DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning", arXiv 2024.11. [Paper] [Website]
"Learning World Models for Unconstrained Goal Navigation", NIPS 2024. [Paper]
"How Far is Video Generation from World Model: A Physical Law Perspective", arXiv 2024.11. [Paper] [Website] [Code]
Adaptive World Models: "Adaptive World Models: Learning Behaviors by Latent Imagination Under Non-Stationarity", NIPS 2024 Workshop Adaptive Foundation Models. [Paper]
LLMCWM: "Language Agents Meet Causality -- Bridging LLMs and Causal World Models", arXiv 2024.10. [Paper] [Code]
"Reward-free World Models for Online Imitation Learning", arXiv 2024.10. [Paper]
"Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation", arXiv 2024.10. [Paper]
AVID: "AVID: Adapting Video Diffusion Models to World Models", arXiv 2024.10. [Paper] [Code]
SMAC: "Grounded Answers for Multi-agent Decision-making Problem through Generative World Model", NeurIPS 2024. [Paper]
OSWM: "One-shot World Models Using a Transformer Trained on a Synthetic Prior", arXiv 2024.09. [Paper]
"Making Large Language Models into World Models with Precondition and Effect Knowledge", arXiv 2024.09. [Paper]
"Efficient Exploration and Discriminative World Model Learning with an Object-Centric Abstraction", arXiv 2024.08. [Paper]
"Not All Actions Are Equal: Rethinking Conditioning for Dexterous World Model", arxiv 2026.06. [Paper]
Tactile-WAM: "Tactile-WAM: Touch-Aware World Action Model with Tactile Asymmetric Attention", arxiv 2026.06. [Paper]
"In-Context World Modeling for Robotic Control", arxiv 2026.06. [Paper]
"World Value Models for Robotic Manipulation", arxiv 2026.06. [Paper]
NavWM: "NavWM: A Unified Navigation World Model for Foresight-Driven Planning", arxiv 2026.06. [Paper]
DynaWM: "DynaWM: Dynamics-Aware Distillation with World Model and Momentum Targets for Smooth Locomotion over Continuous Stairs", arxiv 2026.06. [Paper]
"A Watermark for Vision-Language-Action and World Action Models", arxiv 2026.06. [Paper]
SkyJEPA: "SkyJEPA: Learning Long-Horizon World Models for Zero-Shot Sim-to-Real Control of Quadrotors", arxiv 2026.06. [Paper]
IOI: "IOI: Decoupling Kinematics and Physics for Interactive World Models", arxiv 2026.06. [Paper]
"Causal Reward World Models: Zero-shot Reward Design for Automated Skill Generation", arxiv 2026.06. [Paper]
Foresight: "Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents", arxiv 2026.06. [Paper]
AdaReP: "AdaReP:Adaptive Re-Planning under Model Mismatch for Neural World-Model Predictive Control", arxiv 2026.06. [Paper]
"Attacking the Trusted Imagination: Oracle-Level Integrity Attacks on Imagine-then-Act World Models", arxiv 2026.06. [Paper]
Wh0: "Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data", arxiv 2026.06. [Paper]
MemoryWAM: "MemoryWAM: Efficient World Action Modeling with Persistent Memory", arxiv 2026.06. [Paper]
"Reward as An Agent for Embodied World Models", arxiv 2026.06. [Paper]
Mem-World: "Mem-World: Memory-Augmented Action-Conditioned World Models for Persistent Robot Manipulation", arxiv 2026.06. [Paper]
DREAM-Chunk: "DREAM-Chunk: Reactive Action Chunking with Latent World Model", arxiv 2026.06. [Paper]
PAIWorld: "PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation", arxiv 2026.06. [Paper]
WAM-RL: "WAM-RL: World-Action Model Reinforcement Learning with Reconstruction Rewards and Online Video SFT", arxiv 2026.06. [Paper]
Qwen-RobotWorld: "Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation", arxiv 2026.06. [Paper]
Kairos: "Kairos: A Native World Model Stack for Physical AI", arxiv 2026.06. [Paper]
BRICKS-WM: "BRICKS-WM: Building Reusability via Interface Composition Kinetics for Structured World Models", arxiv 2026.06. [Paper]
"LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies", arxiv 2026.06. [Paper]
RepWAM: "RepWAM: World Action Modeling with Representation Visual-Action Tokenizers", arxiv 2026.06. [Paper]
"$\texttt{WEAVER}$, Better, Faster, Longer: An Effective World Model for Robotic Manipulation", arxiv 2026.06. [Paper]
MaskWAM: "MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models", arxiv 2026.06. [Paper]
NavWAM: "NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation", arxiv 2026.06. [Paper]
EA-WM: "EA-WM: Event-Aware World Models with Task-Specification Grounding for Long-Horizon Manipulation", arxiv 2026.06. [Paper]
EWAM: "EWAM: An Enhanced World Action Model for Closed-Loop Online Adaptation in Embodied Intelligence", arxiv 2026.06. [Paper]
"Making Foresight Actionable: Repurposing Representation Alignment in World Action Models", arxiv 2026.06. [Paper]
PLUME: "PLUME: Probabilistic Latent Unified World Modeling and Parameter Estimation for Multi-Finger Manipulation", arxiv 2026.06. [Paper]
Next Forcing: "Next Forcing: Causal World Modeling with Multi-Chunk Prediction", arxiv 2026.06. [Paper]
TacForeSight: "TacForeSight: Force-Guided Tactile World Model for Contact-Rich Manipulation", arxiv 2026.06. [Paper]
HiMem-WAM: "HiMem-WAM: Hierarchical Memory-Gated World Action Models for Robotic Manipulation", arxiv 2026.06. [Paper]
Efficient-WAM: "Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination", arxiv 2026.06. [Paper]
iMaC: "iMaC: Translating Actions into Motion and Contact Images for Embodied World Models", arxiv 2026.06. [Paper]
WoVR: "WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL", arXiv 2026.02. [Paper]
"Visual Foresight for Robotic Stow: A Diffusion-Based World Model from Sparse Snapshots", arXiv 2026.02. [Paper]
VLAW: "VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model", arXiv 2026.02. [Paper] [Website]
HAIC: "HAIC: Humanoid Agile Object Interaction Control via Dynamics-Aware World Model", arXiv 2026.02. [Paper] [Website]
H-WM: "H-WM: Robotic Task and Motion Planning Guided by Hierarchical World Model", arXiv 2026.02. [Paper]
RISE: "RISE: Self-Improving Robot Policy with Compositional World Model", arXiv 2026.02. [Paper] [Website]
"ContactGaussian-WM: Learning Physics-Grounded World Model from Videos", arXiv 2026.02. [Paper]
"Scaling World Model for Hierarchical Manipulation Policies", arXiv 2026.02. [Paper] [Website]
Say, Dream, and Act: "Say, Dream, and Act: Learning Video World Models for Instruction-Driven Robot Manipulation", arXiv 2026.02. [Paper]
"Affordances Enable Partial World Modeling with LLMs", arXiv 2026.02. [Paper]
VLA-JEPA: "VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model", arXiv 2026.02. [Paper]
MVISTA-4D: "MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation", arXiv 2026.02. [Paper]
Hand2World: "Hand2World: Autoregressive Egocentric Interaction Generation via Free-Space Hand Gestures", arXiv 2026.02. [Paper] [Website]
World-VLA-Loop: "World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy", arXiv 2026.02. [Paper] [Website]
"Coupled Local and Global World Models for Efficient First Order RL", arXiv 2026.02. [Paper]
"Visuo-Tactile World Models", arXiv 2026.02. [Paper]
BridgeV2W: "BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks", arXiv 2026.02. [Paper] [Website]
World-Gymnast: "World-Gymnast: Training Robots with Reinforcement Learning in a World Model", arXiv 2026.02. [Paper] [Website]
MetaWorld: "MetaWorld: Skill Transfer and Composition in a Hierarchical World Model for Grounding High-Level Instructions", arXiv 2026.01. [Paper] [Code]
"Walk through Paintings: Egocentric World Models from Internet Priors", arXiv 2026.01. [Paper]
"Aligning Agentic World Models via Knowledgeable Experience Learning", arXiv 2026.01. [Paper]
ReWorld: "ReWorld: Multi-Dimensional Reward Modeling for Embodied World Models", arXiv 2026.01. [Paper]
"An Efficient and Multi-Modal Navigation System with One-Step World Model", arXiv 2026.01. [Paper]
PointWorld: "PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation", arXiv 2026.01. [Paper] [Website]
Dream2Flow: "Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow", arXiv 2025.12. [Paper] [Website]
"What Drives Success in Physical Planning with Joint-Embedding Predictive World Models?", arXiv 2025.12. [Paper] [Code]
Act2Goal: "Act2Goal: From World Model To General Goal-conditioned Policy", arXiv 2025.12. [Paper] [Website]
AstraNav-World: "AstraNav-World: World Model for Foresight Control and Consistency", arXiv 2025.12. [Paper]
ChronoDreamer: "ChronoDreamer: Action-Conditioned World Model as an Online Simulator for Robotic Planning", arXiv 2025.12. [Paper]
STORM: "STORM: Search-Guided Generative World Models for Robotic Manipulation", arXiv 2025.12. [Paper]
"World Models Can Leverage Human Videos for Dexterous Manipulation", arXiv 2025.12. [Paper]
"Latent Action World Models for Control with Unlabeled Trajectories", arXiv 2025.12. [Paper]
PRISM-WM: "Prismatic World Model: Learning Compositional Dynamics for Planning in Hybrid Systems", arXiv 2025.12. [Paper]
"Learning Robot Manipulation from Audio World Models", arXiv 2025.12. [Paper]
"Embodied Tree of Thoughts: Deliberate Manipulation Planning with Embodied World Model", arXiv 2025.12. [Paper] [Website]
"World Models That Know When They Don't Know: Controllable Video Generation with Calibrated Uncertainty", arXiv 2025.12. [Paper]
"Real-World Robot Control by Deep Active Inference With a Temporally Hierarchical World Model", IEEE Robotics and Automation Letters. [Paper]
"Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling", arXiv 2025.12. [Paper]
IGen: "IGen: Scalable Data Generation for Robot Learning from Open-World Images", arXiv 2025.12. [Paper] [Website]
NavForesee: "NavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation Prediction", arXiv 2025.12. [Paper]
TraceGen: "TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videos", arXiv 2025.11. [Paper] [Website]
ENACT: "ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction", arXiv 2025.11. [Paper] [Website] [Code]
"Learning Massively Multitask World Models for Continuous Control", arXiv 2025.11. [Paper] [Website]
UNeMo: "UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World Model", arXiv 2025.11. [Paper]
"MindForge: Empowering Embodied Agents with Theory of Mind for Lifelong Cultural Learning", NeurIPS 2025. [Paper]
"Towards High-Consistency Embodied World Model with Multi-View Trajectory Videos", arXiv 2025.11. [Paper]
WMPO: "WMPO: World Model-based Policy Optimization for Vision-Language-Action Models", arXiv 2025.11. [Paper] [Website]
"Robot Learning from a Physical World Model", arXiv 2025.11. [Paper] [Website]
"When Object-Centric World Models Meet Policy Learning: From Pixels to Policies, and Where It Breaks", arXiv 2025.11. [Paper]
WorldPlanner: "WorldPlanner: Monte Carlo Tree Search and MPC with Action-Conditioned Visual World Models", arXiv 2025.11. [Paper]
"Learning Interactive World Model for Object-Centric Reinforcement Learning", NIPS 2025. [Paper]
"Scaling Cross-Embodiment World Models for Dexterous Manipulation", arXiv 2025.11. [Paper]
"Co-Evolving Latent Action World Models", arXiv 2025.10. [Paper]
"Ego-Vision World Model for Humanoid Contact Planning", arXiv 2025.10. [Paper] [Website]
Ctrl-World: "Ctrl-World: A Controllable Generative World Model for Robot Manipulation", arXiv 2025.10. [Paper] [Website] [Code]
iMoWM: "iMoWM: Taming Interactive Multi-Modal World Model for Robotic Manipulation", arXiv 2025.10. [Paper] [Website]
WristWorld: "WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation", arXiv 2025.10. [Paper]
"A Recipe for Efficient Sim-to-Real Transfer in Manipulation with Online Imitation-Pretrained World Models", arXiv 2025.10. [Paper]
"Kinodynamic Motion Planning for Mobile Robot Navigation across Inconsistent World Models", RSS 2025 Workshop on Resilient Off-road Autonomous Robotics (ROAR). [Paper]
LongScape: "LongScape: Advancing Long-Horizon Embodied World Models with Context-Aware MoE", arXiv 2025.09. [Paper]
KeyWorld: "KeyWorld: Key Frame Reasoning Enables Effective and Efficient World Models", arXiv 2025.09. [Paper]
DAWM: "DAWM: Diffusion Action World Models for Offline Reinforcement Learning via Action-Inferred Transitions", ICML 2025 Workshop. [Paper]
World4RL: "World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation", arXiv 2025.09. [Paper]
SAMPO: "SAMPO:Scale-wise Autoregression with Motion PrOmpt for generative world models", arXiv 2025.09. [Paper]
PhysicalAgent: "PhysicalAgent: Towards General Cognitive Robotics with Foundation World Models", arXiv 2025.09. [Paper]
SeqWM: "Empowering Multi-Robot Cooperation via Sequential World Models", ICLR 2026. [Paper] [Code]
M3W: "Learning and Planning Multi-Agent Tasks via an MoE-based World Model", NeurIPS 2025. [Paper] [Code]
"World Model Implanting for Test-time Adaptation of Embodied Agents", ICML 2025. [Paper]
"Learning Primitive Embodied World Models: Towards Scalable Robotic Learning", arxiv 2025.08. [Paper] [Website]
GWM: "GWM: Towards Scalable Gaussian World Models for Robotic Manipulation", ICCV 2025. [Paper] [Website]
"Imaginative World Modeling with Scene Graphs for Embodied Agent Navigation", arxiv 2025.08. [Paper]
"Bounding Distributional Shifts in World Modeling through Novelty Detection", arxiv 2025.08. [Paper]
Genie Envisioner: "Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation", arxiv 2025.08. [Paper] [Website]
DiWA: "DiWA: Diffusion Policy Adaptation with World Models", CoRL 2025. [Paper] [Code]
CoEx: "CoEx -- Co-evolving World-model and Exploration", arxiv 2025.07. [Paper]
"Latent Policy Steering with Embodiment-Agnostic Pretrained World Models", arxiv 2025.07. [Paper]
MindJourney: "MindJourney: Test-Time Scaling with World Models for Spatial Reasoning", arxiv 2025.07. [Paper] [Website]
FOUNDER: "FOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision Making", ICML 2025. [Paper] [Website]
EmbodieDreamer: "EmbodieDreamer: Advancing Real2Sim2Real Transfer for Policy Training via Embodied World Modeling", arxiv 2025.07. [Paper] [Website]
World4Omni: "World4Omni: A Zero-Shot Framework from Image Generation World Model to Robotic Manipulation", arxiv 2025.06. [Paper] [Website]
RoboScape: "RoboScape: Physics-informed Embodied World Model", arxiv 2025.06. [Paper] [Code]
ParticleFormer: "ParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic Manipulation", arxiv 2025.06. [Paper] [Website]
ManiGaussian++: "ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model", arxiv 2025.06. [Paper] [Code]
ReOI: "Reimagination with Test-time Observation Interventions: Distractor-Robust World Model Predictions for Visual Model Predictive Control", arxiv 2025.06. [Paper]
GAF: "GAF: Gaussian Action Field as a Dynamic World Model for Robotic Mlanipulation", arxiv 2025.06. [Paper] [Website]
"Prompting with the Future: Open-World Model Predictive Control with Interactive Digital Twins", RSS 2025. [Paper] [Website]
V-JEPA 2 and V-JEPA 2-AC: "V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning", arxiv 2025.06. [Paper] [Website] [Code]
"Time-Aware World Model for Adaptive Prediction and Control", ICML 2025. [Paper]
3DFlowAction: "3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model", arxiv 2025.06. [Paper]
PIN-WM: "PIN-WM: Learning Physics-INformed World Models for Non-Prehensile Manipulation", arXiv 2025.04. [Paper]
"Offline Robotic World Model: Learning Robotic Policies without a Physics Simulator", arXiv 2025.04. [Paper]
ManipDreamer: "ManipDreamer: Boosting Robotic Manipulation World Model with Action Tree and Visual Guidance", arXiv 2025.04. [Paper]
UWM: "Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets", arXiv 2025.04. [Paper] [Website]
"Perspective-Shifted Neuro-Symbolic World Models: A Framework for Socially-Aware Robot Navigation", arXiv 2025.03. [Paper]
AdaWorld: "AdaWorld: Learning Adaptable World Models with Latent Actions", arXiv 2025.03. [Paper] [Website]
DyWA: "DyWA: Dynamics-adaptive World Action Model for Generalizable Non-prehensile Manipulation", arXiv 2025.03. [Paper] [Website]
"Towards Suturing World Models: Learning Predictive Models for Robotic Surgical Tasks", arXiv 2025.03. [Paper] [Website]
"World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning", arXiv 2025.03. [Paper]
LUMOS: "LUMOS: Language-Conditioned Imitation Learning with World Models", ICRA 2025. [Paper] [Website]
"Object-Centric World Model for Language-Guided Manipulation", arXiv 2025.03. [Paper]
DEMO^3: "Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning", arXiv 2025.03. [Paper] [Website]
"Accelerating Model-Based Reinforcement Learning with State-Space World Models", arXiv 2025.02. [Paper]
"Learning Humanoid Locomotion with World Model Reconstruction", arXiv 2025.02. [Paper]
"Strengthening Generative Robot Policies through Predictive World Modeling", arXiv 2025.02. [Paper] [Website]
Robotic World Model: "Robotic World Model: A Neural Network Simulator for Robust Policy Optimization in Robotics", arXiv 2025.01. [Paper]
RoboHorizon: "RoboHorizon: An LLM-Assisted Multi-View World Model for Long-Horizon Robotic Manipulation", arXiv 2025.01. [Paper]
Dream to Manipulate: "Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination", arXiv 2024.12. [Paper] [Website]
R-AIF: "R-AIF: Solving Sparse-Reward Robotic Tasks from Pixels with Active Inference and World Models", arXiv 2024.09. [Paper]
"Representing Positional Information in Generative World Models for Object Manipulation" arXiv 2024.09 [Paper]
DexSim2Real$^2$: "DexSim2Real$^2: Building Explicit World Model for Precise Articulated Object Dexterous Manipulation", arXiv 2024.09. [Paper]
DWL: "Advancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model Learning", RSS 2024 (Best Paper Award Finalist). [Paper]
"Physically Embodied Gaussian Splatting: A Realtime Correctable World Model for Robotics", arXiv 2024.06. [Paper] [Website]
HRSSM: "Learning Latent Dynamic Robust Representations for World Models", ICML 2024. [Paper] [Code]
RoboDreamer: "RoboDreamer: Learning Compositional World Models for Robot Imagination", ICML 2024. [Paper] [Code]
COMBO: "COMBO: Compositional World Models for Embodied Multi-Agent Cooperation", ECCV 2024. [Paper] [Website] [Code]
World Pilot: "World Pilot: Steering Vision-Language-Action Models with World-Action Priors", arxiv 2026.06. [Paper]
"Vision-Language-Action Models Meet World Models: Embodied Agentic AI for Low-Altitude Wireless Networks", arxiv 2026.06. [Paper]
"World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis", arxiv 2026.06. [Paper]
PiL-World: "PiL-World: A Chunk-Wise World Model for VLA Policy-in-the-Loop Evaluation", arxiv 2026.06. [Paper]
"Intercepting the Future: Latent-Space Predictive World Model for Dynamic VLA Manipulation", arxiv 2026.06. [Paper]
Pre-VLA: "Pre-VLA: Preemptive Runtime Verification for Reliable Vision-Language-Action and World-Model Rollouts", arxiv 2026.05. [Paper]
"Reinforcing VLAs in Task-Agnostic World Models", arxiv 2026.05. [Paper]
DIAL: "DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA", arxiv 2026.03. [Paper] [Website]
"Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models", arxiv 2026.03. [Paper]
"Scaling Sim-to-Real Reinforcement Learning for Robot VLAs with Generative 3D Worlds", arxiv 2026.03. [Paper]
Fast-WAM: "Fast-WAM: Do World Action Models Need Test-time Future Imagination?", arxiv 2026.03. [Paper] [Website]
StructVLA: "Beyond Dense Futures: World Models as Structured Planners for Robotic Manipulation", arxiv 2026.03. [Paper]
World2Act: "World2Act: Latent Action Post-Training via Skill-Compositional World Models", arxiv 2026.03. [Paper] [Website]
AtomVLA: "AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World Models", arxiv 2026.03. [Paper]
Chain of World: "Chain of World: World Model Thinking in Latent Motion", CVPR 2026. [Paper] [Website]
"Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation", arxiv 2026.03. [Paper]
World Guidance: "World Guidance: World Modeling in Condition Space for Action Generation", arxiv 2026.02. [Paper] [Website]
SC-VLA: "Self-Correcting VLA: Online Action Refinement via Sparse World Imagination", arxiv 2026.02. [Paper] [Website]
Motus: "Motus: A Unified Latent Action World Model", arxiv 2025.12. [Paper] [Website] [Code]
RoboScape-R: "RoboScape-R: Unified Reward-Observation World Models for Generalizable Robotics Training via RL", arxiv 2025.12. [Paper]
AdaPower: "AdaPower: Specializing World Foundation Models for Predictive Manipulation", arxiv 2025.12. [Paper]
RynnVLA-002: "RynnVLA-002: A Unified Vision-Language-Action and World Model", arxiv 2025.11. [Paper] [Code]
NORA-1.5: "NORA-1.5: A Vision-Language-Action Model Trained using World Model- and Action-based Preference Rewards", arxiv 2025.11. [Paper] [Website] [Code]
"Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model", arxiv 2025.10. [Paper]
VLA-RFT: "VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators", arxiv 2025.10. [Paper]
World-Env: "World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training", arxiv 2025.09. [Paper]
MoWM: "MoWM: Mixture-of-World-Models for Embodied Planning via Latent-to-Pixel Feature Modulation", arxiv 2025.09. [Paper]
LAWM: "Latent Action Pretraining Through World Modeling", arxiv 2025.09. [Paper] [Code]
PAR: "Physical Autoregressive Model for Robotic Manipulation without Action Pretraining", arxiv 2025.08. [Paper] [Website]
DreamVLA: "DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge", arxiv 2025.07. [Paper] [Code] [Website]
WorldVLA: "WorldVLA: Towards Autoregressive Action World Model", arxiv 2025.06. [Paper] [Code]
UP-VLA: "UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent", ICML 2025. [Paper] [Code]
3D-VLA: "3D-VLA: A 3D Vision-Language-Action Generative World Model", ICML 2024. [Paper]
World Models for Visual Understanding
"Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators", arxiv 2026.06. [Paper]
"Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning", arxiv 2026.05. [Paper]
GeoWorld-VLM: "GeoWorld-VLM: Geometry from World Models for Vision-Language Models", arxiv 2026.05. [Paper]
World2VLM: "World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning", arxiv 2026.03. [Paper] [Website] [Code]
DILLO: "Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World Models", arxiv 2026.03. [Paper]
WorldVLM: "WorldVLM: Combining World Model Forecasting and Vision-Language Reasoning", arxiv 2026.03. [Paper]
"When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning", arxiv 2026.01. [Paper] [Website]
"Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models", arxiv 2026.01. [Paper] [Website]
"Semantic World Models", arxiv 2025.10. [Paper] [Website]
DyVA: "Can World Models Benefit VLMs for World Dynamics?", arxiv 2025.10. [Paper] [Website]
"Video models are zero-shot learners and reasoners", arxiv 2025.09. [Paper]
"From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models", arxiv 2025.06. [Paper]
World Models for Autonomous Driving
GeoWorldAD: "GeoWorldAD: Geometry World Action Model for Autonomous Driving", arxiv 2026.07. [Paper]
Orbis 2: "Orbis 2: A Hierarchical World Model for Driving", arxiv 2026.07. [Paper]
M$^\text{4}$World: "M$^\text{4}$World: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming", arxiv 2026.07. [Paper]
"Ego-Dynamics-Augmented World Model for Autonomous Driving with Zero-Shot Cross-Chassis Adaptation", arxiv 2026.07. [Paper]
"Is Energy Guidance All You Need? Training-Free Norm Injection for Driving World Models", arxiv 2026.07. [Paper]
"World Models as Adversaries: Multi-Agent Self-Play Fine-Tuning for Robust Motion Planning", arxiv 2026.07. [Paper]
Valdi: "Valdi: Value Diffusion World Models", arxiv 2026.07. [Paper]
OWMDrive: "OWMDrive: Causality-Aware End-to-End Autonomous Driving via 4D Occupancy World Model", arxiv 2026.06. [Paper]
LWDrive: "LWDrive: Layer-Wise World-Model-Guided Vision-Language Model Planning for Autonomous Driving", arxiv 2026.06. [Paper]
"X-Mind: Efficient Visual Chain-of-Thought via Predictive World Model for End-to-End Driving", arxiv 2026.06. [Paper]
"Risk-Aware Selective Multimodal Driver Monitoring with Driver-State World Modeling", arxiv 2026.06. [Paper]
"Future Dynamic 3D Reconstruction: A 3D World Model with Disentangled Ego-Motion", arxiv 2026.06. [Paper]
OmniDrive: "OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation", arxiv 2026.06. [Paper]
GraphWorld: "GraphWorld: Long-Horizon Planning with World Models for End-to-End Autonomous Driving", arxiv 2026.06. [Paper]
"Metis: A Generalizable and Efficient World-Action Model for Autonomous Driving and Urban Navigation", arxiv 2026.06. [Paper]
"A Tutorial on World Models and Physical AI", arxiv 2026.06. [Paper]
PLAN-S: "PLAN-S: Bridging Planning with Latent Style Dynamics for Autonomous Driving World Models", arxiv 2026.06. [Paper]
Unified Driving Tokens: "Unified Driving Tokens: Representation- and Geometry-Guided Discrete Tokenizer for Driving World Models and Planning", arxiv 2026.06. [Paper]
DriveWAM: "DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving", arxiv 2026.05. [Paper]
X-Foresight: "X-Foresight: A Joint Vision-Action Causal Forecasting Network via Predictive World Modeling", arxiv 2026.05. [Paper]
SparseWorld: "SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation", arxiv 2026.05. [Paper] [Website]
"Reason--Imagine--Act: Closed-Loop LLM Decision Making with World Models for Autonomous Driving", IEEE ITSC 2026. [Paper]
HEAT: "HEAT: Heterogeneous End-to-End Autonomous Driving via Trajectory-Guided World Models", arxiv 2026.05. [Paper]
EponaV2: "EponaV2: Driving World Model with Comprehensive Future Reasoning", arxiv 2026.05. [Paper]
HorizonDrive: "HorizonDrive: Self-Corrective Autoregressive World Model for Long-horizon Driving Simulation", arxiv 2026.05. [Paper]
DeepSight: "DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving", ICML 2026. [Paper] [Code]
CoWorld-VLA: "CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving", arxiv 2026.05. [Paper] [Code]
DriveFuture: "DriveFuture: Future-Aware Latent World Models for Autonomous Driving", arxiv 2026.05. [Paper] [Website]
GEM: "GEM: Generating LiDAR World Model via Deformable Mamba", arxiv 2026.05. [Paper] [Website]
Driver-WM: "Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout", arxiv 2026.05. [Paper]
HERMES++: "HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation", arxiv 2026.04. [Paper] [Website] [Code]
X-Cache: "X-Cache: Cross-Chunk Block Caching for Few-Step Autoregressive World Models Inference", arxiv 2026.04. [Paper]
"Active World-Model with 4D-informed Retrieval for Exploration and Awareness", arxiv 2026.04. [Paper]
"Learning Vision-Language-Action World Models for Autonomous Driving", CVPR 2026 Findings. [Paper]
LMGenDrive: "LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving", arxiv 2026.04. [Paper]
"Beyond Static Forecasting: Unleashing the Power of World Models for Mobile Traffic Extrapolation", arxiv 2026.04. [Paper]
DeltaWorld: "A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens", CVPR 2026. [Paper] [Code]
DriveDreamer-Policy: "DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning", arxiv 2026.04. [Paper] [Website]
DLWM: "DLWM: Dual Latent World Models enable Holistic Gaussian-centric Pre-training in Autonomous Driving", CVPR 2026. [Paper]
AutoWorld: "AutoWorld: Scaling Multi-Agent Traffic Simulation with Self-Supervised World Models", arxiv 2026.03. [Paper]
OccSim: "OccSim: Multi-kilometer Simulation with Long-horizon Occupancy World Models", arxiv 2026.03. [Paper]
Uni-World VLA: "Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving", arxiv 2026.03. [Paper]
DreamerAD: "DreamerAD: Efficient Reinforcement Learning via Latent World Model for Autonomous Driving", arxiv 2026.03. [Paper]
Latent-WAM: "Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving", arxiv 2026.03. [Paper]
"Toward Physically Consistent Driving Video World Models under Challenging Trajectories", arxiv 2026.03. [Paper] [Website]
CounterScene: "CounterScene: Counterfactual Causal Reasoning in Generative World Models for Safety-Critical Closed-Loop Evaluation", arxiv 2026.03. [Paper]
X-World: "X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving", arxiv 2026.03. [Paper]
DynFlowDrive: "DynFlowDrive: Flow-Based Dynamic World Modeling for Autonomous Driving", arxiv 2026.03. [Paper] [Code]
Enactor: "Enactor: From Traffic Simulators to Surrogate World Models", arxiv 2026.03. [Paper]
VectorWorld: "VectorWorld: Efficient Streaming World Model via Diffusion Flow on Vector Graphs", arxiv 2026.03. [Paper] [Code]
"Bridging Scene Generation and Planning: Driving with World Model via Unifying Vision and Motion Representation", arxiv 2026.03. [Paper] [Code]
"Latent World Models for Automated Driving: A Unified Taxonomy, Evaluation Framework, and Open Challenges", arxiv 2026.03. [Paper]
"Kinematics-Aware Latent World Models for Data-Efficient Autonomous Driving", arxiv 2026.03. [Paper]
ShareVerse: "ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling", arxiv 2026.03. [Paper]
"Risk-Aware World Model Predictive Control for Generalizable End-to-End Autonomous Driving", arxiv 2026.02. [Paper]
RAYNOVA: "RAYNOVA: Scale-Temporal Autoregressive World Modeling in Ray Space", CVPR 2026. [Paper] [Website]
"When World Models Dream Wrong: Physical-Conditioned Adversarial Attacks against World Models", arxiv 2026.02. [Paper]
"Factored Latent Action World Models", arxiv 2026.02. [Paper]
ResWorld: "ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving", arxiv 2026.02. [Paper] [Code]
DriveWorld-VLA: "DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving", arxiv 2026.02. [Paper] [Code]
"Safe Urban Traffic Control via Uncertainty-Aware Conformal Prediction and World-Model Reinforcement Learning", arxiv 2026.02. [Paper]
InstaDrive: "InstaDrive: Instance-Aware Driving World Models for Realistic and Consistent Video Generation", arxiv 2026.02. [Paper] [Website]
ConsisDrive: "ConsisDrive: Identity-Preserving Driving World Models for Video Generation by Instance Mask", arxiv 2026.02. [Paper] [Website]
MAD: "MAD: Motion Appearance Decoupling for efficient Driving World Models", arxiv 2026.01. [Paper] [Website]
UniDrive-WM: "UniDrive-WM: Unified Understanding, Planning and Generation World Model For Autonomous Driving", arxiv 2026.01. [Paper] [Website]
DriveLaW: "DriveLaW:Unifying Planning and Video Generation in a Latent Driving World", arxiv 2025.12. [Paper]
GaussianDWM: "GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation", arxiv 2025.12. [Paper] [Code]
WorldRFT: "WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous Driving", AAAI 2026. [Paper]
InDRiVE: "InDRiVE: Reward-Free World-Model Pretraining for Autonomous Driving via Latent Disagreement", arxiv 2025.12. [Paper]
GenieDrive: "GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation", arxiv 2025.12. [Paper] [Website]
FutureX: "FutureX: Enhance End-to-End Autonomous Driving via Latent Chain-of-Thought World Model", arxiv 2025.12. [Paper]
"Latent Chain-of-Thought World Modeling for End-to-End Driving", arxiv 2025.12. [Paper]
MindDrive: "MindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous Driving", arxiv 2025.12. [Paper]
"Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles", arxiv 2025.12. [Paper]
U4D: "U4D: Uncertainty-Aware 4D World Modeling from LiDAR Sequences", arxiv 2025.12. [Paper]
"Vehicle Dynamics Embedded World Models for Autonomous Driving", arXiv 2025.12. [Paper]
"World Model Robustness via Surprise Recognition", arXiv 2025.12. [Paper]
SparseWorld-TC: "SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model", arXiv 2025.11. [Paper]
AD-R1: "AD-R1: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving with Impartial World Models", arXiv 2025.11. [Paper]
Map-World: "Map-World: Masked Action planning and Path-Integral World Model for Autonomous Driving", arXiv 2025.11. [Paper]
WPT: "WPT: World-to-Policy Transfer via Online World Model Distillation", arXiv 2025.11. [Paper]
Percept-WAM: "Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving", arXiv 2025.11. [Paper]
Thinking Ahead: "Thinking Ahead: Foresight Intelligence in MLLMs and World Models", arXiv 2025.11. [Paper]
LiSTAR: "LiSTAR: Ray-Centric World Models for 4D LiDAR Sequences in Autonomous Driving", arXiv 2025.11. [Paper] [Website]
"Dual-Mind World Models: A General Framework for Learning in Dynamic Wireless Networks", arXiv 2025.10. [Paper]
"Addressing Corner Cases in Autonomous Driving: A World Model-based Approach with Mixture of Experts and LLMs", arXiv 2025.10. [Paper]
"From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction", NIPS 2025. [Paper] [Code]
"Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks", arXiv 2025.10. [Paper] [Website]
SparseWorld: "SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic Queries", arXiv 2025.10. [Paper] [Code]
"Vision-Centric 4D Occupancy Forecasting and Planning via Implicit Residual World Models", arXiv 2025.10. [Paper]
DriveVLA-W0: "DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving", arXiv 2025.10. [Paper] [Code]
CoIRL-AD: "CoIRL-AD: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous Driving", arXiv 2025.10. [Paper] [Code]
TeraSim-World: "TeraSim-World: Worldwide Safety-Critical Data Synthesis for End-to-End Autonomous Driving", arXiv 2025.09. [Paper] [Website]
"Enhancing Physical Consistency in Lightweight World Models", arXiv 2025.09. [Paper]
OccTENS: "OccTENS: 3D Occupancy World Model via Temporal Next-Scale Prediction", arXiv 2025.09. [Paper]
IRL-VLA: "IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model", arXiv 2025.08. [Paper] [Website] [Code]
LiDARCrafter: "LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences", arXiv 2025.08. [Paper] [Website] [Code]
FASTopoWM: "FASTopoWM: Fast-Slow Lane Segment Topology Reasoning with Latent World Models", arXiv 2025.07. [Paper] [Code]
Orbis: "Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models", arXiv 2025.07. [Paper] [Code]
"World Model-Based End-to-End Scene Generation for Accident Anticipation in Autonomous Driving", arXiv 2025.07. [Paper]
NRSeg: "NRSeg: Noise-Resilient Learning for BEV Semantic Segmentation via Driving World Models", arXiv 2025.07. [Paper] [Code]
World4Drive: "World4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model", ICCV2025. [Paper] [Code]
Epona: "Epona: Autoregressive Diffusion World Model for Autonomous Driving", ICCV2025. [Paper] [Code]
"Towards foundational LiDAR world models with efficient latent flow matching", arXiv 2025.06. [Paper]
SceneDiffuser++: "SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model", CVPR 2025. [Paper]
COME: "COME: Adding Scene-Centric Forecasting Control to Occupancy World Model", arXiv 2025.06. [Paper] [Code]
STAGE: "STAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation", arXiv 2025.06. [Paper]
ReSim: "ReSim: Reliable World Simulation for Autonomous Driving", arXiv 2025.06. [Paper] [Code] [Project Page]
"Ego-centric Learning of Communicative World Models for Autonomous Driving", arXiv 2025.06. [Paper]
Dreamland: "Dreamland: Controllable World Creation with Simulator and Generative Models", arXiv 2025.06. [Paper] [Project Page]
LongDWM: "LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model", arXiv 2025.06. [Paper] [Project Page]
GeoDrive: "GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control", arXiv 2025.05. [Paper] [Code]
FutureSightDrive: "FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving", NeurIPS 2025. [Paper] [Code]
Raw2Drive: "Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2)", arXiv 2025.05. [Paper]
VL-SAFE: "VL-SAFE: Vision-Language Guided Safety-Aware Reinforcement Learning with World Models for Autonomous Driving", arXiv 2025.05. [Paper] [Project Page]
PosePilot: "PosePilot: Steering Camera Pose for Generative World Models with Self-supervised Depth", arXiv 2025.05. [Paper]
"World Model-Based Learning for Long-Term Age of Information Minimization in Vehicular Networks", arXiv 2025.05. [Paper]
"Learning to Drive from a World Model", arXiv 2025.04. [Paper]
DriVerse: "DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment", arXiv 2025.04. [Paper]
"End-to-End Driving with Online Trajectory Evaluation via BEV World Model", arXiv 2025.04. [Paper] [Code]
"Knowledge Graphs as World Models for Semantic Material-Aware Obstacle Handling in Autonomous Vehicles", arXiv 2025.03. [Paper]
MiLA: "MiLA: Multi-view Intensive-fidelity Long-term Video Generation World Model for Autonomous Driving", arXiv 2025.03. [Paper] [Project Page]
SimWorld: "SimWorld: A Unified Benchmark for Simulator-Conditioned Scene Generation via World Model", arXiv 2025.03. [Paper] [Project Page]
UniFuture: "Seeing the Future, Perceiving the Future: A Unified Driving World Model for Future Generation and Perception", arXiv 2025.03. [Paper] [Project Page]
EOT-WM: "Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latent Space", arXiv 2025.03. [Paper]
"Temporal Triplane Transformers as Occupancy World Models", arXiv 2025.03. [Paper]
InDRiVE: "InDRiVE: Intrinsic Disagreement based Reinforcement for Vehicle Exploration through Curiosity Driven Generalized World Model", arXiv 2025.02. [Paper]
MaskGWM: "MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction", arXiv 2025.02. [Paper]
Dream to Drive: "Dream to Drive: Model-Based Vehicle Control Using Analytic World Models", arXiv 2025.02. [Paper]
"Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous Driving", ICLR 2025. [Paper]
"Dream to Drive with Predictive Individual World Model", IEEE TIV. [Paper] [Code]
HERMES: "HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation", arXiv 2025.01. [Paper]
AdaWM: "AdaWM: Adaptive World Model based Planning for Autonomous Driving", ICLR 2025. [Paper]
AD-L-JEPA: "AD-L-JEPA: Self-Supervised Spatial World Models with Joint Embedding Predictive Architecture for Autonomous Driving with LiDAR Data", arXiv 2025.01. [Paper]
DrivingWorld: "DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT", arXiv 2024.12. [Paper] [Code] [Project Page]
DrivingGPT: "DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers", arXiv 2024.12. [Paper] [Project Page]
"An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training", arXiv 2024.12. [Paper]
GEM: "GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control", arXiv 2024.12. [Paper] [Project Page]
GaussianWorld: "GaussianWorld: Gaussian World Model for Streaming 3D Occupancy Prediction", arXiv 2024.12. [Paper] [Code]
Doe-1: "Doe-1: Closed-Loop Autonomous Driving with Large World Model", arXiv 2024.12. [Paper] [Project Page] [Code]
"Pysical Informed Driving World Model", arXiv 2024.12. [Paper] [Project Page]
InfiniCube: "InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models", arXiv 2024.12. [Paper] [Project Page]
InfinityDrive: "InfinityDrive: Breaking Time Limits in Driving World Models", arXiv 2024.12. [Paper] [Project Page]
ReconDreamer: "ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration", arXiv 2024.11. [Paper] [Project Page]
Imagine-2-Drive: "Imagine-2-Drive: High-Fidelity World Modeling in CARLA for Autonomous Vehicles", ICRA 2025. [Paper] [Project Page]
"Mitigating Covariate Shift in Imitation Learning for Autonomous Vehicles Using Latent Space Generative World Models", arXiv 2024.09. [Paper]
LatentDriver: "Learning Multiple Probabilistic Decisions from Latent World Model in Autonomous Driving", arXiv 2024.09. [Paper] [Code]
RenderWorld: "World Model with Self-Supervised 3D Label", arXiv 2024.09. [Paper]
OccLLaMA: "An Occupancy-Language-Action Generative World Model for Autonomous Driving", arXiv 2024.09. [Paper]
DriveGenVLM: "Real-world Video Generation for Vision Language Model based Autonomous Driving", arXiv 2024.08. [Paper]
Drive-OccWorld: "Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous Driving", arXiv 2024.08. [Paper]
CarFormer: "Self-Driving with Learned Object-Centric Representations", ECCV 2024. [Paper] [Code]
BEVWorld: "A Multimodal World Model for Autonomous Driving via Unified BEV Latent Space", arXiv 2024.07. [Paper] [Code]
TOKEN: "Tokenize the World into Object-level Knowledge to Address Long-tail Events in Autonomous Driving", arXiv 2024.07. [Paper]