70 systems · 36 organisations · last checked 10 Oct 2026

World models

A world model is an AI model that predicts how a scene will change, often in response to an action. This page explains how world models developed, the main types, who builds them and where they are used.

01

What a world model is

A world model takes what a camera sees now, and often an action, and predicts what the camera will see next.

  • A robot can practise a task inside a world model instead of in the real world.
  • A world model can test a robot’s policy before the policy runs on a real robot.
  • A world model can generate videos and scenes that are used as training data.
  • A world model is learned from data, such as videos and robot recordings. A physics simulator is written by engineers from the equations of physics.
NOW Action Worldmodel PREDICTED NEXT FED BACK IN TO LOOK FURTHER AHEAD
02

How world models developed

The chart shows 70 systems released between 2015 and 2026. Each row is one type of world model. A line joins a system to an earlier system that it builds on.

Used for
20152016’17201820192020’2120222023202420252026Latent dynamics11 systemsPredictive representation4 systemsAction-conditioned video21 systemsWorld-action6 systemsVideo generation11 systemsInteractive worlds10 systems3D and 4D worlds3 systemsLearned physics4 systemsAction-conditional video prediction (Atari)Physical interaction video predictionDeep Visual ForesightInteraction NetworksWorld ModelsPlaNetMuZeroDreamerGraph Network Simulator (GNS)DreamerV2MeshGraphNetsJEPA proposalTD-MPCVideo Diffusion ModelsDayDreamerIRISDreamerV3UniPiGAIA-1DriveDreamerTD-MPC2UniSimStable Video DiffusionGenieSoraV-JEPADIAMONDVistaGameNGen1X World ModelOasisDINO-WMGenie 2Veo 2Cosmos world foundation modelsMuse (WHAM)WanGAIA-2DreamGenOdyssey-1Veo 3V-JEPA 2 and V-JEPA 2-ACHunyuanWorld 1.0Genie 3Genie EnvisionerMatrix-Game 2.0NeRD (Neural Robot Dynamics)Dreamer 4MarbleSora 2UnifoLM-WMA-0Cosmos-Predict2.5Ctrl-WorldGAIA-3GWM-1HY-World 1.5 (WorldPlay)Veo world simulator for Gemini Robotics1XWM as NEO policyCosmos PolicyLingBot-WorldDreamDojoDreamZeroWaymo World ModelMatrix-Game 3.0Cosmos 3GE-Sim 2.0Oasis 3GAIA-4AtlasOdyssey-3
Weights are publicWeights are not publicNot knownBuilds on an earlier systemLink inferred by usYears with more releases are drawn wider. Select a system to see its details and sources.
  1. 2015–2017

    Predicting what happens next

    Researchers trained networks to predict game frames or robot camera images from actions, and to predict object physics with graph networks.

  2. 2018–2022

    Learning inside a world model

    Agents such as World Models, PlaNet, Dreamer, MuZero and TD-MPC learned compact models of their environment and planned or trained inside them.

  3. 2023–2024

    Video models as simulators

    Large video generators such as UniSim, GAIA-1, Sora and Genie were trained on internet or driving video and presented as simulators of the world.

  4. 2025–2025

    Foundation and real-time world models

    Companies released open world foundation models (Cosmos, Wan) and real-time interactive worlds (Genie 3, Matrix-Game 2.0), and robot teams used them for data and evaluation.

  5. 2026–2026

    World models as robot policies and test tracks

    World models were used directly as robot policies (DreamZero, Cosmos Policy, 1XWM) and as closed-loop simulators for robots and cars (GAIA-4, Oasis 3, Waymo World Model).

03

Types of world models

There are 8 main types. They differ in what they take in, what they predict and what they are used for.

Latent dynamics models

A latent dynamics model compresses each observation, such as a camera image, into a short list of numbers and learns to predict how that list changes after each action. It takes past observations and actions as input and outputs the predicted next state, usually together with the reward the agent will receive. It is used to plan actions or to train a decision-making program (a policy) on predicted experience, so the agent needs fewer real trials.567+6 Inferred

How it works

  1. The agent acts in a game, a simulator or the real world and records what it observed, which action it took and what reward it received.
  2. An encoder network compresses each observation into a short list of numbers called the latent state.
  3. A dynamics network learns to predict the next latent state and the reward from the current latent state and the chosen action.
  4. The agent plans, or trains its policy, by running many predicted futures inside the model instead of in the real environment.
  5. The agent then acts with the improved policy, records new data and updates the model, and the cycle repeats.

Takes actions

Yes

Used for

planningtraining policiesrobot controlgames

Strengths

  • Predicting a short list of numbers is cheaper than predicting full images, so a planner can test many candidate actions quickly. The 2026 robotic manipulation survey (arXiv 2606.00113) gives this as the main appeal of latent models.
  • Learning needs fewer real trials because most practice happens inside the model. DayDreamer reports a four-legged robot learning to roll off its back, stand up and walk from scratch in 1 hour, without resets.
  • One method can cover many tasks with one setting. DreamerV3 reports results on over 150 diverse tasks with a single configuration, and TD-MPC2 reports a single 317M parameter agent performing 80 tasks.
  • The approach works without being told the rules. MuZero learned Go, chess, shogi and 57 Atari games without knowledge of their underlying dynamics.

Limits

  • Errors add up when the model predicts many steps ahead, because each step starts from the previous prediction. The 2025 embodied AI survey (arXiv 2510.16732) names this error accumulation as the main weakness of step-by-step prediction.
  • Compressing an image into a short list of numbers can drop small visual details that matter for the task. The DIAMOND authors make this point about compact discrete latent states.
  • People cannot easily read the internal numbers, so it is hard to check whether a prediction is physically plausible. The 2026 robotic manipulation survey calls this a loss of auditability.
  • A model trained for one task or one robot often transfers poorly to others. The 2025 embodied AI survey and the 2026 robotic manipulation survey both report this for task-specific world models.

Examples

  • World ModelsGoogle Brain; NNAISENSE; Swiss AI Lab IDSIA (USI & SUPSI) · 2018
  • PlaNetGoogle Brain; DeepMind; Google Research; University of Toronto; University of Michigan · 2018
  • MuZeroDeepMind · 2019
  • DayDreamerUniversity of California, Berkeley · 2022
  • DreamerV3Google DeepMind; University of Toronto · 2023
  • TD-MPC2University of California San Diego · 2023

Predictive representation models

A predictive representation model hides part of an image or video and learns to predict a numeric description (features) of the hidden or future part, without drawing any pixels. It takes images or video, sometimes with robot actions, and outputs predicted features. It is used as a general model of what video shows, and with actions added it is used to plan robot movements toward a goal image.1094159+1 Inferred

How it works

  1. An encoder network turns images or video frames into features.
  2. Parts of the input are hidden, or lie in the future, and a predictor network estimates the features of those parts from the visible parts.
  3. Training compares the estimate with the encoder's features of the real hidden parts and adjusts both networks. Extra rules stop the networks from giving every input the same features, a failure called collapse.
  4. For robot control, an action-conditioned predictor is trained on robot video. V-JEPA 2-AC used less than 62 hours of unlabeled robot videos from the Droid dataset.
  5. To plan, the system tries candidate action sequences, predicts the resulting features and picks the sequence whose prediction is closest to the features of a goal image.

Takes actions

Optional

Used for

pretraining visual modelsplanningrobot control

Strengths

  • No computing is spent on drawing texture and lighting. Meta's V-JEPA 2 and NYU's DINO-WM both predict features without reconstructing pixels.
  • Most training uses video without action labels. V-JEPA 2 was pretrained on over 1 million hours of internet video and then needed less than 62 hours of robot video for its action-conditioned part.
  • Robots can be controlled in new places without new data. Meta deployed V-JEPA 2-AC zero-shot on Franka arms in two different labs for picking and placing objects, without collecting data from the robots in these environments.
  • Small models can plan quickly. The LeWorldModel authors report about 15M parameters, training on a single GPU in a few hours, and planning up to 48x faster than foundation-model-based world models.

Limits

  • The predictions are lists of numbers, so people cannot watch them to check whether they are right.
  • Long plans are hard. The V-JEPA 2 authors report that prediction accuracy falls in longer rollouts and that the number of possible action sequences grows exponentially with the planning horizon.
  • Results depend on the physical setup. V-JEPA 2-AC was sensitive to camera position, and its authors tried camera positions by hand before finding one that worked.
  • Goals are usually given as images. The V-JEPA 2 authors note that language would be a more natural way to state goals for robots in everyday settings.
  • Training can fail through collapse. The LeWorldModel authors describe earlier methods as fragile and dependent on extra losses, moving-average tricks, pretrained encoders or extra supervision to avoid it.

Examples

  • I-JEPAMeta AI (FAIR), with McGill University, Mila and New York University · 2023
  • DINO-WMNew York University (Courant Institute); Meta AI · 2024
  • V-JEPA 2 and V-JEPA 2-ACFAIR at Meta, with Mila and Polytechnique Montréal · 2025
  • LeWorldModelMila & Université de Montréal; New York University; Samsung SAIL; Brown University · 2026

Action-conditioned video models

An action-conditioned video model predicts the next video frames of a scene given the actions that a robot, a car or a game agent will take. It takes recent camera frames and a sequence of planned actions as input and outputs the video those actions would produce. Teams use it to test robot and driving software without the real machine, to create training data and to train policies inside the predictions.212537+4 Inferred

How it works

  1. Collect videos from robots, cars or games together with the action taken at each moment, such as arm positions, steering or game controls.
  2. Start from a video model, often one pretrained on general video, and train it to predict the next frames from past frames plus the recorded actions.
  3. To test a policy, feed the policy's actions into the model, show the policy the predicted frames, and repeat step by step.
  4. Judge whether the predicted video shows the task succeeding, and compare policies by their predicted success.
  5. Successful predicted runs can also be used as extra training examples for the policy.

Takes actions

Yes

Used for

evaluating policiestraining policiesgenerating training datatesting self-driving software

Strengths

  • Policies can be compared without running real robots. The Ctrl-World authors report that their model can accurately rank policy performance without real-world robot rollouts, and Google DeepMind checked its Veo-based evaluator against 1600+ real-world evaluations of eight Gemini Robotics policy checkpoints and five tasks.
  • Rare or unsafe situations can be produced on request. Wayve built GAIA-2 to simulate both common and rare driving scenarios, and the Veo-based evaluator edits scenes to add new objects, backgrounds and distractors.
  • Predicted runs can improve a policy. Ctrl-World reports that fine-tuning on successful trajectories generated in the model improved policy success by 44.7%.
  • Policies trained inside the model can work on real robots. The UniSim authors report vision-language and reinforcement learning policies that were deployed in the real world zero-shot after training purely in their simulator.

Limits

  • The model can show events that cannot happen. The UniSim authors report hallucinations when an action does not fit the scene, and 1X reports many generations that fail to adhere to physical laws.
  • Objects can change shape or colour or disappear during interaction, as 1X reports for its world model.
  • Memory is short. The UniSim authors give the example of an apple that disappears from a drawer when the frames showing it being put there are no longer in the model's input.
  • Accuracy drops for robots, scenes and tasks that are missing from the training data, which the UniSim authors list as limited out-of-domain generalization.
  • A video that looks right can still show a physically wrong result for the robot. The 2026 robotic manipulation survey states that visual plausibility is not equivalent to action validity.

Examples

World-action models

A world-action model predicts what the scene will look like next and also outputs the actions a robot should take, in one model. It takes camera frames, a task instruction and the robot's state, and outputs robot actions and, in the examples here, predicted future frames as well. It is used directly as the robot's controller.19111112+4 Inferred

How it works

  1. Start from a model pretrained to generate video. GR-1 uses large-scale video generative pretraining, and DreamZero builds on a pretrained video diffusion backbone.
  2. Train it on robot data so it predicts future frames and robot actions together.
  3. At run time the robot feeds in its camera view and instruction, and the model outputs the next actions.
  4. Some designs skip drawing the video at run time to save time. UVA decodes video and actions separately so that action output stays fast.
  5. An earlier design generates a video plan first and then extracts actions from it with a separate model. UniPi works this way.

Takes actions

Yes

Used for

robot control

Strengths

  • NVIDIA reports that DreamZero generalised to new tasks and environments over 2x better than the vision-language-action models it was compared with, in real robot experiments.
  • Video of other robots or of humans can help. DreamZero reports a relative improvement of over 42% on unseen tasks from 10-20 minutes of video-only demonstrations.
  • ByteDance Research reports that GR-1 raised the success rate on the CALVIN benchmark from 88.9% to 94.9%.

Limits

  • Large models are slow. DreamZero needed model and system optimizations to run a 14B model for real-time closed-loop control at 7Hz.
  • Very fine precision is hard. The DreamZero authors note limitations on tasks requiring sub-centimeter precision, such as key insertion.
  • The UVA authors state that video-generation-based methods have struggled to match direct policy learning in action accuracy and inference speed.
  • When prediction is hidden inside the controller, it is hard to tell whether the model learned useful dynamics or only better image features, as the 2026 robotic manipulation survey notes.

Examples

Video generators

A video generator creates a new video clip from a text description, a starting image or a short video, and it does not receive step-by-step actions from a player or a robot. Its output is the video clip. It is used to make video content and as a base model that builders later fine-tune into action-conditioned world models, and OpenAI describes scaling such models as a path towards general-purpose simulators of the physical world.1154554+1 Inferred

How it works

  1. Builders collect very large sets of videos and images and attach a text description to each. NVIDIA reports about 20M hours of raw video for Cosmos, and OpenAI trained a captioning model to describe every video in Sora's training set.
  2. A compression network shrinks each video into a smaller grid of numbers so the generator can process long, high-resolution clips.
  3. The generator learns to turn random noise into a clip that matches the description, or to continue a given clip.
  4. At use time a person supplies a prompt, a start image or a clip, and the model produces the new clip.
  5. Builders can fine-tune the same model on robot or driving video with camera paths or actions as extra inputs, which turns it into an action-conditioned model. The Cosmos paper shows such fine-tuning for camera control, robotic manipulation and autonomous driving.

Takes actions

No

Used for

content creationbase model for other world modelsgenerating training data

Strengths

  • Training needs only videos, images and text descriptions, which exist in very large amounts. NVIDIA built Cosmos from about 20M hours of raw video.
  • Large models show some 3D consistency and object permanence without being designed for it. OpenAI reports that Sora keeps people and scene elements consistent as the camera moves, and often, though not always, keeps objects that are hidden or leave the frame.
  • One model can perform visual tasks it was not trained for. Google DeepMind reports Veo 3 segmenting objects, detecting edges, editing images and solving mazes without task-specific training.
  • A general model can be specialised for a robot or a car with less data. The Cosmos paper states that the dataset for post-training can be much smaller than the pretraining data.
  • Generated videos can become robot training data. NVIDIA's DreamGen adapts image-to-video models to a robot, recovers actions from the generated videos, and reports a humanoid robot performing 22 new behaviors while its teleoperation data covered a single pick-and-place task in one environment.

Limits

  • Generated clips can break physics. OpenAI states that Sora does not accurately model the physics of many basic interactions, such as glass shattering.
  • Long clips can drift. OpenAI lists incoherence in long samples and objects that appear spontaneously, and NVIDIA reports objects unexpectedly appearing from below in some outputs of its autoregressive Cosmos models.
  • The model has no input for a specific robot action, so it cannot directly show what happens if a robot moves its arm a certain way.
  • Uses beyond video content are often proposals. The Cosmos paper lists policy evaluation, policy training, planning and synthetic data generation as uses, and states that it does not include empirical results for them.
  • Builders often do not publish the model details. OpenAI's Sora report states that model and implementation details are not included.

Examples

Real-time interactive worlds

A real-time interactive world model draws a game-like world one frame at a time while a person or an AI agent controls it with a keyboard, a mouse or text commands. It takes the latest control input and the frames it has already drawn, and it outputs the next frame fast enough for live play. It is used in games and interactive media research and as a place to train and test AI agents.283236+3 Inferred

How it works

  1. Collect many hours of gameplay or video. GameNGen recorded an AI agent playing DOOM, Matrix-Game 2.0 produced about 1200 hours of video from Unreal Engine and GTA5 environments, and Genie learned from unlabelled Internet videos.
  2. Train a model to predict the next frame from recent frames and the control pressed at that moment. Genie learns its own small set of controls from videos that have no control labels.
  3. Make each frame fast enough for live play, for example by reducing the number of generation steps per frame.
  4. During play, each new frame joins the model's recent history, the player reacts, and the loop repeats.

Takes actions

Yes

Used for

gamescontent creationtraining policiesevaluating policies

Strengths

  • Worlds can be started from a prompt without hand-built 3D assets. Genie can be prompted with text, synthetic images, photographs and sketches, and Genie 3 generates worlds from a text prompt.
  • Live play is possible on current hardware. GameNGen runs at 20 frames per second on a single TPU, Oasis at 20 frames per second, Matrix-Game 2.0 at 25 FPS and Genie 3 at 24 frames per second at 720p.
  • Short clips can be hard to tell apart from the real game. The GameNGen authors report that human raters are only slightly better than random chance at distinguishing short clips of the game from clips of the simulation.
  • The worlds can host AI agents. Google DeepMind ran its SIMA agent inside Genie 3 worlds and gave it goals to pursue.

Limits

  • Memory is short. GameNGen has access to a little over 3 seconds of history, the first Genie model is limited to 16 frames of memory, and Oasis lists limited memory over long horizons.
  • Sessions last minutes. Google DeepMind states that Genie 3 supports a few minutes of continuous interaction.
  • Control is limited. Genie 3 lists a limited action space and difficulty simulating other agents, and Oasis lists difficulty with precise inventory control.
  • Speed is hard to reach. The first Genie model ran at around 1FPS, and the Matrix-Game 2.0 authors say earlier interactive models were held back by lengthy inference steps.
  • Access can be restricted. Genie 3 was announced as a limited research preview, and the Genie authors chose not to release model checkpoints or training data.

Examples

  • GenieGoogle DeepMind (one author also at University of British Columbia) · 2024
  • DIAMONDUniversity of Geneva; University of Edinburgh; Microsoft Research · 2024
  • GameNGenGoogle Research; Google DeepMind; Tel Aviv University · 2024
  • OasisDecart; Etched · 2024
  • Genie 3Google DeepMind · 2025
  • Matrix-Game 2.0Skywork AI · 2025

3D and 4D world models

A 3D world model builds a scene as 3D structure, such as points, surfaces or a grid of occupied space, so the scene stays the same when seen from a new angle. It takes text, images, video or recorded sensor data and outputs a 3D scene that can be explored, rendered from any viewpoint or exported to games and simulators. Some versions also predict how the 3D scene will change over the next seconds, which self-driving research uses to forecast the road ahead.116117118+4 Inferred

How it works

  1. Take a prompt or a real recording. Marble accepts text, images, video or coarse 3D layouts, WonderWorld starts from a single image, and Waabi's UniSim starts from one recorded drive with camera and LiDAR data.
  2. Estimate depth and shape. Some systems first generate a 360-degree panorama, and HunyuanWorld 1.0 uses panoramic images as stand-ins for the whole world.
  3. Store the result in a 3D format such as Gaussian splats (a large set of small semi-transparent particles), meshes, or grids that mark which spaces are occupied.
  4. Render new views, let a user move through the scene, export it to a game engine, or edit it, for example by adding or removing cars.
  5. In forecasting versions such as OccWorld, a second model predicts the next occupancy grid and the car's own path from the previous grids.

Takes actions

Optional

Used for

virtual worlds and 3D contentgamestesting self-driving softwarerobot simulationplanning

Strengths

  • The scene is consistent from every angle because it is stored in 3D. The HunyuanWorld 1.0 authors give geometric consistency as the advantage of 3D-based methods over video-based ones.
  • Results move into existing tools. HunyuanWorld 1.0 exports meshes, and Marble exports Gaussian splats, meshes and videos.
  • A recorded drive can be replayed with changes. Waabi's UniSim converts a single recorded log into a closed-loop multi-sensor simulation in which actors can be added, removed or moved.
  • Generation can be fast enough for interactive editing. WonderWorld generates connected 3D scenes in less than 10 seconds on a single A6000 GPU.

Limits

  • There is less 3D training data than video. The HunyuanWorld 1.0 authors say 3D-based methods struggle with limited training data and memory-inefficient representations.
  • Fast movement is hard to represent. The 2025 embodied AI survey reports that scenes built from renderable pieces such as Gaussian splats have limited ability to handle rapid dynamics or changes in shape.
  • Physics support is basic in current products. World Labs describes Marble's collider meshes as low-fidelity meshes intended for coarse physics simulation.
  • Text and single-image prompts give limited control over details, as World Labs states for Marble.
  • Accurate 3D needs depth sensors or many camera views, which the 2026 robotic manipulation survey notes are less available than ordinary video.

Examples

Learned and hybrid physics simulators

A learned physics simulator is a neural network trained to predict the next physical state of a system, such as the positions and speeds of particles, mesh points or robot joints. It takes the current state and any forces or commands and outputs the state a short time later, and repeating this step produces a full motion. It is used to replace or speed up a traditional physics engine, and in robotics to train controllers inside the learned simulator and to adjust the simulator with real-world data.4911+1 Inferred

How it works

  1. Run a traditional physics simulator, or record real measurements, to collect many examples of a state and the state one time step later.
  2. Describe the system as a graph in which particles, mesh points or robot parts are nodes and nearby or connected pieces are linked.
  3. Train a network that passes messages along the links and predicts how each node moves in the next step.
  4. Apply the network repeatedly to roll the simulation forward for hundreds or thousands of steps. GNS adds noise to its training data to limit the build-up of errors.
  5. In robotics, plug the learned model into a simulator in place of the hand-written dynamics and contact code, train policies there and fine-tune the model with real robot data, as NeRD does.

Takes actions

Optional

Used for

speeding up physics simulationscience and engineering simulationrobot simulationtraining policies

Strengths

  • It can run faster than the solver it learned from. MeshGraphNets runs 1-2 orders of magnitude faster than the simulation on which it is trained.
  • It can handle larger systems than it was trained on. GNS generalises to thousands of timesteps and at least an order of magnitude more particles at test time.
  • It can learn from real measurements. The NeRD authors report that their learned simulators can be fine-tuned from real-world data, unlike most classical simulators.
  • The outputs are physical quantities such as positions and speeds, so they can be compared directly with measurements.

Limits

  • Errors build up over long simulations, so training needs extra measures such as the added noise used in GNS.
  • Many learned simulators need training for each application and do not transfer to new tasks or environments, as the NeRD authors state.
  • The examples listed here take the system as particles, meshes or joint states and do not start from raw camera images.
  • Training examples usually come from a classical simulator, so the learned model can be no more accurate than that data.

Examples

How recent surveys group world models
SurveyTheir groupsHow they match the types on this page
A Comprehensive Survey on World Models for Embodied AIFunctionality axis: Decision-Coupled (task-specific) and General-Purpose (task-agnostic) world models, Temporal axis: Sequential Simulation and Inference (one step at a time) and Global Difference Prediction (many future steps at once), Spatial representation axis: Global Latent Vector; Token Feature Sequence; Spatial Latent Grid (bird's-eye-view or voxel grids); Decomposed Rendering Representation (NeRF, 3D Gaussian splatting)The survey sorts each model along three separate axes instead of naming families, so each of our families matches a combination of cells. Our latent-dynamics family matches decision-coupled, step-by-step models with a global latent vector, where the survey lists World Models 2018, PlaNet and Dreamer. Our video-generation family matches general-purpose models that predict many steps at once with token features, where the survey lists Sora. Our action-video family matches general-purpose step-by-step models with token features, where the survey lists iVideoGPT, Genie and Vid2World, and it moves to the decision-coupled side when a video model is tied to one robot task. The survey lists V-JEPA and V-JEPA 2 in the same cell as Sora, so it does not separate feature prediction from pixel prediction at the top level. Our 3d-world family matches two of its representation classes, spatial latent grids (for example the occupancy model OccWorld) and decomposed rendering representations (for example GaussianWorld). The survey has no class for real-time interactive worlds and does not discuss GameNGen or DIAMOND. It also has no class for learned physics simulators, and physics-based models such as PIN-WM and ParticleFormer appear inside its decision-coupled cells.
Agentic World Modeling: Foundations, Capabilities, Laws, and BeyondCapability levels: L1 Predictor (one-step prediction); L2 Simulator (multi-step, action-conditioned rollouts that respect domain laws); L3 Evolver (revises its own model when predictions fail against new evidence), Governing-law regimes: physical, digital, social, scientific, Physical-world system types: physics simulation (classical engines); video generation models, split into appearance-first video generation, action-conditioned and interactive video worlds, and decision-oriented video world models; robotics and sim-to-real transfer; spatial reasoning; 3D-structured world models; autonomous driving world models; game world models (placed between the physical and digital regimes), Architecture axes: representation (symbolic, latent continuous, structured 3D, discrete tokens); dynamics (stochastic latent, deterministic value-aware, autoregressive token, diffusion-based); control interfaceAll seven of our families fall inside its physical regime, with games placed between the physical and digital regimes. Its L1 to L3 levels describe how capable a model is and cut across our families, so any of our families can contain L1 or L2 systems. Its group of action-conditioned and interactive video worlds covers both our action-video and interactive-world families, and lists Genie, GAIA-1, Oasis and Matrix-Game 3.0 together. Its appearance-first video generation group (Sora, Lumiere, VideoPoet) matches our video-generation family. Its architecture section separates stochastic latent dynamics (DreamerV3) from deterministic value-aware dynamics (MuZero, TD-MPC2), which are both in our latent-dynamics family, and it groups V-JEPA2 with DreamerV3 as latent continuous representations. Its 3D-structured world models (Marble, RTFM, TesserAct, RoboOccWorld) and its driving occupancy models (OccWorld) match our 3d-world family. Its physics simulation heading covers classical engines such as MuJoCo and Isaac Lab, which it says are not learned simulators, so our learned-simulator family has no direct match; the closest material is its fine-grained physical representations (ParticleFormer, GWM) and its scientific-regime operator learning. Its digital, social and scientific regimes (web and GUI agents, social simulation, AI for science) are outside our scope.
World Models for Robotic Manipulation: A SurveyRepresentation families: image and video; learned latent; motion fields and scene flow; geometric and spatiotemporal; physics-informed dynamics, Functional taxonomy: integrated prediction-action models; explicit predictive planners, Infrastructure roles: synthetic experience generation; candidate-action filtering and refinement; search-based action evaluation; learned environments for policy evaluation and improvement; outcome scoring and feasibility verification, Learning lifecycle: pretraining; post-training; inference adaptationIts image and video family covers our video-generation and action-video families, and lists UniPi, SuSIE, GR-1, DreamGen, WorldGym and Genie Envisioner there. Its learned latent family covers our latent-dynamics family and also contains V-JEPA 2, so it does not separate our predictive-representation family. Its geometric and spatiotemporal family matches our 3d-world family, with a focus on predicting how a 3D scene changes under robot actions (TesserAct, PointWorld, GWM). Its physics-informed dynamics family overlaps our learned-simulator family and also includes hybrids that add learned parts to differentiable physics (PIN-WM). Its motion fields and scene flow family (FLIP, FlowVLA) has no counterpart in our list. Its integrated prediction-action models (GR-1, GR-2, WorldVLA) have no counterpart either, and we propose a world-action family for them. The survey covers robot manipulation only, so real-time interactive worlds for games are absent. Its infrastructure roles match our used_for values. Synthetic experience generation corresponds to generating training data, and learned environments correspond to evaluating and training policies.
04

Who builds world models

36 organisations are listed here. The table shows which types of world model each one has released. A filled dot means at least one of those models has public weights.

OrganisationLatent dynamicsPredictive representationAction-conditioned videoWorld-actionVideo generationInteractive worlds3D and 4D worldsLearned physicsLargest disclosed round
Large technology companies
Alibaba (Tongyi Wan, DAMO Academy)ChinaNot disclosed
Ant Group (Robbyant / LingBot)ChinaNot disclosed
ByteDance (Seed)ChinaNot disclosed
Google DeepMindEuropeNot disclosed
Kuaishou (Kling)ChinaNot disclosed
Manycore Tech (SpatialVerse)ChinaNot disclosed
Meta (FAIR)North AmericaNot disclosed
Microsoft (Research and Xbox)North AmericaNot disclosed
NVIDIANorth AmericaNot disclosed
OpenAINorth AmericaNot disclosed
SenseTimeChinaNot disclosed
Skywork AI (Kunlun Tech)ChinaNot disclosed
Tencent (Hunyuan)ChinaNot disclosed
TeslaNorth AmericaNot disclosed
Start-ups
AMI Labs (Advanced Machine Intelligence)Europe$1.03 billionMar 2026
DecartSeveral regions$300 millionMay 2026
Dynamics Lab—
General IntuitionNorth America$320 millionJun 2026
GigaAI (极佳科技)China—
Luma AINorth America—
OdysseyNorth America$310 millionJun 2026
RunwayNorth America$315 millionFeb 2026
Shengshu Technology (生数科技, Vidu)China—
WaabiNorth America$750 millionJan 2026
WayveEurope$1.2 billionFeb 2026
World LabsNorth America$1 billionFeb 2026
Robot companies
1X TechnologiesSeveral regions$100 millionJan 2024
AgiBot (智元机器人)China—
Figure AINorth America$1 billionSep 2025
Galbot (银河通用)China—
Physical IntelligenceNorth America$600 millionNov 2025
Skild AINorth America$1.4 billionJan 2026
Unitree Robotics (宇树科技)China—
WaymoNorth America$16 billionFeb 2026
Research labs
Beijing Academy of Artificial Intelligence (BAAI)China—
Shanghai AI LaboratoryChina—

Disclosed funding

Each bar is the largest funding round an organisation has disclosed in US dollars. This is money raised by the whole company, not money spent on world models. Select a bar to see every round and its exact wording. The first bar is shortened, shown by the break, so that the others stay readable. A round marked Secondary was found only in news reports. Large technology companies do not disclose how much they spend on world models.

Funding rounds of the Chinese start-ups in the table (AgiBot, Galbot, Unitree, GigaAI, Shengshu Technology and Manycore Tech) have not been checked yet, so their funding is left blank.

05

Where world models are used

14 uses are listed. 2 are at the research stage, 6 are in pilots and 6 are offered as products.

Research2

Shown in papers or demos.

Pilot6

Used by a company in its own work, or tested with customers.

Product6

Offered to customers.

06

How world models are tested

Benchmarks for world models check how realistic the generated video is, whether it follows physics, or whether a policy gets the same result inside the world model as on a real robot.

Do world models rank robot policies the way real robots do?

Each row is one published comparison between scores inside a world model and scores on real robots. The value is the Pearson correlation r: 1.0 means the two sets of scores rise and fall together exactly. A hollow dot is the figure reported in the world model’s own paper. A filled dot is a figure measured by another team. A line shows a range of per-task values. Select a row to open its source.

0.40.50.60.70.80.91.0Reported in the world model’s own paperMIDDLE VALUE 0.87DreamDojo vs real-robot fruit-packing success: Pearson r = 0.995; MMRV = 0.003DreamDojo6 checkpoints of one policy0.995Interactive World Simulator vs real-robot task scores: r = 0.8553 (T Pushing), 0.8455 (Rope Routing), 0.8869 (Mug Grasping), 0.9908 (Pile Sweeping) (Fig. 7; the paper labels the coefficient r without naming it)Interactive World Simulator4 tasks, no overall valueRoboWorld vs RoboArena leaderboard: Pearson r = 0.989; Spearman rho = 0.970RoboWorld8 policies, against RoboArena0.989GWM-Robotics vs RoboArena real-world results: Pearson correlation = 0.95 (figure label: Pearson r = 0.952); MMRV = 0.033Runway GWM-Robotics8 policies0.95WorldEval vs real-robot success rates: Table 1 (3 tasks): Pearson r = 0.958 (Place Cup), 0.887 (Strike Block), 0.980 (Handover Block), average 0.942; MMRV = 0.000, 0.133, 0.000, average 0.044. Figure 4 (5 tasks): Bussing Table r = 0.935, MMRV = 0.000; Collect Toy r = 0.885, MMRV = 0.133; Place Cup r = 0.958, MMRV = 0.000; Handover Block r = 0.980, MMRV = 0.000; Strike Block r = 0.887, MMRV = 0.000.WorldEval4 policies, 3 tasks (average)0.942dWorldEval vs real AgileX robot success: r = 0.918, MMRV = 0.02; no-memory ablation r = 0.829, MMRV = 0.033 (Fig. 7d)dWorldEval3 policies, 5 tasks0.918Veo (Robotics) nominal scenes vs real ALOHA 2 evaluations: Pearson = 0.88; MMRV = 0.03 (Figure 4)Veo (Robotics)8 checkpoints0.88PlayWorld vs real-robot success (18 policies): PlayWorld: Pearson r = 0.8766, RMSE = 0.171. Human-demo-trained model: r = 0.6636, RMSE = 0.275. Human-play-trained model: r = 0.6619, RMSE = 0.297 (Fig. 7).PlayWorld18 policies0.8766WEAVER vs real-robot success (five tasks): WEAVER-FT (Table 8): Pearson = 0.863, Spearman = 0.870, MMRV = 0.035, RMSE = 0.188; abstract and Fig. 6 report rho = 0.870. Pretrained WEAVER: Pearson = 0.563, Spearman = 0.594, MMRV = 0.155, RMSE = 0.359.WEAVER (fine-tuned)2 policies, 5 tasks0.863OSCAR vs RoboArena real-world success: Skeleton: MMRV 0.571 (scale 0-6), Spearman rho +0.750, Pearson r +0.852, SISR_delta 1.73 pp. Latent action: MMRV 1.429, rho +0.643, r +0.867, SISR_delta 1.98 pp. Mesh: MMRV 0.714, rho +0.679, r +0.781, SISR_delta 3.04 pp (Table 4).OSCAR (skeleton)7 policies, against RoboArena0.852DexTouch-WM vs real dexterous-robot scores: Per task (Place Shoes / Place Phone / Stack Bowls / Stand Bottle): WM-Robot Pearson 0.400 / 0.945 / 0.906 / 0.334, MMRV 0.033 / 0 / 0 / 0.067 (mean Pearson 0.646, mean MMRV 0.025); WM-Mix Pearson 0.972 / 0.939 / 0.889 / 0.577, MMRV 0 / 0 / 0 / 0.067 (mean Pearson 0.844, mean MMRV 0.017).DexTouch-WM (WM-Mix)3 policies per task, 4 tasks (mean)0.844PersistWorld vs real-robot task progress: Pearson r = 0.822 (p = 0.007); MMRV = 0.006PersistWorld3 policies, 3 tasks0.822WorldGym vs OpenVLA Bridge real-robot results: Pearson r = 0.78 between per-task success rates in WorldGym and in the real world (each point is one task-policy pair). Mean success rate, real vs WorldGym: RT-1-X 18.5% vs 15.5%, Octo 20.0% vs 23.82%, OpenVLA 70.6% vs 67.4%; average difference 3.3%. The order of the three policies by mean success rate is the same in both.WorldGym3 policies0.78Cosmos-Surg-dVRK (human labels) vs real dVRK success rates: Pooled Pearson r = 0.718 (p < 0.001) across all tasks and training regimes. Per task Pearson / MMRV: Handover 0.468 / 0.217; Throw 0.716 / 0.183; Knot Tie 0.840 / 0.050; Pickup 0.806 / 0.067; average 0.707 / 0.129 (Table 2: Pearson 0.71 +/- 0.17, MMRV 0.13 +/- 0.08). Mean bias error 0.140 (95% CI 0.081-0.199).Cosmos-Surg-dVRK6 checkpoints (human labels)0.718Cosmos-based world model vs real Bridge-setup success (NVIDIA): Pearson = 0.687; MMRV = 0.171 (Fig. 6)NVIDIA Cosmos-based model3 policies, 4 tasks0.687Measured in another team’s paperMIDDLE VALUE 0.58Ctrl-World vs real-robot task progress (measured in the PersistWorld paper): Pearson r = 0.796 (p = 0.010); MMRV = 0.053Ctrl-WorldIn the PersistWorld paper0.796IRASim vs real Bridge-setup success (measured by NVIDIA): Pearson = 0.613; MMRV = 0.611 (Fig. 6)IRASimIn the NVIDIA paper0.613Ctrl-World vs real-robot success (measured in the WEAVER paper): Pearson = 0.552; Spearman = 0.523; MMRV = 0.215; RMSE = 0.410 (Table 8; Fig. 6 shows rho = 0.523, MMRV = 0.215)Ctrl-WorldIn the WEAVER paper0.552Ctrl-World as a baseline evaluator in the PolaRiS study: Pearson r = 0.53; MMRV = 0.22 (PolaRiS Figure 7). Same figure, other evaluators: PolaRiS r = 0.90, MMRV = 0.03; LIBERO-90 fine-tuned checkpoints at 1k / 10k / 50k steps r = 0.66 / 0.70 / 0.66, MMRV = 0.19 / 0.04 / 0.15; action MSE r = -0.55 (train) and -0.53 (validation), MMRV = 0.40 for both.Ctrl-WorldIn the PolaRiS paper, which shares an author with Ctrl-World0.53
Sources: the papers linked from each row383639385+41. Values are copied from each paper. Studies compare different policies and tasks, so the values are not directly comparable.
5 more comparisons report no correlation

Ctrl-World vs real-robot rollouts on the authors' DROID setup. No correlation coefficient or MMRV reported. Linear fits of world-model rate on real rate across policy-task pairs (Figure 7): instruction following y = 0.87x - 0.04; success rate y = 0.81x - 0.11.40240181

1X World Model outcome-prediction alignment. Alignment 63.06% when trained on about 216M Shelf video tokens; 71.17% with an added about 1.46B Arcade video tokens391

GigaWorld-1 closed-loop success-rate alignment with real robots. No correlation statistic. Fitted line of generated vs real success rate: y = 1.134x - 0.091 (Fig. 16). Generated minus real success per subtask ranges from -0.13 to +0.07 across 12 subtasks (Fig. 17).452645646+1

EnerVerse-AC vs real-robot success (four tasks, three training steps). No statistic; bar charts show the same task ranking and the same training-step trend. Real vs EVAC per task: Take a Bottle 28% vs 25%; Take a Toast 100% vs 90%; Take a Bacon 85% vs 88%; Take a Leaf 55% vs 50%. Per training step (Take a Bottle): 4K 40% vs 40%; 8K 61% vs 63%; 13K 79% vs 76% (Fig. 7).373374372

RoboWM-Bench real-to-sim outcome consistency. Success consistency 10/10 and failure consistency 10/10 for every task (Pick Object, Pull Object, Push Object, Put on Plate, Discard Trash, Close Drawer, Put in Drawer): 140 of 140 outcomes matched482483

Benchmarks for world models 73

BenchmarkWhat it measuresBuilt byReleased
1X World Model Challenge
Entry on this site
How well models predict future first-person frames of 1X's EVE humanoid from past frames and actions.3523531X Technologies (2025 phases with OpenDriveLab)Jun 2024
EVA-Bench (Embodied Video Anticipation Benchmark)Embodied video anticipation by world models that combine a vision-language model and a video generator, across four meta-tasks: Action-Description, How-To, Finish-Thinking and Next-Step, on real robots, simulated robots and egocentric human activities, with in-domain and out-of-distribution samples.354355356Hong Kong University of Science and Technology, Peking University (State Key Laboratory of Multimedia Information Processing)20 Oct 2024
WorldSimBenchVideo generation models used as world simulators for embodied agents, in three scenarios: an open-ended embodied environment (Minecraft via MineRL), autonomous driving (CARLA) and robot manipulation (CALVIN). It checks perceived visual quality per embodied dimension and whether generated videos can be turned into correct control signals.357358359+3arXiv and project page: The Chinese University of Hong Kong, Shenzhen; Shanghai Artificial Intelligence Laboratory; Beihang University; The University of Hong Kong. PMLR version adds Sun Yat-sen University, University of Oxford and the Guangdong Key Laboratory of Big Data Analysis and Processing.23 Oct 2024
EWMBench
Entry on this site
How closely robot-manipulation videos from a video generator match real AgiBot World episodes in scene, motion and task meaning.370AgiBot; Shanghai Jiao Tong University; MMLab-CUHK; Harbin Institute of Technology (arXiv header)May 2025
AgiBot World Challenge, World Model track
Entry on this site
How well models predict robot head-camera video from actions on held-out AgiBot World episodes.371AgiBot (2025 with OpenDriveLab)May 2025
EnerVerse-AC (EVAC) as policy evaluatorWhether success rates of a Go-1 policy evaluated inside the EVAC action-conditioned world model match real-robot success rates across four retrieval tasks and across three training steps.372373374+3AgiBot; Shanghai Jiao Tong University; MMLab, CUHK14 May 2025
DreamGen BenchHow well image-to-video world models adapt to a target robot embodiment and generalise to unseen objects, behaviours and environments: whether generated robot videos follow the instruction and obey physics. Setups: RoboCasa (simulated Franka) and three real Fourier GR1 humanoid splits (Object, Behavior, Environment).54377378+4NVIDIA (GEAR lab, lead), University of Washington, KAIST, UCLA, UCSD, Caltech, NTU, University of Maryland, UT Austin19 May 2025
WorldEvalWhether success rates of real-robot manipulation policies, obtained by rolling each policy out inside an action-conditioned video world model, track and rank the same policies' real-robot success rates.298383299+1Midea Group; East China Normal University25 May 2025
WorldGymWhether success rates of VLA policies, obtained by Monte Carlo rollouts in an autoregressive action-conditioned video world model started from real first frames, match real-robot success rates and keep policy rankings. Also used to test policies on edited out-of-distribution scenes and instructions.297385386+3Stanford University; NYU; Google DeepMind31 May 2025
Real-robot goal-image planning test (V-JEPA 2-AC)Whether an action-conditioned world model can drive a real robot by planning: the robot is given goal images and the world model is used to search for actions, with no task-specific training or reward.390FAIR at Meta (V-JEPA 2 paper)Jun 2025
1X World Model evaluationHow often the world model's success/failure prediction matches real outcomes, and whether its scores pick the same checkpoints and architectures as real double-blind A/B evaluations on humanoid robots.383911X Technologies16 Jun 2025
IRASim policy evaluation (LIBERO)Whether success rates of diffusion-policy checkpoints judged in IRASim-generated rollouts match the success rates measured in the LIBERO MuJoCo simulator.396397398+1Hong Kong University of Science and Technology; ByteDance Seed29 Jul 2025
Ctrl-WorldWhether the instruction-following rate and the task success rate of generalist DROID policies, rolled out closed-loop inside a multi-view action-conditioned video world model, match the same policies' rates on a real robot in a new DROID setup.80401402+5Stanford University; Tsinghua University11 Oct 2025
Cosmos-Surg-dVRKWhether success rates of surgical robot policies, rolled out online in a fine-tuned Cosmos world foundation model, match success rates of the same policies on a real da Vinci Research Kit (dVRK Si).406407408NVIDIA; Johns Hopkins University; Stanford University17 Oct 2025
World-in-WorldWhether visual world models help an embodied agent succeed in closed-loop tasks: Active Recognition (AR), Image-Goal Navigation (ImageNav), Active Embodied Question Answering (A-EQA) and robotic manipulation. Each world model is plugged into the same proposal-simulation-revision planning loop through a unified action API (text prompt, camera trajectory or low-level actions).409410411+4Johns Hopkins University (lead; corresponding author Jieneng Chen), Peking University, Princeton University, MIT, Harvard University20 Oct 2025
Scalable Policy Evaluation with Video World Models (NVIDIA)Whether success rates predicted by action-conditioned video models (post-trained Cosmos-Predict2-2B) with a VLM success judge match simulator success rates on four RoboMimic tasks and real success rates of three generalist policies on four Bridge-setup tasks.416417418NVIDIA Research; University of Toronto; Vector Institute14 Nov 2025
Veo (Robotics) policy evaluationWhether success rates of Gemini Robotics On-Device policy checkpoints, predicted by closed-loop rollouts in a Veo-based action-conditioned multi-view video model, match real ALOHA 2 success rates in nominal scenes and in edited out-of-distribution scenes, and whether the model finds unsafe behaviour that also occurs on the real robot.87419420Google DeepMind (Gemini Robotics Team)11 Dec 2025
PolaRiS
Entry on this site
Whether policy scores in simulated copies of real scenes, built from short video scans with 2D Gaussian splatting, match real-robot scores of generalist DROID policies.421422423+2University of Washington; Princeton University; UC Berkeley; Stanford University; Toyota Research Institute; University of Southern California; Cornell University; Physical Intelligence18 Dec 2025
RBench
Entry on this site
Whether video generators produce correct and physically plausible robot task videos across five task types and four robot body types.426Peking University; ByteDance SeedJan 2026
WoW-World-Eval (Wow, wo, val)Image-plus-text-to-video generation for robot manipulation, framed as an 'Embodied Turing Test' over five abilities: perception, planning, prediction, generalisation and execution. Includes a human Turing test (can people tell generated from real video) and an inverse-dynamics-model (IDM) Turing test (can actions recovered from generated video be executed on a real robot).427428429Peking University (State Key Laboratory of Multimedia Information Processing, School of Computer Science), Beijing Innovation Center of Humanoid Robotics, The Hong Kong University of Science and Technology7 Jan 2026
DreamDojo policy evaluation (AgiBot fruit packing)Whether the success rates of policy checkpoints simulated inside the DreamDojo world model match the success rates of the same checkpoints on a real robot, on one long-horizon fruit-packing task.93430431+2NVIDIA (lead); co-author affiliations also list HKUST, UC Berkeley, KAIST, University of Toronto, UC San Diego, University of Washington, Stanford, UT Austin6 Feb 2026
WorldArenaHow well embodied world models (text- or action-conditioned robot video models) predict bimanual manipulation video, and whether they are useful for three downstream jobs: generating training data for a policy (data engine), standing in for a simulator when scoring policies (policy evaluator), and planning actions (action planner). Human ratings are collected as a third view.433434435+10Tsinghua University (lead; corresponding author Yong Li), Shanghai Jiao Tong University, The University of Hong Kong, Princeton University, Chinese Academy of Sciences, University of Science and Technology of China, Peking University, National University of Singapore9 Feb 2026
GWM-Robotics policy evaluationWhether progress scores of VLA policies rolled out inside Runway's GWM-Robotics world model rank the policies in the same order as their real-world RoboArena evaluations.290224Runway27 Feb 2026
GigaBrain Challenge 2026 - World Model Track (CVPR 2026 workshop)World models as evaluators of a VLA policy (GigaBrain) on 8 real-robot manipulation tasks: video quality when replaying teleoperation actions, and whether closed-loop rollouts driven by policy actions reach the same outcome as the real-robot reference video.446447448+4GigaAI (organizers Zheng Zhu, Xiaofeng Wang) with co-organizers from University of Hong Kong, Peking University, Shanghai Jiao Tong University, RoboChallenge/Dexmal and Horizon RoboticsMar 2026
WorldArena Challenge @ CVPR 2026Two tracks built on WorldArena: Track 1 video perception quality of embodied world models; Track 2 whether generated worlds work as data engines and as policy evaluators.453454437+4Organisers listed on the challenge page: Amap CV Lab (AMAP), Manifold AI, Tsinghua University, Princeton University, National University of Singapore, The University of Hong KongMar 2026
PlayWorld policy evaluationWhether success rates predicted by a video world model trained on autonomous robot play match real-robot success rates for many manipulation policies on contact-rich tasks, and whether predicted failure modes match observed ones.456457458+4Princeton University9 Mar 2026
Interactive World Simulator sim-to-real policy evaluationWhether task scores of imitation policies (final and intermediate checkpoints of DP, ACT, pi0 and pi0.5) evaluated closed-loop inside the learned world model match their real-robot scores on four ALOHA tasks.463464465+4Columbia University; Toyota Research Institute; Amazon; University of Illinois Urbana-Champaign9 Mar 2026
PersistWorld world-model-to-real policy evaluationWhether the task progress of robot policies rolled out inside an action-conditioned video world model matches their real-robot task progress, comparing PersistWorld (Ctrl-World post-trained with reinforcement learning on its own rollouts) with the base Ctrl-World.470471472+3Czech Institute of Informatics, Robotics and Cybernetics, Czech Technical University in Prague26 Mar 2026
Cosmos-H-Surgical-Simulator evaluation on Open-H-EmbodimentHow closely an action-conditioned surgical world model's generated video follows the recorded kinematic actions of held-out real surgical-robot episodes (open-loop replay), as groundwork for in-silico policy evaluation.476477478+2NVIDIA with the Open-H-Embodiment Consortium (50+ institutions in the paper's author list); model card by NVIDIAApr 2026
RoboWM-BenchWhether manipulation behaviour in videos generated by world models can be executed: generated human-hand or robot-arm videos are converted to robot actions and run in simulated scenes, including real-to-sim reconstructions of real tabletop scenes.481482483+2Peking University (lead; corresponding author Ruihai Wu), Tsinghua University, Lightwheel21 Apr 2026
dWorldEval policy evaluation proxyWhether success rates estimated inside a discrete-diffusion world model (with a progress token that marks task completion) match ground-truth success for pi0 checkpoints on LIBERO, heterogeneous policies on RoboTwin, and three policies on five real bimanual AgileX tasks; it also compares three video-diffusion world models on LIBERO.486487488+2Current Robotics; University of Toronto24 Apr 2026
WorldArena 2.0Extends WorldArena along three axes: visuotactile prediction (tactile plus video) for contact-rich tasks, world models used as interactive reinforcement-learning environments for policy improvement, and evaluation across two simulators and a real robot (RoboTwin 2.0, LIBERO, AgileX split-type ALOHA).492493494+4Tsinghua University (lead; corresponding author Yong Li), Shanghai Jiao Tong University, Zhejiang University, Stanford University, The University of Hong Kong, Princeton University, Chinese Academy of Sciences, University of Science and Technology of China, Peking University, National University of Singapore18 May 2026
OSCAR policy evaluation on RoboArenaWhether OSCAR-generated replays of RoboArena episodes, judged by GPT-5, reproduce the real per-policy success rates and ranking of seven open-source DROID policies.499500Peking University; University of Michigan; NVIDIA3 Jun 2026
WEAVER policy evaluation on real hardwareWhether success judged on rollouts imagined by the world model matches the real success rates of two pi0.5-based policies on five real manipulation tasks, compared with Ctrl-World and with the WEAVER model before task fine-tuning.501502503+5Mila - Quebec AI Institute; Universite de Montreal; Carnegie Mellon University; McGill University11 Jun 2026
RoboWorld
Entry on this site
Whether scores from closed-loop rollouts of DROID policies in an autoregressive video world model match the RoboArena real-world leaderboard.509510511KAIST; Config1 Jul 2026
WMBench (GigaWorld-1)How well video world models serve as surrogate evaluators of robot policies, by comparing generated rollouts with matched real-robot executions on eight manipulation tasks built from teleoperation data and GigaBrain policy rollouts.295452512+3GigaAI; Tsinghua University2 Jul 2026
WorldArena 2.0 Challenge @ IROS 2026Three tracks: Track 1 video quality of embodied world models with harder tasks and new out-of-distribution scenes; Track 2 world models as RL environments for policy optimisation; Track 3 real-world manipulation by world action models (WAM), in tactile and vision-only settings.514498497+2WorldArena 2.0 team (organisers not named on the challenge page; contact worldarenav2@outlook.com); paper team led by Tsinghua University10 Jul 2026
TriWorldBenchWhether an embodied world model's synchronized head, left-wrist and right-wrist videos describe one consistent manipulation event, plus task alignment, physical and 3D coherence, motion quality, temporal consistency and visual quality.515516517+4Peking University, Tsinghua University, Beihang University, Shanghai Jiao Tong University, University of Science and Technology of China, Shanghai AI Laboratory (arXiv affiliation block); the website lists OpenCompass in place of Shanghai AI Laboratory28 Jul 2026
Pelican-Sim 1.0 policy evaluation and ranking (RoboTwin)Whether the Pelican-Sim world model combined with a fine-tuned VLM judge reproduces RoboTwin simulator success rates and the ranking of five checkpoints from one VLA training run.522523524+1Beijing Innovation Center of Humanoid Robotics (X-Humanoid), WFM System Group10 Sep 2026
DexTouch-WM world models as policy evaluatorsWhether scores of three VLA policies rolled out closed-loop inside two task-adapted visuo-tactile world models match their scores on a real dexterous-hand robot on four tasks.526527HKUST (Guangzhou); Xspark AI; Peking University; Tsinghua University; University of Hong Kong17 Sep 2026

Known problems

  • Almost every published agreement number between a world-model evaluator and real robots was measured by the world model's own builders, on 1 to 18 policies or checkpoints, often 3 to 8. Third-party checks without shared authors exist only for Ctrl-World (PersistWorld and WEAVER papers on real robots; dWorldEval and WorldArena against simulators), WorldGym (dWorldEval, simulator only), Cosmos-Predict 2.5 (WorldArena, simulator only) and IRASim (NVIDIA, real robots). Several studies reuse real-robot numbers from earlier work instead of running new paired trials (WorldGym from OpenVLA, RoboWorld and Runway from RoboArena).385510290+1 Inferred
  • The same world model gets very different agreement scores depending on who tests it and where. Ctrl-World's agreement with ground truth was Pearson r = 0.53 / MMRV 0.22 in the PolaRiS real-robot study, r = 0.552 / MMRV 0.215 in the WEAVER real-robot study, r = 0.796 / MMRV 0.053 in the PersistWorld real-robot study, r = 0.841 against the LIBERO simulator in the dWorldEval study, and r = 0.986 against the RoboTwin simulator in WorldArena. Its own paper reports no coefficient.422501470+3 Inferred
  • The evaluators that third parties cannot run are the ones built by companies. Google DeepMind's Veo (Robotics) evaluator, Runway's GWM-Robotics, 1X's world model and Wayve's GAIA-3 have no released code or weights for evaluation, so their agreement numbers cannot be reproduced; RoboWorld's code was marked 'coming soon' on its project page.420290391+1 Inferred
  • World-in-World finds that visual quality does not predict closed-loop task success: on its leaderboard of 16 Active Recognition entries, zero-shot Cosmos-Predict2 has the highest generation-quality score (0.481732) but the lowest success rate (55.35%, tied with Wan2.2 5B), while controllability (1 - LPIPS) tracks success more closely.410647412
  • Vision-language judges, which several world-model evaluators use to decide whether a task succeeded, are error-prone. On FailBench (2,197 manipulation attempts from 14 public sources, 12 real and 2 simulated), the best of 13 VLM-based detectors reached 0.77 mean balanced accuracy, fell below 0.60 on contact-intensive assembly, and leaned toward predicting success when evidence was ambiguous.648
  • World models can predict higher success than real robots achieve. DreamDojo's authors report that its absolute success rates are often higher than real ones; in its Fig. 5a, DreamDojo rates span about 0.06-0.81 while real rates span 0.0-0.44.430431
  • Generated videos can outscore real videos on automated benchmarks. On the PAI-Bench generation leaderboard, Cosmos3-Super (83.9) and Cosmos3-Nano (83.7) score above the real source videos (82.6) overall, and in the robot domain Cosmos3-Super (90.0), Cosmos3-Nano (90.2) and Veo-3 (86.9) score above the source videos (86.2).649
  • Contact-rich, granular and deformable interactions remain hard to simulate. WEAVER names bean pouring and bag and towel handling as the hardest tasks; Veo (Robotics) reports objects appearing spontaneously during contact and its weakest out-of-distribution correlation on novel objects (Pearson 0.56).502419650

86 problems are recorded in total. All of them are in the downloadable data.

07

Media to follow

Blogs, newsletters, channels and accounts to follow on world models. Each one posted something in the 90 days before 10 Oct 2026. These are our picks.

BlogEN

Runway Research

Runway

Research hub for Runway's general world models (GWM), real-time interactive worlds, an interface world model (Solaris) and a world action model for robots (Praxis-1).

Start withIntroducing Praxis-1

runwayml.com30 Sep 2026 · Irregular
08

Sources 650

  1. 1Action-Conditional Video Prediction using Deep Networks in Atari GamesPaper · 31 Jul 2015 · checked 10 Oct 2026
  2. 2Unsupervised Learning for Physical Interaction through Video PredictionPaper · 23 May 2016 · checked 10 Oct 2026
  3. 3Deep Visual Foresight for Planning Robot MotionPaper · 3 Oct 2016 · checked 10 Oct 2026
  4. 4Interaction Networks for Learning about Objects, Relations and PhysicsPaper · 1 Dec 2016 · checked 10 Oct 2026
  5. 5World ModelsPaper · 27 Mar 2018 · checked 10 Oct 2026
  6. 6Learning Latent Dynamics for Planning from PixelsPaper · 12 Nov 2018 · checked 10 Oct 2026
  7. 7Mastering Atari, Go, Chess and Shogi by Planning with a Learned ModelPaper · 19 Nov 2019 · checked 10 Oct 2026
  8. 8Dream to Control: Learning Behaviors by Latent ImaginationPaper · 3 Dec 2019 · checked 10 Oct 2026
  9. 9Learning to Simulate Complex Physics with Graph NetworksPaper · 21 Feb 2020 · checked 10 Oct 2026
  10. 10Mastering Atari with Discrete World ModelsPaper · 5 Oct 2020 · checked 10 Oct 2026
  11. 11Learning Mesh-Based Simulation with Graph NetworksPaper · 7 Oct 2020 · checked 10 Oct 2026
  12. 12Yann LeCun on a vision to make AI systems learn and reason like animals and humansOfficial blog · 23 Feb 2022 · checked 11 Oct 2026
  13. 13Temporal Difference Learning for Model Predictive ControlPaper · 9 Mar 2022 · checked 10 Oct 2026
  14. 14Video Diffusion ModelsPaper · 7 Apr 2022 · checked 10 Oct 2026
  15. 15DayDreamer: World Models for Physical Robot LearningPaper · 28 Jun 2022 · checked 10 Oct 2026
  16. 16Transformers are Sample-Efficient World ModelsPaper · 1 Sep 2022 · checked 10 Oct 2026
  17. 17eloialonso/iris pretrained modelsRepository · May 2024 · checked 11 Oct 2026
  18. 18Mastering Diverse Domains through World ModelsPaper · 10 Jan 2023 · checked 10 Oct 2026
  19. 19Learning Universal Policies via Text-Guided Video GenerationPaper · 31 Jan 2023 · checked 10 Oct 2026
  20. 20Introducing GAIA-1: A Cutting-Edge Generative AI Model for AutonomyOfficial blog · 17 Jun 2023 · checked 11 Oct 2026
  21. 21GAIA-1: A Generative World Model for Autonomous DrivingPaper · 29 Sep 2023 · checked 10 Oct 2026
  22. 22DriveDreamer: Towards Real-world-driven World Models for Autonomous DrivingPaper · 18 Sep 2023 · checked 10 Oct 2026
  23. 23TD-MPC2: Scalable, Robust World Models for Continuous ControlPaper · 25 Oct 2023 · checked 10 Oct 2026
  24. 24nicklashansen/tdmpc2 checkpointsRepository · Oct 2023 · checked 11 Oct 2026
  25. 25Learning Interactive Real-World SimulatorsPaper · 9 Oct 2023 · checked 10 Oct 2026
  26. 26Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large DatasetsPaper · 25 Nov 2023 · checked 10 Oct 2026
  27. 27stabilityai/stable-video-diffusion-img2vidRepository · Nov 2023 · checked 11 Oct 2026
  28. 28Genie: Generative Interactive EnvironmentsPaper · 23 Feb 2024 · checked 10 Oct 2026
  29. 29Video generation models as world simulatorsOfficial blog · 15 Feb 2024 · checked 11 Oct 2026
  30. 30Revisiting Feature Prediction for Learning Visual Representations from VideoPaper · 15 Feb 2024 · checked 10 Oct 2026
  31. 31facebookresearch/jepa (README)Repository · Feb 2024 · checked 11 Oct 2026
  32. 32Diffusion for World Modeling: Visual Details Matter in AtariPaper · 20 May 2024 · checked 10 Oct 2026
  33. 33eloialonso/diamond pretrained modelsRepository · Oct 2024 · checked 11 Oct 2026
  34. 34Vista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilityPaper · 27 May 2024 · checked 10 Oct 2026
  35. 35OpenDriveLab/VistaRepository · Jun 2024 · checked 11 Oct 2026
  36. 36Diffusion Models Are Real-Time Game EnginesPaper · 27 Aug 2024 · checked 10 Oct 2026
  37. 371X World ModelOfficial blog · 17 Sep 2024 · checked 11 Oct 2026
  38. 381X World Model (policy evaluation update)Official blog · 16 Jun 2025 · checked 11 Oct 2026
  39. 39Oasis: A Universe in a TransformerOfficial site · 31 Oct 2024 · checked 11 Oct 2026
  40. 40Etched/oasis-500m model repositoryRepository · 31 Oct 2024 · checked 11 Oct 2026
  41. 41DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot PlanningPaper · 7 Nov 2024 · checked 10 Oct 2026
  42. 42gaoyuezhou/dino_wm (README)Repository · Jan 2025 · checked 11 Oct 2026
  43. 43Genie 2: A large-scale foundation world modelOfficial blog · 4 Dec 2024 · checked 11 Oct 2026
  44. 44State-of-the-art video and image generation with Veo 2 and Imagen 3Official blog · 16 Dec 2024 · checked 11 Oct 2026
  45. 45Cosmos World Foundation Model Platform for Physical AIPaper · 7 Jan 2025 · checked 10 Oct 2026
  46. 46nvidia/Cosmos-1.0-Diffusion-7B-Text2WorldRepository · 7 Jan 2025 · checked 11 Oct 2026
  47. 47World and Human Action Models towards gameplay ideation (Nature)Paper · 19 Feb 2025 · checked 11 Oct 2026
  48. 48Introducing Muse: Our first generative AI model designed for gameplay ideationOfficial blog · 19 Feb 2025 · checked 11 Oct 2026
  49. 49microsoft/wham model repositoryRepository · Feb 2025 · checked 11 Oct 2026
  50. 50Wan: Open and Advanced Large-Scale Video Generative ModelsPaper · 26 Mar 2025 · checked 10 Oct 2026
  51. 51Wan-AI/Wan2.1-T2V-14BRepository · Feb 2025 · checked 11 Oct 2026
  52. 52GAIA-2: A Controllable Multi-View Generative World Model for Autonomous DrivingPaper · 26 Mar 2025 · checked 10 Oct 2026
  53. 53GAIA-2: Pushing the Boundaries of Video Generative Models for Safer Assisted and Automated DrivingOfficial blog · 26 Mar 2025 · checked 11 Oct 2026
  54. 54DreamGen: Unlocking Generalization in Robot Learning through Video World ModelsPaper · 19 May 2025 · checked 10 Oct 2026
  55. 55NVIDIA/GR00T-Dreams (README)Repository · Jun 2025 · checked 11 Oct 2026
  56. 56Introducing Odyssey-1: A Playable World ModelOfficial blog · 28 May 2025 · checked 11 Oct 2026
  57. 57Fuel your creativity with new generative media models and toolsOfficial blog · 20 May 2025 · checked 11 Oct 2026
  58. 58Video models are zero-shot learners and reasonersPaper · 24 Sep 2025 · checked 10 Oct 2026
  59. 59V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and PlanningPaper · 11 Jun 2025 · checked 10 Oct 2026
  60. 60facebookresearch/vjepa2 (README)Repository · 25 Jun 2025 · checked 11 Oct 2026
  61. 61HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or PixelsPaper · 29 Jul 2025 · checked 10 Oct 2026
  62. 62tencent/HunyuanWorld-1Repository · Jul 2025 · checked 11 Oct 2026
  63. 63Genie 3: A new frontier for world modelsOfficial blog · 5 Aug 2025 · checked 11 Oct 2026
  64. 64Project Genie: AI world model now available for Ultra users in U.S.Official blog · 29 Jan 2026 · checked 11 Oct 2026
  65. 65Genie Envisioner: A Unified World Foundation Platform for Robotic ManipulationPaper · 7 Aug 2025 · checked 10 Oct 2026
  66. 66agibot-world/Genie-EnvisionerRepository · Aug 2025 · checked 11 Oct 2026
  67. 67Matrix-game 2.0: An open-source, real-time, and streaming interactive world modelPaper · 18 Aug 2025 · checked 10 Oct 2026
  68. 68Skywork/Matrix-Game-2.0Repository · Aug 2025 · checked 11 Oct 2026
  69. 69SkyworkAI/Matrix-Game (README news)Repository · 27 Mar 2026 · checked 11 Oct 2026
  70. 70Neural Robot DynamicsPaper · 21 Aug 2025 · checked 10 Oct 2026
  71. 71NVlabs/neural-robot-dynamics (README)Repository · Sep 2025 · checked 11 Oct 2026
  72. 72Training Agents Inside of Scalable World ModelsPaper · 29 Sep 2025 · checked 10 Oct 2026
  73. 73Generating Bigger and Better WorldsOfficial blog · 16 Sep 2025 · checked 11 Oct 2026
  74. 74Marble: A Multimodal World ModelOfficial blog · 12 Nov 2025 · checked 11 Oct 2026
  75. 75Sora 2 is hereOfficial blog · 30 Sep 2025 · checked 11 Oct 2026
  76. 76unitreerobotics/unifolm-world-model-action (README)Repository · 15 Sep 2025 · checked 11 Oct 2026
  77. 77unitreerobotics/UnifoLM-WMA-0 model repositoryRepository · Sep 2025 · checked 11 Oct 2026
  78. 78World Simulation with Video Foundation Models for Physical AIPaper · 28 Oct 2025 · checked 10 Oct 2026
  79. 79nvidia/Cosmos-Predict2.5-2BRepository · 2025 · checked 11 Oct 2026
  80. 80Ctrl-World: A Controllable Generative World Model for Robot ManipulationPaper · 11 Oct 2025 · checked 10 Oct 2026
  81. 81Robert-gyj/Ctrl-World (README)Repository · Oct 2025 · checked 11 Oct 2026
  82. 82GAIA-3: Scaling World Models to Power Safety and EvaluationOfficial blog · 2 Dec 2025 · checked 11 Oct 2026
  83. 83Introducing Runway GWM-1Official blog · 11 Dec 2025 · checked 11 Oct 2026
  84. 84WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World ModelingPaper · 16 Dec 2025 · checked 10 Oct 2026
  85. 85HY-World 1.5 Technical ReportPaper · 17 Dec 2025 · checked 11 Oct 2026
  86. 86Tencent-Hunyuan/HY-WorldPlay (README)Repository · 17 Dec 2025 · checked 11 Oct 2026
  87. 87Evaluating Gemini Robotics Policies in a Veo World SimulatorPaper · 11 Dec 2025 · checked 10 Oct 2026
  88. 881X World Model | From Video to Action: A New Way Robots LearnOfficial blog · 12 Jan 2026 · checked 11 Oct 2026
  89. 89Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and PlanningPaper · 22 Jan 2026 · checked 10 Oct 2026
  90. 90nvlabs/cosmos-policy (README)Repository · Dec 2025 · checked 11 Oct 2026
  91. 91Advancing Open-source World ModelsPaper · 28 Jan 2026 · checked 10 Oct 2026
  92. 92robbyant/lingbot-world (README)Repository · 29 Jan 2026 · checked 11 Oct 2026
  93. 93DreamDojo: A Generalist Robot World Model from Large-Scale Human VideosPaper · 6 Feb 2026 · checked 10 Oct 2026
  94. 94NVIDIA/DreamDojo (README)Repository · 18 Feb 2026 · checked 11 Oct 2026
  95. 95World Action Models are Zero-shot PoliciesPaper · 17 Feb 2026 · checked 10 Oct 2026
  96. 96dreamzero0/dreamzero (README)Repository · Feb 2026 · checked 11 Oct 2026
  97. 97The Waymo World Model: A New Frontier For Autonomous Driving SimulationOfficial blog · 6 Feb 2026 · checked 11 Oct 2026
  98. 98Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon MemoryPaper · 10 Apr 2026 · checked 10 Oct 2026
  99. 99Cosmos 3: Omnimodal World Models for Physical AIPaper · 1 Jun 2026 · checked 10 Oct 2026
  100. 100nvidia/Cosmos3-Super model cardRepository · 31 May 2026 · checked 11 Oct 2026
  101. 101GE-Sim 2.0: A Roadmap Towards Comprehensive Closed-loop Video World Simulators for Robotic ManipulationPaper · 26 May 2026 · checked 10 Oct 2026
  102. 102agibot-world/Genie-Envisioner-Sim-v2.0 model cardRepository · May 2026 · checked 11 Oct 2026
  103. 103Introducing Oasis 3: First Interactive World Model for Physical AIOfficial blog · 10 Jun 2026 · checked 11 Oct 2026
  104. 104GAIA-4: Multimodal World Models Powering Closed-Loop Simulation for Safe and Scalable AutonomyOfficial blog · 3 Aug 2026 · checked 11 Oct 2026
  105. 105Atlas: A World Model for Spatial IntelligenceOfficial blog · 1 Sep 2026 · checked 11 Oct 2026
  106. 106Introducing Odyssey-3: A General-Purpose Physical IntelligenceOfficial blog · 15 Sep 2026 · checked 11 Oct 2026
  107. 107A Comprehensive Survey on World Models for Embodied AIPaper · 19 Oct 2025 · checked 10 Oct 2026
  108. 108World Models for Robotic Manipulation: A SurveyPaper · 27 May 2026 · checked 10 Oct 2026
  109. 109Self-Supervised Learning from Images with a Joint-Embedding Predictive ArchitecturePaper · 19 Jan 2023 · checked 10 Oct 2026
  110. 110LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from PixelsPaper · 13 Mar 2026 · checked 10 Oct 2026
  111. 111Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationPaper · 20 Dec 2023 · checked 10 Oct 2026
  112. 112Unified Video Action ModelPaper · 28 Feb 2025 · checked 10 Oct 2026
  113. 113Agentic World Modeling: Foundations, Capabilities, Laws, and BeyondPaper · 24 Apr 2026 · checked 10 Oct 2026
  114. 114World Models and World-Action Models: An Accessible and Comprehensive SurveyPaper · 12 Aug 2026 · checked 10 Oct 2026
  115. 115Video generation models as world simulatorsOfficial blog · 15 Feb 2024 · checked 10 Oct 2026
  116. 116UniSim: A Neural Closed-Loop Sensor SimulatorPaper · Jun 2023 · checked 10 Oct 2026
  117. 117OccWorld: Learning a 3D Occupancy World Model for Autonomous DrivingPaper · 27 Nov 2023 · checked 10 Oct 2026
  118. 118WonderWorld: Interactive 3D Scene Generation from a Single ImagePaper · 13 Jun 2024 · checked 10 Oct 2026
  119. 119Genie 3 model pageOfficial site · checked 11 Oct 2026
  120. 120Project GenieOfficial site · checked 11 Oct 2026
  121. 121Veo 3 Model CardOfficial site · 23 May 2025 · checked 11 Oct 2026
  122. 122Introducing Veo 3.1 and advanced capabilities in FlowOfficial blog · 15 Oct 2025 · checked 11 Oct 2026
  123. 123Veo model pageOfficial site · checked 11 Oct 2026
  124. 124SIMA 2: An agent that plays, reasons, and learns with you in virtual 3D worldsOfficial blog · 13 Nov 2025 · checked 11 Oct 2026
  125. 125Sora 2 Model | OpenAI APIOfficial site · checked 11 Oct 2026
  126. 126Deprecations: 2026-03-24 Sora 2 video generation models and Videos APIOfficial site · 24 Mar 2026 · checked 11 Oct 2026
  127. 127OpenAI API changelog (Oct 6 entry: v1/videos with Sora 2)Official site · checked 11 Oct 2026
  128. 128Sora (text-to-video model)Secondary · checked 11 Oct 2026
  129. 129V-JEPA: The next step toward Yann LeCun's vision of advanced machine intelligence (AMI)Official blog · 15 Feb 2024 · checked 11 Oct 2026
  130. 130Introducing the V-JEPA 2 world model and new benchmarks for physical reasoningOfficial blog · 11 Jun 2025 · checked 11 Oct 2026
  131. 131Navigation World ModelsPaper · 4 Dec 2024 · checked 11 Oct 2026
  132. 132facebookresearch/nwmRepository · 12 Mar 2025 · checked 11 Oct 2026
  133. 133Introducing Muse Image and Muse Video (Meta Superintelligence Labs)Official blog · 7 Jul 2026 · checked 11 Oct 2026
  134. 134Microsoft Research License Terms (WHAM)Repository · checked 11 Oct 2026
  135. 135WHAMM! Real-time world modelling of interactive environmentsOfficial blog · 4 Apr 2025 · checked 11 Oct 2026
  136. 136Cosmos World Foundation Model Platform for Physical AI (HTML v3)Paper · 9 Jul 2025 · checked 11 Oct 2026
  137. 137nvidia-cosmos GitHub organisation (repository list)Repository · checked 11 Oct 2026
  138. 138Cosmos 3 (HTML v4)Paper · 23 Jun 2026 · checked 11 Oct 2026
  139. 139NVIDIA Cosmos product pageOfficial site · checked 11 Oct 2026
  140. 140Tesla Q2 2025 Update (8-K Exhibit 99.1)filing · 23 Jul 2025 · checked 11 Oct 2026
  141. 141Tesla Q3 2025 Update (8-K Exhibit 99.1)filing · 22 Oct 2025 · checked 11 Oct 2026
  142. 142Tesla Q4 and FY 2025 Update (8-K Exhibit 99.1)filing · 28 Jan 2026 · checked 11 Oct 2026
  143. 143Tesla Q1 2026 Update (8-K Exhibit 99.1)filing · 22 Apr 2026 · checked 11 Oct 2026
  144. 144Tesla Q2 2026 Update (8-K Exhibit 99.1)filing · 22 Jul 2026 · checked 11 Oct 2026
  145. 145Tesla Form 10-K for fiscal year 2025filing · 29 Jan 2026 · checked 11 Oct 2026
  146. 146Investing to bring the Waymo Driver to more ridersOfficial blog · 25 Oct 2024 · checked 11 Oct 2026
  147. 147Accelerating our global growth: Waymo raises $16 billion investment roundOfficial blog · 2 Feb 2026 · checked 11 Oct 2026
  148. 148Waymo Closes $5 Billion Debt Financing to Accelerate Business ExpansionOfficial blog · 8 Oct 2026 · checked 11 Oct 2026
  149. 149GAIA: generative world models for autonomy (Wayve Labs page)Official site · checked 11 Oct 2026
  150. 150Wayve releases GAIA-1 technical reportpress-release · 3 Oct 2023 · checked 11 Oct 2026
  151. 151Wayve unveils GAIA-2press-release · 26 Mar 2025 · checked 11 Oct 2026
  152. 152Wayve launches GAIA-3, advancing world models from simulation to evaluationpress-release · 2 Dec 2025 · checked 11 Oct 2026
  153. 153Wayve investors pageOfficial site · checked 11 Oct 2026
  154. 154Wayve announces $200 Million in Funding to Accelerate AV2.0press-release · 18 Jan 2022 · checked 11 Oct 2026
  155. 155Wayve Raises Over $1 Billion Led by SoftBank to Develop Embodied AI Products for Automated Drivingpress-release · 7 May 2024 · checked 11 Oct 2026
  156. 156Wayve secures $1.5B to deploy its global autonomy platformpress-release · 25 Feb 2026 · checked 11 Oct 2026
  157. 157Wayve Broadens Silicon Backing with $60M Investment from AMD, Arm and Qualcommpress-release · 15 Apr 2026 · checked 11 Oct 2026
  158. 158Waabi press and insights indexOfficial site · checked 11 Oct 2026
  159. 159Waabi WorldOfficial blog · 9 Feb 2022 · checked 11 Oct 2026
  160. 160Introducing UniSim, one of the core technologies powering Waabi WorldOfficial blog · 14 Jun 2023 · checked 11 Oct 2026
  161. 161Introducing Copilot4DOfficial blog · 15 Mar 2024 · checked 11 Oct 2026
  162. 162The ultimate driving test for AI: Mixed Reality Testing pushes the boundaries of AV safetyOfficial blog · 14 Jul 2025 · checked 11 Oct 2026
  163. 163Welcoming Volvo Group Venture Capital as a strategic investor in WaabiOfficial blog · 18 Jan 2023 · checked 11 Oct 2026
  164. 164Waabi raises $200M to launch fully driverless trucks in 2025press-release · 18 Jun 2024 · checked 11 Oct 2026
  165. 165Waabi secures $1 Billion in new funding to lead Physical AI revolutionpress-release · 28 Jan 2026 · checked 11 Oct 2026
  166. 1661X World Model: Sampling Challenge UpdateOfficial blog · 5 Nov 2024 · checked 11 Oct 2026
  167. 167The 1X World Model Lab | Est. 2026Official blog · 4 Jun 2026 · checked 11 Oct 2026
  168. 1681X Raises $23.5M in Series A2 Funding led by OpenAIpress-release · 23 Mar 2023 · checked 11 Oct 2026
  169. 169Series B: 1X Secures $100M Fundingpress-release · Jan 2024 · checked 11 Oct 2026
  170. 170Norwegian Robotics Startup 1X Secures $100M in Series B Funding Led by EQT VenturesOfficial blog · 8 Jun 2024 · checked 11 Oct 2026
  171. 171Physical Intelligence (π) – Blog (archived copy, 2026-10-03; live site returned HTTP 429)Official blog · Oct 2026 · checked 11 Oct 2026
  172. 172Jeff Bezos and OpenAI invest in robot startup Physical Intelligence at $2.4 billion valuation (CNBC)Secondary · 4 Nov 2024 · checked 11 Oct 2026
  173. 173Jeff Bezos-backed Physical Intelligence raises $600M to improve AI robot brainsSecondary · 20 Nov 2025 · checked 11 Oct 2026
  174. 174Skild AI blog indexOfficial blog · checked 11 Oct 2026
  175. 175The case for an omni-bodied robot brainOfficial blog · 24 Sep 2025 · checked 11 Oct 2026
  176. 176Introducing S1: In-Context Learning for RoboticsOfficial blog · checked 11 Oct 2026
  177. 177Announcing Series COfficial blog · 14 Jan 2026 · checked 11 Oct 2026
  178. 178Announcing our $300M Series A FundingOfficial blog · 9 Jul 2024 · checked 11 Oct 2026
  179. 179SoftBank, Nvidia looking to invest in Skild AI at $14 billion valuation, sources say (Reuters)Secondary · 8 Dec 2025 · checked 11 Oct 2026
  180. 180Figure news indexOfficial site · checked 11 Oct 2026
  181. 181Project Go-BigOfficial blog · 18 Sep 2025 · checked 11 Oct 2026
  182. 182Helix 02Official blog · 27 Jan 2026 · checked 11 Oct 2026
  183. 183Introducing Index: Building The World's Largest and Most Diverse Physical DatasetOfficial blog · 25 Aug 2026 · checked 11 Oct 2026
  184. 184Helix 2.5: zero-shot generalization in 30 homesOfficial blog · 17 Sep 2026 · checked 11 Oct 2026
  185. 185Figure Exceeds $1B in Series C Funding at $39B Post-Money ValuationOfficial blog · 16 Sep 2025 · checked 11 Oct 2026
  186. 186World Labs blog and news indexOfficial blog · checked 10 Oct 2026
  187. 187Generating WorldsOfficial blog · 2 Dec 2024 · checked 10 Oct 2026
  188. 188RTFM: A Real-Time Frame ModelOfficial blog · 16 Oct 2025 · checked 10 Oct 2026
  189. 189Announcing the World APIOfficial blog · 21 Jan 2026 · checked 10 Oct 2026
  190. 190About World LabsOfficial site · checked 10 Oct 2026
  191. 191World Labs is Joining AMDOfficial blog · 28 Sep 2026 · checked 10 Oct 2026
  192. 192Fei-Fei Li's World Labs comes out of stealth with $230M in fundingSecondary · 13 Sep 2024 · checked 10 Oct 2026
  193. 193World Labs Announces New FundingOfficial blog · 18 Feb 2026 · checked 10 Oct 2026
  194. 194Autodesk invests $200 million in World Labs, secures strategic advisor roleOfficial blog · 18 Feb 2026 · checked 10 Oct 2026
  195. 195AMI Labs: Real World. Real Intelligence.Official site · checked 10 Oct 2026
  196. 196AMI Labs - Updates: Official launchOfficial blog · 10 Mar 2026 · checked 10 Oct 2026
  197. 197Yann LeCun's AMI Labs raises $1.03B to build world modelsSecondary · 9 Mar 2026 · checked 10 Oct 2026
  198. 198General IntuitionOfficial site · checked 10 Oct 2026
  199. 199MIRA: a playable multiplayer world modelOfficial blog · checked 10 Oct 2026
  200. 200mira-wm/mira: Code for MIRA: Multiplayer Interactive World Models with Representation AutoencodersRepository · 5 Jul 2026 · checked 10 Oct 2026
  201. 201General Intuition's $2.3B bet that video games can train AI agents for the real worldSecondary · 25 Jun 2026 · checked 10 Oct 2026
  202. 202World model startup General Intuition closes $220M investmentSecondary · 29 Sep 2026 · checked 10 Oct 2026
  203. 203Decart AI Lab | Resources (publications and news index)Official blog · checked 10 Oct 2026
  204. 204Oasis: A Universe in a TransformerOfficial blog · 31 Oct 2024 · checked 10 Oct 2026
  205. 205Introducing Lucy 2: SOTA realtime world transformation modelOfficial blog · 26 Jan 2026 · checked 10 Oct 2026
  206. 206Partnering with Decart: The Future of AI-Generated ExperiencesOfficial blog · 31 Oct 2024 · checked 10 Oct 2026
  207. 207Decart adds another $32M at a $500M valuationSecondary · 19 Dec 2024 · checked 10 Oct 2026
  208. 208Exclusive: Decart raises $100 million at a $3.1 billion valuationSecondary · 7 Aug 2025 · checked 10 Oct 2026
  209. 209Decart Raises $300M: Tech Leaders Back the Company as Both Customers and InvestorsOfficial blog · 18 May 2026 · checked 10 Oct 2026
  210. 210Decart: $300 Million Raised At Nearly $4 Billion ValuationSecondary · 19 May 2026 · checked 10 Oct 2026
  211. 211The latest from Odyssey (post index)Official blog · checked 10 Oct 2026
  212. 212World Models for Film, Gaming, and BeyondOfficial blog · 18 Dec 2024 · checked 10 Oct 2026
  213. 213Introducing Odyssey-2: A General-Purpose World ModelOfficial blog · 27 Oct 2025 · checked 10 Oct 2026
  214. 214Introducing Odyssey-2 Max: Scaled World SimulationOfficial blog · 21 Apr 2026 · checked 10 Oct 2026
  215. 215Starchild-1: The First Real-Time Multimodal World ModelOfficial blog · 17 May 2026 · checked 10 Oct 2026
  216. 216Introducing Agora-2: Advancing Multi-Agent World SimulationOfficial blog · 21 Sep 2026 · checked 10 Oct 2026
  217. 217Meet Odyssey-3: Our Most Powerful Foundation World ModelOfficial blog · 8 Oct 2026 · checked 10 Oct 2026
  218. 218Odyssey Announces Investment from NVentures and Samsung NextOfficial blog · 12 Feb 2026 · checked 10 Oct 2026
  219. 219Our $310 Million Fundraise to Accelerate World SimulationOfficial blog · 17 Jun 2026 · checked 10 Oct 2026
  220. 220World model maker Odyssey nabs $1.45B valuation backed by Amazon and other big namesSecondary · 17 Jun 2026 · checked 10 Oct 2026
  221. 221Runway ResearchOfficial site · checked 11 Oct 2026
  222. 222Introducing General World ModelsOfficial blog · 11 Dec 2023 · checked 11 Oct 2026
  223. 223Runway Gen-4.5: State-of-the-Art AI Video GenerationOfficial blog · 1 Dec 2025 · checked 11 Oct 2026
  224. 224Introducing GWM-1Official blog · 11 Dec 2025 · checked 11 Oct 2026
  225. 225Introducing SolarisOfficial blog · 31 Aug 2026 · checked 11 Oct 2026
  226. 226Introducing GWM Worlds 2Official blog · 3 Sep 2026 · checked 11 Oct 2026
  227. 227Introducing Praxis-1Official blog · Sep 2026 · checked 11 Oct 2026
  228. 228Towards a new media ecosystem with world simulatorsOfficial blog · 3 Apr 2025 · checked 11 Oct 2026
  229. 229Crunchbase News on Runway's Series ESecondary · 10 Feb 2026 · checked 11 Oct 2026
  230. 230New Funding to Scale World SimulationOfficial blog · 10 Feb 2026 · checked 11 Oct 2026
  231. 231Luma AI launches Ray3press-release · 18 Sep 2025 · checked 11 Oct 2026
  232. 232Luma Introduces Ray3.2 Model & APIOfficial blog · 9 Jun 2026 · checked 11 Oct 2026
  233. 233Luma is Partnering with HUMAIN to Accelerate the Arrival of Multimodal AGIOfficial blog · 15 May 2025 · checked 11 Oct 2026
  234. 234AGI is multimodal and reality is the dataset of AGIOfficial blog · 19 Nov 2025 · checked 11 Oct 2026
  235. 235Saudi-Based Humain Leads $900M Series C Round for Luma AISecondary · 20 Nov 2025 · checked 11 Oct 2026
  236. 236Mirage: AI UGC game engine (archived copy of the official blog, 2025-07-02)Official blog · Jul 2025 · checked 10 Oct 2026
  237. 237Magica: AI UGC game engine (archived copy of the official blog, 2025-12-02)Official blog · Dec 2025 · checked 10 Oct 2026
  238. 238Mirage 2 allows users to turn sketches and photos into interactive game worldsSecondary · 22 Aug 2025 · checked 10 Oct 2026
  239. 239Tencent-Hunyuan/HunyuanWorld-1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or PixelsRepository · 26 Jul 2025 · checked 11 Oct 2026
  240. 240Tencent-Hunyuan/Hunyuan-GameCraft-1.0: High-dynamic Interactive Game Video GenerationRepository · 14 Aug 2025 · checked 11 Oct 2026
  241. 241Tencent-Hunyuan/HunyuanWorld-Voyager: Interactive RGBD video generation conditioned on camera inputRepository · 2 Sep 2025 · checked 11 Oct 2026
  242. 242Tencent-Hunyuan/HunyuanWorld-Mirror: WorldMirror: universal 3D reconstructionRepository · 22 Oct 2025 · checked 11 Oct 2026
  243. 243Tencent-Hunyuan/HY-World-2.0: HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D WorldsRepository · 16 Apr 2026 · checked 11 Oct 2026
  244. 244Wan-Video/Wan2.1: Wan: Open and Advanced Large-Scale Video Generative ModelsRepository · 25 Feb 2025 · checked 11 Oct 2026
  245. 245Wan-Video/Wan2.2: Wan: Open and Advanced Large-Scale Video Generative ModelsRepository · 28 Jul 2025 · checked 11 Oct 2026
  246. 246alibaba-damo-academy/RynnVLA-002: RynnVLA-002 (formerly WorldVLA): A Unified Vision-Language-Action and World ModelRepository · 23 Jun 2025 · checked 11 Oct 2026
  247. 247alibaba-damo-academy/RynnWorld-4D: RynnWorld-4D: 4D Embodied World Models for Robotic ManipulationRepository · 7 Jul 2026 · checked 11 Oct 2026
  248. 248alibaba-damo-academy/RynnWorld-Teleop: RynnWorld-Teleop: An Action-Conditioned World Model for Digital TeleoperationRepository · 2 Jul 2026 · checked 11 Oct 2026
  249. 249ByteDance-Seed/VideoWorld: VideoWorld: learning world models from unlabeled videos (CVPR 2025)Repository · 15 Jan 2025 · checked 11 Oct 2026
  250. 250Seedance (ByteDance Seed model page)Official site · checked 11 Oct 2026
  251. 251Kling AI Release Notes (content did not load)Official site · checked 11 Oct 2026
  252. 252Robbyant/lingbot-world: LingBot-World: Advancing Open-source World ModelsRepository · 29 Jan 2026 · checked 11 Oct 2026
  253. 253Robbyant/lingbot-va: LingBot-VA: Causal World Modeling for Robot ControlRepository · 29 Jan 2026 · checked 11 Oct 2026
  254. 254Robbyant/lingbot-world-v2: LingBot-World-Infinity: Infinite Worlds with Versatile InteractionsRepository · 9 Jul 2026 · checked 11 Oct 2026
  255. 255SenseTime-FVG/OpenDWM: Open Driving World Models (OpenDWM)Repository · 15 Jan 2025 · checked 11 Oct 2026
  256. 256InternRobotics/Aether: Aether: Geometric-Aware Unified World ModelingRepository · 28 Mar 2025 · checked 11 Oct 2026
  257. 257InternRobotics/InternW0-Delta: InternW0-Δ: A World Action Model with 20K+ Hours of Open DataRepository · 24 Sep 2026 · checked 11 Oct 2026
  258. 258OpenDriveLab/Vista: Vista: A Generalizable Driving World ModelRepository · 28 May 2024 · checked 11 Oct 2026
  259. 259baaivision/Emu3: Emu3: Next-Token Prediction is All You NeedRepository · 26 Sep 2024 · checked 11 Oct 2026
  260. 260baaivision/Emu3.5: Emu3.5: Native Multimodal Models are World LearnersRepository · 29 Oct 2025 · checked 11 Oct 2026
  261. 261SkyworkAI/Matrix-3D: Matrix-3D: explorable 3D scenes from panorama videosRepository · 11 Aug 2025 · checked 11 Oct 2026
  262. 262manycore-research/SpatialGen: SpatialGen: Layout-guided 3D Indoor Scene GenerationRepository · 24 Aug 2025 · checked 11 Oct 2026
  263. 263AgibotTech/EnerVerse-AC: EnerVerse-AC: Envisioning Embodied Environments with Action ConditionRepository · 14 May 2025 · checked 11 Oct 2026
  264. 264AgibotTech/Genie-Envisioner-V1: Genie Envisioner: A Unified World Foundation Platform for Robotic ManipulationRepository · 8 Aug 2025 · checked 11 Oct 2026
  265. 265AgibotTech/GE-Sim-V2: GE-Sim 2.0: closed-loop video world simulator for robotic manipulationRepository · 28 May 2026 · checked 11 Oct 2026
  266. 266AgiBot news page (动态速递)Official site · checked 11 Oct 2026
  267. 267Galbot news page (lists 'WAM-TTT' item dated 2026-07-16)Official site · checked 11 Oct 2026
  268. 268unitreerobotics/unifolm-wla: UnifoLM-WLA-1.0Repository · 28 Sep 2026 · checked 11 Oct 2026
  269. 269open-gigaai/giga-world-0: GigaWorld-0: World Models as Data Engine to Empower Embodied AIRepository · 25 Nov 2025 · checked 11 Oct 2026
  270. 270open-gigaai/giga-world-policy: GigaWorld-Policy: An Efficient Action-Centered World–Action ModelRepository · 3 Mar 2026 · checked 11 Oct 2026
  271. 271open-gigaai/giga-world-1: GigaWorld-1: A Roadmap to World Models for Robot Policy EvaluationRepository · Jul 2026 · checked 11 Oct 2026
  272. 272thu-ml/vidar: Vidar and Vidarc: video foundation model for roboticsRepository · Jul 2025 · checked 11 Oct 2026
  273. 273shengshu-ai/Motubrain: Motubrain: An Advanced World Action Model for Robot ControlRepository · Apr 2026 · checked 11 Oct 2026
  274. 274shengshu-ai/Motus2: Motus2: A Self-Evolving General World Model for Dexterous ManipulationRepository · Sep 2026 · checked 11 Oct 2026
  275. 275shengshu-ai/Vidu-S: Vidu S: Real-Time Interactive, Editable, and Spatial Video GenerationRepository · Jul 2026 · checked 11 Oct 2026
  276. 276NVIDIA Announces Major Release of Cosmos World Foundation Models and Physical AI Data Toolspress-release · 18 Mar 2025 · checked 10 Oct 2026
  277. 277NVIDIA Launches Cosmos World Foundation Model Platform to Accelerate Physical AI Developmentpress-release · 6 Jan 2025 · checked 10 Oct 2026
  278. 278How Cosmos 3 Helps Physical AI Think Before It ActsOfficial blog · 31 May 2026 · checked 11 Oct 2026
  279. 279NVIDIA and Global Robotics Leaders Take Physical AI to the Real Worldpress-release · 16 Mar 2026 · checked 10 Oct 2026
  280. 280NVIDIA Opens Portals to World of Robotics With New Omniverse Libraries, Cosmos Physical AI Models and AI Computing Infrastructurepress-release · 11 Aug 2025 · checked 10 Oct 2026
  281. 281NVIDIA Launches Cosmos 3, the Open Frontier Foundation Model for Physical AIpress-release · 31 May 2026 · checked 10 Oct 2026
  282. 282Runway Robotics: Build smarter robots with general world modelsOfficial site · undated · checked 10 Oct 2026
  283. 283NVIDIA Powers Humanoid Robot Industry With Cloud-to-Robot Computing Platforms for Physical AIpress-release · 18 May 2025 · checked 10 Oct 2026
  284. 284GR00T N1: An Open Foundation Model for Generalist Humanoid RobotsPaper · 18 Mar 2025 · checked 10 Oct 2026
  285. 285GR00T N1 (HTML v2, sections 2.2, 3.2 and 4.4)Paper · 27 Mar 2025 · checked 10 Oct 2026
  286. 286GigaWorld-0: World Models as Data Engine to Empower Embodied AIPaper · 25 Nov 2025 · checked 10 Oct 2026
  287. 287GE-Sim2 | Genie Envisioner World Simulator 2.0Official site · Apr 2026 · checked 10 Oct 2026
  288. 288Hugging Face Hub API: NVIDIA Cosmos models with download countsindex · 10 Oct 2026 · checked 10 Oct 2026
  289. 289Japan's Robotics and Manufacturing Leaders Build on NVIDIA Cosmos to Advance Physical AI Frontierpress-release · 15 Jul 2026 · checked 10 Oct 2026
  290. 290Accelerating Robot Policy Evaluation with General World ModelsOfficial blog · 27 Feb 2026 · checked 10 Oct 2026
  291. 291NVIDIA Releases New Physical AI Models as Global Partners Unveil Next-Generation Robotspress-release · 5 Jan 2026 · checked 10 Oct 2026
  292. 292Evaluating Gemini Robotics Policies in a Veo World Simulator (HTML v2, section 4)Paper · 6 Jan 2026 · checked 10 Oct 2026
  293. 293AGIBOT Unveils Genie Envisioner 2.0, Advancing World Models into Scalable World Simulators for Embodied AIpress-release · 10 Apr 2026 · checked 10 Oct 2026
  294. 294AGIBOT's Genie Envisioner-Sim 2.0 Ranks No. 1 on WorldArena Benchmarkpress-release · 29 May 2026 · checked 10 Oct 2026
  295. 295GigaWorld-1: A Roadmap to Build World Models for Robot Policy EvaluationPaper · 2 Jul 2026 · checked 10 Oct 2026
  296. 296Ctrl-World project pageOfficial site · Oct 2025 · checked 11 Oct 2026
  297. 297WorldGym: World Model as An Environment for Policy EvaluationPaper · 31 May 2025 · checked 10 Oct 2026
  298. 298WorldEval: World Model as Real-World Robot Policies EvaluatorPaper · 25 May 2025 · checked 10 Oct 2026
  299. 299WorldEval project pageOfficial site · May 2025 · checked 11 Oct 2026
  300. 300NEO Factory | Building Your NEOOfficial blog · 30 Apr 2026 · checked 10 Oct 2026
  301. 301GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot ManipulationPaper · 8 Oct 2024 · checked 10 Oct 2026
  302. 302Order NEOOfficial site · undated · checked 10 Oct 2026
  303. 303Foretellix Expands Data Automation Toolchain for AI-Powered AV Development with Breakthrough Simulation Capabilities Using NVIDIA Omniverse and Cosmos TransferOfficial blog · 18 Mar 2025 · checked 11 Oct 2026
  304. 304Uber to Deploy One of the World's Largest Networks of Autonomous Vehicles, Powered by NVIDIA AI Architecturepress-release · 28 Oct 2025 · checked 11 Oct 2026
  305. 305Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation ModelsPaper · 10 Jun 2025 · checked 11 Oct 2026
  306. 306Introducing Oasis 3: First Interactive World Model for Physical AIOfficial blog · 10 Jun 2026 · checked 11 Oct 2026
  307. 307Pricing - Decart API Platform DocumentationOfficial site · undated · checked 11 Oct 2026
  308. 308Helm.ai Sets New Full HD (2MP) Standard for Generative Simulationpress-release · 27 May 2026 · checked 11 Oct 2026
  309. 309The Waymo World Model: A New Frontier For Autonomous Driving SimulationOfficial blog · 6 Feb 2026 · checked 11 Oct 2026
  310. 310XPENG Releases World Model Technical Report, Powering VLA 2.0 Model R&D and Verificationpress-release · 29 Apr 2026 · checked 11 Oct 2026
  311. 311X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End DrivingPaper · 20 Mar 2026 · checked 11 Oct 2026
  312. 312Li Auto Inc. Annual Results Announcement for the Year Ended December 31, 2025 (Form 6-K exhibit 99.2)press-release · Mar 2026 · checked 11 Oct 2026
  313. 3132026 华为乾崑技术大会在京举行 (2026 Huawei Qiankun Technology Conference)press-release · 23 Apr 2026 · checked 11 Oct 2026
  314. 314乾崑智驾 ADS 5 product pageOfficial site · undated · checked 11 Oct 2026
  315. 315Pony.ai Launches PonyWorld 2.0, a Self-Improving Physical AI Engine for Autonomous Drivingpress-release · 10 Apr 2026 · checked 11 Oct 2026
  316. 316openpilot 0.11Official blog · 17 Mar 2026 · checked 11 Oct 2026
  317. 317Learning to Drive from a World ModelPaper · 27 Apr 2025 · checked 11 Oct 2026
  318. 318Tesla AI Chief Details Unified 'World Simulator' for FSD and Optimus (Humanoids Daily)Secondary · 24 Oct 2025 · checked 11 Oct 2026
  319. 319XPENG Unveils the World Model Accelerator X-Cache, Which Requires No Training, Is Plug-and-Play, and Boosts Inference Speed by 2.7 Timespress-release · 6 May 2026 · checked 11 Oct 2026
  320. 320NIO Inc. Provides January 2026 Delivery Update (Form 6-K exhibit 99.1)press-release · 1 Feb 2026 · checked 11 Oct 2026
  321. 321X-Mind: Empowering Autonomous Driving with a Future-Foresight Brainpress-release · 29 Jun 2026 · checked 11 Oct 2026
  322. 322Google AI plans and subscriptionsOfficial site · undated · checked 11 Oct 2026
  323. 323Agora-1: The Multi-Agent World ModelOfficial blog · 18 May 2026 · checked 11 Oct 2026
  324. 324Oasis: A Universe in a Transformer (Decart)Official blog · 31 Oct 2024 · checked 11 Oct 2026
  325. 325Multiplayer Interactive World Models with Representation AutoencodersPaper · Jul 2026 · checked 11 Oct 2026
  326. 326Empowering Creators and Players With Muse, a Generative AI Model for GameplayOfficial blog · 19 Feb 2025 · checked 11 Oct 2026
  327. 327Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History ConditionPaper · 20 Jun 2025 · checked 11 Oct 2026
  328. 328GameNGen project pageOfficial site · Aug 2024 · checked 11 Oct 2026
  329. 329Decart Raises $300M: Tech Leaders Back the Company as Both Customers and Investorspress-release · 18 May 2026 · checked 11 Oct 2026
  330. 330Runway Partners with Lionsgatepress-release · 18 Sep 2024 · checked 11 Oct 2026
  331. 331Runway and Lionsgate Expand Partnershippress-release · 11 Jun 2026 · checked 11 Oct 2026
  332. 332Runway Partners with AMC Networks Across Marketing and TV Developmentpress-release · 4 Jun 2025 · checked 11 Oct 2026
  333. 333How House of David Used Runway to Become Amazon's Latest Hit SeriesOfficial site · undated · checked 11 Oct 2026
  334. 334The Walt Disney Company and OpenAI Reach Landmark Agreement to Bring Beloved Characters From Across Disney's Brands to Sorapress-release · 11 Dec 2025 · checked 11 Oct 2026
  335. 335From Image to Immersive: VIVE Mars x MarbleOfficial site · 12 Nov 2025 · checked 11 Oct 2026
  336. 336Framing Worlds: How Marble Helps Creators Bring Consistency to AI FilmmakingOfficial site · 12 Nov 2025 · checked 11 Oct 2026
  337. 337Generating the World Layer: How Lightwheel and World Labs Scale Robotics EvaluationOfficial site · 12 Nov 2025 · checked 11 Oct 2026
  338. 338Reframing Space: How Marble Is Transforming Architectural and Interior VisualizationOfficial site · 12 Nov 2025 · checked 11 Oct 2026
  339. 339WorldGen: From Text to Traversable and Interactive 3D WorldsPaper · 20 Nov 2025 · checked 11 Oct 2026
  340. 340NVIDIA Omniverse Physical AI Operating System Expands to More Industries and Partnerspress-release · 18 Mar 2025 · checked 11 Oct 2026
  341. 341Building Worlds Together: How VIVERSE and Marble Empowered Creators to Build Interactive 3D WorldsOfficial site · 12 Nov 2025 · checked 11 Oct 2026
  342. 342Immersive Exposure: Generative Worlds for OCD TherapyOfficial site · 12 Nov 2025 · checked 11 Oct 2026
  343. 343Splat World (World Labs showcase)Official site · undated · checked 11 Oct 2026
  344. 344Mastering diverse control tasks through world models (Nature)Paper · 2 Apr 2025 · checked 11 Oct 2026
  345. 345DIAMOND project pageOfficial site · May 2024 · checked 11 Oct 2026
  346. 346nvidia/Cosmos-H-Surgical-Simulator (Hugging Face model card)Repository · Feb 2026 · checked 11 Oct 2026
  347. 347nvidia/Cosmos-H-Dreams (Hugging Face model card)Repository · Jul 2026 · checked 11 Oct 2026
  348. 348MuZero's first step from research into the real worldOfficial blog · 11 Feb 2022 · checked 11 Oct 2026
  349. 349MuZero with Self-competition for Rate Control in VP9 Video CompressionPaper · 14 Feb 2022 · checked 11 Oct 2026
  350. 350Safety-first AI for autonomous data centre cooling and industrial controlOfficial blog · 17 Aug 2018 · checked 11 Oct 2026
  351. 351DeepMind AI Reduces Google Data Centre Cooling Bill by 40%Official blog · 20 Jul 2016 · checked 11 Oct 2026
  352. 3521x-technologies/1xgptRepository · checked 10 Oct 2026
  353. 3531xgpt README (challenge rules)Repository · checked 10 Oct 2026
  354. 354EVA: An Embodied World Model for Future Video Anticipation (arXiv abs; v1 2024-10-20, v2 2025-06-10)Paper · 20 Oct 2024 · checked 10 Oct 2026
  355. 355EVA full text (arXiv HTML v2), Section 5 and Appendix C on EVA-BenchPaper · 10 Jun 2025 · checked 11 Oct 2026
  356. 356EVA demo site (labelled 'ICML2025 Submission')project-page · checked 11 Oct 2026
  357. 357WorldSimBench: Towards Video Generation Models as World Simulators (arXiv abs, 13 authors)Paper · 23 Oct 2024 · checked 10 Oct 2026
  358. 358WorldSimBench full text (arXiv HTML v1)Paper · 23 Oct 2024 · checked 11 Oct 2026
  359. 359WorldSimBench, Proceedings of ICML 2025, PMLR 267:50338-50362 (12 authors)Paper · Jul 2025 · checked 11 Oct 2026
  360. 360WorldSimBench PMLR PDF (affiliation footnote, Table 3)Paper · Jul 2025 · checked 11 Oct 2026
  361. 361WorldSimBench project pageproject-page · checked 11 Oct 2026
  362. 362GitHub repository search for WorldSimBench (only the homepage repo IranQin/WorldSimBench.github.io, no licence)index · checked 11 Oct 2026
  363. 363WorldModelBench: Judging Video Generation Models As World Models (arXiv abstract page)Paper · 28 Feb 2025 · checked 10 Oct 2026
  364. 364WorldModelBench (arXiv HTML v1)Paper · 28 Feb 2025 · checked 10 Oct 2026
  365. 365WorldModelBench project pageproject-page · checked 10 Oct 2026
  366. 366WorldModelBench leaderboard data fileLeaderboard · checked 10 Oct 2026
  367. 367WorldModelBench repository READMERepository · 29 Jul 2025 · checked 10 Oct 2026
  368. 368WorldModelBench on OpenReview (NeurIPS 2025 Datasets and Benchmarks Track poster; read via api2.openreview.net search)Paper · checked 10 Oct 2026
  369. 369worldmodelbench dataset metadatadataset-card · 20 Dec 2024 · checked 10 Oct 2026
  370. 370EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World ModelsPaper · 14 May 2025 · checked 10 Oct 2026
  371. 371AgiBot World Challenge ICRA 2026 World Model trackLeaderboard · checked 10 Oct 2026
  372. 372EnerVerse-AC: Envisioning Embodied Environments with Action Condition (arXiv abstract, v1)Paper · 14 May 2025 · checked 10 Oct 2026
  373. 373EnerVerse-AC full text (arXiv HTML v1), Sec. 3 'Evaluator for Policy Model', Sec. 4.3, Appendix A.4.1Paper · 14 May 2025 · checked 10 Oct 2026
  374. 374EnerVerse-AC Fig. 7: success rate per task and per learning step, real robot vs EVACPaper · 14 May 2025 · checked 11 Oct 2026
  375. 375Hugging Face API: agibot-world/EnerVerse-AC model (licence cc-by-nc-sa-4.0)dataset-card · 15 May 2025 · checked 11 Oct 2026
  376. 376EnerVerse-AC project pageproject-page · checked 11 Oct 2026
  377. 377DreamGen full text (arXiv HTML v2)Paper · 17 Jun 2025 · checked 10 Oct 2026
  378. 378DreamGen project page (NVIDIA GEAR)project-page · 20 May 2025 · checked 10 Oct 2026
  379. 379NVIDIA/GR00T-Dreams README, section 5 DreamGen BenchRepository · checked 10 Oct 2026
  380. 380GitHub licence record for NVIDIA/GR00T-Dreams (Apache License 2.0)Repository · checked 10 Oct 2026
  381. 381Cosmos-Predict2 doc: Video2World post-training for DreamGen Bench (dataset links)Repository · checked 11 Oct 2026
  382. 382HF dataset card nvidia/EVAL-175 (resolves to nvidia/PhysicalAI-Robotics-GR00T-Eval; CC-BY-4.0)dataset-card · 14 Jun 2025 · checked 11 Oct 2026
  383. 383WorldEval full text (arXiv HTML v1): Sections 3-4, Appendices A-CPaper · 25 May 2025 · checked 10 Oct 2026
  384. 384WorldEval repository (LICENSE, README, apache-original/LICENSE)Repository · 19 May 2025 · checked 10 Oct 2026
  385. 385WorldGym full text (arXiv HTML v3): Sections 3-4, Appendices B and EPaper · 30 Sep 2025 · checked 10 Oct 2026
  386. 386Evaluating Robot Policies in a World Model (WorldGym arXiv v1 full text)Paper · 31 May 2025 · checked 11 Oct 2026
  387. 387WorldGym project page (abstract page)project-page · checked 10 Oct 2026
  388. 388WorldGym repository (README, root file listing)Repository · 9 Jun 2025 · checked 10 Oct 2026
  389. 389WorldGym on OpenReview, listed as ICLR 2026 Poster (checked through the api2.openreview.net notes search)Paper · checked 10 Oct 2026
  390. 390V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and PlanningPaper · 11 Jun 2025 · checked 10 Oct 2026
  391. 3911X World Model: Evaluating Bits, not Atoms (technical progress report)Report · 16 Jun 2025 · checked 10 Oct 2026
  392. 392Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation (arXiv abs; comment: ACL 2025 (Findings))Paper · 27 Jun 2025 · checked 11 Oct 2026
  393. 393WM-ABench full text (arXiv HTML v1)Paper · 27 Jun 2025 · checked 11 Oct 2026
  394. 394WM-ABench project page with per-category leaderboardproject-page · checked 11 Oct 2026
  395. 395HF dataset card maitrix-org/WM-ABench (license: apache-2.0)dataset-card · 29 Aug 2025 · checked 11 Oct 2026
  396. 396IRASim: A Fine-Grained World Model for Robot Manipulation (arXiv abstract; v1 2024-06-20, v2 2025-07-29)Paper · 29 Jul 2025 · checked 10 Oct 2026
  397. 397IRASim full text (arXiv HTML v2), Sec. 4.2 and Table 4Paper · 29 Jul 2025 · checked 10 Oct 2026
  398. 398IRASim arXiv HTML v1 (no policy-evaluation section, no Pearson value)Paper · 20 Jun 2024 · checked 11 Oct 2026
  399. 399bytedance/IRASim repository (GitHub API licence: Apache-2.0)Repository · 19 Jun 2024 · checked 11 Oct 2026
  400. 400PAI-Bench: A Comprehensive Benchmark For Physical AIPaper · 1 Dec 2025 · checked 10 Oct 2026
  401. 401Ctrl-World full text (arXiv HTML v3): Section 5.3, Appendix B Table 3Paper · 1 Mar 2026 · checked 10 Oct 2026
  402. 402Ctrl-World PDF v3 (Figure 7 on page 8; ICLR 2026 header)Paper · 1 Mar 2026 · checked 10 Oct 2026
  403. 403Ctrl-World model card metadata (license: mit)Repository · 3 Nov 2025 · checked 10 Oct 2026
  404. 404Stable Video Diffusion img2vid model card metadata (license_name: stable-video-diffusion-community)Repository · checked 10 Oct 2026
  405. 405Ctrl-World on OpenReview, listed as ICLR 2026 Poster (checked through the api2.openreview.net notes search)Paper · checked 11 Oct 2026
  406. 406Cosmos-Surg-dVRK: World Foundation Model-based Automated Online Evaluation of Surgical Robot Policy Learning (arXiv abs; v1 2025-10-17, v2 2025-11-03)Paper · 3 Nov 2025 · checked 10 Oct 2026
  407. 407Cosmos-Surg-dVRK full text (arXiv HTML v2): Sections 3-6, Tables 1-3Paper · 3 Nov 2025 · checked 10 Oct 2026
  408. 408Crossref record: IEEE Robotics and Automation Letters 11(5): 5978-5985index · May 2026 · checked 10 Oct 2026
  409. 409World-in-World: World Models in a Closed-Loop World (arXiv abs; v1 2025-10-20, v2 2026-08-16)Paper · 20 Oct 2025 · checked 10 Oct 2026
  410. 410World-in-World full text (arXiv HTML v2)Paper · 16 Aug 2026 · checked 10 Oct 2026
  411. 411World-in-World project pageproject-page · checked 10 Oct 2026
  412. 412World-In-World LeaderboardLeaderboard · checked 10 Oct 2026
  413. 413World-in-World GitHub READMERepository · checked 10 Oct 2026
  414. 414GitHub licence record for World-In-World/world-in-world (MIT License, LICENSE)Repository · checked 10 Oct 2026
  415. 415HF dataset zonszer/WIW_datasets metadata (no license field)dataset-card · 23 Oct 2025 · checked 10 Oct 2026
  416. 416Scalable Policy Evaluation with Video World Models (arXiv abstract; v1 2025-11-14, v3 2025-12-04)Paper · 14 Nov 2025 · checked 11 Oct 2026
  417. 417Scalable Policy Evaluation with Video World Models full text (arXiv HTML v3), Table I, Sec. IV-C, IV-DPaper · 4 Dec 2025 · checked 11 Oct 2026
  418. 418Tseng et al. Fig. 6: policy evaluation on the Bridge setup (Cosmos and IRASim, MMRV and Pearson)Paper · 4 Dec 2025 · checked 11 Oct 2026
  419. 419Veo world simulator report full text (arXiv HTML v2): Sections 2-5 and 7Report · 6 Jan 2026 · checked 10 Oct 2026
  420. 420Veo Robotics project pageproject-page · checked 10 Oct 2026
  421. 421PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies (arXiv abs; v1 2025-12-18, v2 2025-12-30)Paper · 30 Dec 2025 · checked 10 Oct 2026
  422. 422PolaRiS full text (arXiv HTML v2): Sections 3, 5, Appendices C-DPaper · 30 Dec 2025 · checked 10 Oct 2026
  423. 423PolaRiS project pageproject-page · checked 10 Oct 2026
  424. 424PolaRiS repository (LICENSE)Repository · 24 Nov 2025 · checked 10 Oct 2026
  425. 425PolaRiS-Hub dataset card metadata (license: mit)dataset-card · 14 Mar 2026 · checked 10 Oct 2026
  426. 426Rethinking Video Generation Model for the Embodied World (RBench)Paper · 21 Jan 2026 · checked 10 Oct 2026
  427. 427Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Test (arXiv abs)Paper · 7 Jan 2026 · checked 10 Oct 2026
  428. 428WoW-World-Eval full text (arXiv HTML v1)Paper · 7 Jan 2026 · checked 11 Oct 2026
  429. 429WoW-World-Eval Figure 3b: overall score on metric vs overall score on human (r = 0.93, rho = 0.91)Paper · 7 Jan 2026 · checked 11 Oct 2026
  430. 430DreamDojo full text (arXiv HTML v1), Sec. 4.7 'Downstream Applications' and LimitationsPaper · 6 Feb 2026 · checked 10 Oct 2026
  431. 431DreamDojo Fig. 5(a): real vs DreamDojo success rates (points A-F)Paper · 6 Feb 2026 · checked 10 Oct 2026
  432. 432DreamDojo project pageproject-page · checked 11 Oct 2026
  433. 433WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models (arXiv abs, v1 2026-02-09, v2 2026-02-11)Paper · 9 Feb 2026 · checked 10 Oct 2026
  434. 434WorldArena full text (arXiv HTML v2)Paper · 11 Feb 2026 · checked 10 Oct 2026
  435. 435WorldArena Figure 1a: EWMScore of 14 modelsPaper · 11 Feb 2026 · checked 10 Oct 2026
  436. 436WorldArena project pageproject-page · checked 10 Oct 2026
  437. 437WorldArena GitHub README (tsinghua-fib-lab/WorldArena)Repository · checked 10 Oct 2026
  438. 438WorldArena evaluation and submission guidelineRepository · checked 10 Oct 2026
  439. 439GitHub API record for tsinghua-fib-lab/WorldArena (license: null; 269 stars; created 2026-02-11)Repository · checked 10 Oct 2026
  440. 440WorldArena embodied_task/LICENSE (MIT, Copyright 2025 Tsinghua University)Repository · checked 10 Oct 2026
  441. 441WorldArena embodied_task/worldarena_track2/LICENSE (Apache-2.0, Copyright 2026 WorldArena Team)Repository · checked 10 Oct 2026
  442. 442HF dataset card WorldArena/WorldArena_Robotwin2.0 (license: apache-2.0)dataset-card · 18 Jul 2026 · checked 10 Oct 2026
  443. 443HF Space WorldArena/WorldArena metadata (official leaderboard; card license: mit; last modified 2026-07-15)Leaderboard · 15 Jul 2026 · checked 10 Oct 2026
  444. 444WorldArena leaderboard (Gradio app, tables read through its gradio_api reload_data endpoints)Leaderboard · checked 10 Oct 2026
  445. 445WorldArena leaderboard source: src/data_loader.py (EWMScore computation)Leaderboard · 15 Jul 2026 · checked 10 Oct 2026
  446. 446GigaBrain Challenge 2026 @ CVPR 2026 (tracks, schedule, final rankings, organizers)project-page · 2026 · checked 11 Oct 2026
  447. 447GigaBrain Challenge 2026 World Model Track Guideproject-page · 2026 · checked 11 Oct 2026
  448. 448World Model as Evaluator - Scoring Criteria (0-3 rubric and aggregation)project-page · 2026 · checked 11 Oct 2026
  449. 449World Model Track leaderboard results.json (Hugging Face space open-gigaai/CVPR-2026-WorldModel-Track-LeaderBoard)Leaderboard · 1 Jul 2026 · checked 11 Oct 2026
  450. 450open-gigaai/CVPR-2026-Workshop-WM-Track baseline and evaluation code (GitHub API licence: Apache-2.0; README)Repository · 6 Mar 2026 · checked 11 Oct 2026
  451. 451Hugging Face API: open-gigaai/CVPR-2026-WorldModel-Track-Dataset (gated; card holds 'GigaBrain Challenge 2026 Data & Model License Agreement')dataset-card · 5 Mar 2026 · checked 11 Oct 2026
  452. 452GigaWorld-1 full text (arXiv HTML v1), Sec. 4 (WMBench), 5.1, 6.5Paper · 2 Jul 2026 · checked 10 Oct 2026
  453. 453WorldArena Challenge @ CVPR 2026 pageproject-page · checked 10 Oct 2026
  454. 4541st Workshop on Video World Models (CVPR 2026) page, WorldArena Challenge sectionproject-page · checked 10 Oct 2026
  455. 455WorldArena Track 2 policy evaluation with a VLM judge (Policy_eval.md)Repository · checked 10 Oct 2026
  456. 456PlayWorld: Learning Robot World Models from Autonomous Play (arXiv abstract; v1 2026-03-09, v3 2026-04-06)Paper · 9 Mar 2026 · checked 10 Oct 2026
  457. 457PlayWorld full text (arXiv HTML v3), Sec. 4.4 and Fig. 7Paper · 6 Apr 2026 · checked 10 Oct 2026
  458. 458PlayWorld Fig. 7: policy evaluation success-rate correlation (RMSE and r per training-data source)Paper · 6 Apr 2026 · checked 10 Oct 2026
  459. 459PlayWorld arXiv HTML v1 (checked that Pearson 0.8766 and 18 policies already appear)Paper · 9 Mar 2026 · checked 11 Oct 2026
  460. 460PlayWorld project page (links Code and Data)project-page · checked 11 Oct 2026
  461. 461irom-princeton/open-world repository (linked as Code; GitHub API shows no licence)Repository · 10 Feb 2026 · checked 11 Oct 2026
  462. 462Hugging Face API: tennyyyin/playworld_dataset_preview (no licence field, gated)dataset-card · 28 Mar 2026 · checked 11 Oct 2026
  463. 463Interactive World Simulator for Robot Policy Training and Evaluation (arXiv abstract, v1)Paper · 9 Mar 2026 · checked 10 Oct 2026
  464. 464Interactive World Simulator full text (arXiv HTML v1), Sec. IV-A, IV-C, IV-DPaper · 9 Mar 2026 · checked 10 Oct 2026
  465. 465Interactive World Simulator Fig. 7: world-simulator vs real task scores per task (r values)Paper · 9 Mar 2026 · checked 10 Oct 2026
  466. 466RSS XXII (2026) proceedings page p018, DOI 10.15607/RSS.2026.XXII.018Paper · Jul 2026 · checked 10 Oct 2026
  467. 467RSS 2026 paper PDF p018 (Fig. 7 r values)Paper · Jul 2026 · checked 10 Oct 2026
  468. 468Interactive World Simulator project pageproject-page · checked 10 Oct 2026
  469. 469WangYixuan12/interactive_world_sim repository (LICENSE file; GitHub API reports 'Other')Repository · 10 Feb 2026 · checked 11 Oct 2026
  470. 470Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning (arXiv abstract; v1 2026-03-26, v2 2026-09-04; ECCV 2026)Paper · 26 Mar 2026 · checked 10 Oct 2026
  471. 471PersistWorld full text (arXiv HTML v2), 'Policy Evaluation' paragraph and Appendix 0.CPaper · 4 Sep 2026 · checked 10 Oct 2026
  472. 472PersistWorld arXiv HTML v1, Appendix 0.C (same Fig. 8 file as v2)Paper · 26 Mar 2026 · checked 10 Oct 2026
  473. 473PersistWorld Fig. 8: WM-to-real task progression, Ours vs Baseline (Ctrl-World)Paper · 4 Sep 2026 · checked 10 Oct 2026
  474. 474Jai2500/PersistWorld repository (GitHub API licence: MIT)Repository · 25 Mar 2026 · checked 11 Oct 2026
  475. 475PersistWorld project pageproject-page · checked 11 Oct 2026
  476. 476HF model card nvidia/Cosmos-H-Surgical-Simulator (Quantitative Evaluation: FDS, GATC, TCD)dataset-card · 9 Oct 2026 · checked 11 Oct 2026
  477. 477Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics (arXiv abs; v1 2026-04-22, v3 2026-06-04)Paper · 22 Apr 2026 · checked 10 Oct 2026
  478. 478Open-H-Embodiment full text (arXiv HTML v3), Cosmos-H-Surgical-Simulator sectionsPaper · 4 Jun 2026 · checked 11 Oct 2026
  479. 479GitHub API record for NVIDIA-Medtech/Cosmos-H-Surgical-Simulator (license Apache-2.0)Repository · checked 11 Oct 2026
  480. 480HF dataset nvidia/PhysicalAI-Robotics-Open-H-Embodiment metadata (license: cc-by-4.0)dataset-card · checked 11 Oct 2026
  481. 481RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation (arXiv abs; v1 2026-04-21, v2 2026-05-14)Paper · 21 Apr 2026 · checked 10 Oct 2026
  482. 482RoboWM-Bench full text (arXiv HTML v2)Paper · 14 May 2026 · checked 11 Oct 2026
  483. 483RoboWM-Bench project pageproject-page · checked 11 Oct 2026
  484. 484fffstrong/RoboWM-Bench README (links arXiv 2604.19092 and the project page)Repository · checked 11 Oct 2026
  485. 485GitHub API record for fffstrong/RoboWM-Bench (license: null; created 2026-04-15)Repository · checked 11 Oct 2026
  486. 486dWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Model (arXiv abstract, v1)Paper · 24 Apr 2026 · checked 10 Oct 2026
  487. 487dWorldEval full text (arXiv HTML v1), Sec. 4.1, 4.3, Fig. 7, AppendixPaper · 24 Apr 2026 · checked 10 Oct 2026
  488. 488dWorldEval Fig. 7: real vs generated success rates, panels (a)-(d)Paper · 24 Apr 2026 · checked 10 Oct 2026
  489. 489dWorldEval arXiv PDF v1, first page (author-to-affiliation mapping)Paper · 24 Apr 2026 · checked 10 Oct 2026
  490. 490dWorldEval project page (affiliations; 'Accepted by ICML 2026 Spotlight')project-page · checked 10 Oct 2026
  491. 491nvidia/Cosmos-HumanEval-v1 dataset carddataset-card · checked 10 Oct 2026
  492. 492WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform (arXiv abs)Paper · 18 May 2026 · checked 10 Oct 2026
  493. 493WorldArena 2.0 full text (arXiv HTML v1)Paper · 18 May 2026 · checked 10 Oct 2026
  494. 494WorldArena 2.0 project pageproject-page · checked 10 Oct 2026
  495. 495GitHub API record for WorldArena2/WorldArena-2.0 (license MIT; created 2026-05-05)Repository · checked 10 Oct 2026
  496. 496HF dataset WorldArena/WorldArena2.0 metadata (card license: apache-2.0)dataset-card · 29 Jul 2026 · checked 10 Oct 2026
  497. 497HF Space WorldArena/WorldArena2.0 metadata (card license: mit; last modified 2026-09-18)Leaderboard · 18 Sep 2026 · checked 10 Oct 2026
  498. 498WorldArena 2.0 leaderboard (Gradio app, tables read through its gradio_api reload_data endpoints)Leaderboard · checked 10 Oct 2026
  499. 499OSCAR: Omni-Embodiment Action-Conditioned World Model for Robotics (arXiv abstract; v1 2026-06-03, v2 2026-06-04)Paper · 3 Jun 2026 · checked 11 Oct 2026
  500. 500OSCAR full text (arXiv HTML v2), Sec. 5.4 Table 4, Appendix A.10Paper · 4 Jun 2026 · checked 11 Oct 2026
  501. 501WEAVER, Better, Faster, Longer: An Effective World Model for Robotic Manipulation (arXiv abstract; v1 2026-06-11, v2 2026-06-16)Paper · 11 Jun 2026 · checked 10 Oct 2026
  502. 502WEAVER full text (arXiv HTML v2), Sec. 3.4, 4, 5.2.1, Appendix A4.1 Table 8Paper · 16 Jun 2026 · checked 10 Oct 2026
  503. 503WEAVER Fig. 6: policy evaluation scatter plots (rho and MMRV per world model)Paper · 16 Jun 2026 · checked 10 Oct 2026
  504. 504WEAVER arXiv HTML v1 (checked that rho=0.870 and Table 8 values already appear)Paper · 11 Jun 2026 · checked 11 Oct 2026
  505. 505WEAVER project pageproject-page · checked 11 Oct 2026
  506. 506arnavkj1995/WEAVER repository (GitHub API licence: MIT)Repository · 8 May 2026 · checked 11 Oct 2026
  507. 507Hugging Face API: arnavkj1995/WEAVER model (no licence field)dataset-card · 8 Jun 2026 · checked 11 Oct 2026
  508. 508Hugging Face API: yilin-wu/droid_ood_data (no licence field)dataset-card · 10 Jun 2026 · checked 11 Oct 2026
  509. 509RoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy Evaluation (arXiv abs; v1 2026-07-01, v4 2026-07-15)Paper · 15 Jul 2026 · checked 10 Oct 2026
  510. 510RoboWorld full text (arXiv HTML v4)Paper · 15 Jul 2026 · checked 10 Oct 2026
  511. 511RoboWorld project pageproject-page · checked 11 Oct 2026
  512. 512GigaWorld-1 project page (WMBench leaderboard table)project-page · checked 11 Oct 2026
  513. 513Hugging Face API: open-gigaai/Giga-World-1 model (licence apache-2.0)dataset-card · 1 Jul 2026 · checked 11 Oct 2026
  514. 514WorldArena 2.0 Challenge, IROS 2026 Competition pageproject-page · checked 10 Oct 2026
  515. 515TriWorldBench: A Tri-View Consistency Perspective on Embodied World Models (arXiv abs)Paper · 22 Sep 2026 · checked 10 Oct 2026
  516. 516TriWorldBench full text (arXiv HTML v1)Paper · 22 Sep 2026 · checked 11 Oct 2026
  517. 517TriWorldBench GitHub README (news: code, val and test sets released and challenge launched 2026-07-28)Repository · 28 Jul 2026 · checked 11 Oct 2026
  518. 518TriWorldBench metric weights guide (Qwen3-VL-8B-Instruct judge)Repository · checked 11 Oct 2026
  519. 519GitHub API record for TriWorldBench/TriWorldBench (license: null; 235 stars)Repository · checked 11 Oct 2026
  520. 520HF dataset card TriWorldBench/Dataset (no license field)dataset-card · 28 Jul 2026 · checked 11 Oct 2026
  521. 521TriWorldBench website with leaderboard (55 published models)Leaderboard · checked 11 Oct 2026
  522. 522Pelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence (arXiv abstract, v1)Paper · 10 Sep 2026 · checked 10 Oct 2026
  523. 523Pelican-Sim 1.0 full text (arXiv HTML v1), Sec. 4.6, Tables 10-11Paper · 10 Sep 2026 · checked 10 Oct 2026
  524. 524Pelican-Sim 1.0 project pageproject-page · checked 11 Oct 2026
  525. 525ZouShilong1024/Pelican-Sim1.0 repository (GitHub API licence: Apache-2.0; README only)Repository · 11 Sep 2026 · checked 11 Oct 2026
  526. 526DexTouch-WM: Learning Action-Conditioned Tactile World Models from Human Touch for Dexterous Robot Manipulation (arXiv abstract; v1 2026-09-17, v2 2026-09-18)Paper · 17 Sep 2026 · checked 10 Oct 2026
  527. 527DexTouch-WM full text (arXiv HTML v2), Sec. IV-D, Tables II-IIIPaper · 18 Sep 2026 · checked 10 Oct 2026
  528. 528VBench: Comprehensive Benchmark Suite for Video Generative Models (arXiv abstract page)Paper · 29 Nov 2023 · checked 10 Oct 2026
  529. 529VBench (arXiv HTML v1)Paper · 29 Nov 2023 · checked 10 Oct 2026
  530. 530VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models (arXiv abstract page)Paper · 20 Nov 2024 · checked 10 Oct 2026
  531. 531VBench++ (arXiv HTML v1)Paper · 20 Nov 2024 · checked 10 Oct 2026
  532. 532VBench repository README and LICENSERepository · 21 Aug 2026 · checked 10 Oct 2026
  533. 533VBench score weights (scripts/constant.py)Repository · checked 10 Oct 2026
  534. 534VBench Leaderboard (Hugging Face Space; table values read from the app config)Leaderboard · checked 10 Oct 2026
  535. 535VideoPhy: Evaluating Physical Commonsense for Video Generation (arXiv abstract page)Paper · 5 Jun 2024 · checked 10 Oct 2026
  536. 536VideoPhy (arXiv HTML v2)Paper · 3 Oct 2024 · checked 10 Oct 2026
  537. 537VideoPhy repository README (human and automatic leaderboards)Repository · 30 Jan 2026 · checked 10 Oct 2026
  538. 538videophy_test_public dataset metadatadataset-card · 5 Jun 2024 · checked 10 Oct 2026
  539. 539videophy_train_public dataset metadatadataset-card · 9 Jun 2024 · checked 10 Oct 2026
  540. 540Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation (arXiv abstract page)Paper · 7 Oct 2024 · checked 10 Oct 2026
  541. 541PhyGenBench paper (arXiv HTML v1)Paper · 7 Oct 2024 · checked 10 Oct 2026
  542. 542PhyGenBench project pageproject-page · checked 10 Oct 2026
  543. 543PhyGenBench repository (README leaderboard, description)Repository · 25 Oct 2024 · checked 10 Oct 2026
  544. 544Do generative video models understand physical principles? (arXiv abstract page)Paper · 14 Jan 2025 · checked 10 Oct 2026
  545. 545Do generative video models understand physical principles? (arXiv HTML v3)Paper · 27 Feb 2025 · checked 10 Oct 2026
  546. 546Physics-IQ Benchmark project pageproject-page · checked 10 Oct 2026
  547. 547Physics-IQ LICENSERepository · checked 10 Oct 2026
  548. 548Physics-IQ and Physics-IQ Verified leaderboards (repo README)Leaderboard · 8 Oct 2026 · checked 10 Oct 2026
  549. 549VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation (arXiv abstract page)Paper · 9 Mar 2025 · checked 10 Oct 2026
  550. 550VideoPhy-2 (arXiv HTML v1)Paper · 9 Mar 2025 · checked 10 Oct 2026
  551. 551VideoPhy-2 project pageproject-page · checked 10 Oct 2026
  552. 552VideoPhy-2 README and human leaderboardLeaderboard · checked 10 Oct 2026
  553. 553videophy2_test dataset metadatadataset-card · 8 Mar 2025 · checked 10 Oct 2026
  554. 554videophy2_train dataset metadatadataset-card · 8 Mar 2025 · checked 10 Oct 2026
  555. 555Impossible Videos (arXiv abstract page)Paper · 18 Mar 2025 · checked 11 Oct 2026
  556. 556Impossible Videos (arXiv HTML v1)Paper · 18 Mar 2025 · checked 11 Oct 2026
  557. 557Impossible Videos, Proceedings of the 42nd International Conference on Machine Learning (PMLR 267)Paper · checked 11 Oct 2026
  558. 558ImpossibleVideos dataset metadatadataset-card · 21 Mar 2025 · checked 11 Oct 2026
  559. 559VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness (arXiv abstract page)Paper · 27 Mar 2025 · checked 10 Oct 2026
  560. 560VBench-2.0 (arXiv HTML v2)Paper · 20 Aug 2025 · checked 10 Oct 2026
  561. 561VBench-2.0 code folderRepository · checked 10 Oct 2026
  562. 562WorldScore: A Unified Evaluation Benchmark for World Generation (arXiv abstract page)Paper · 1 Apr 2025 · checked 10 Oct 2026
  563. 563WorldScore (arXiv HTML v2)Paper · 29 Nov 2025 · checked 10 Oct 2026
  564. 564WorldScore project pageproject-page · checked 11 Oct 2026
  565. 565WorldScore Leaderboard (leaderboard.csv)Leaderboard · 10 Sep 2026 · checked 11 Oct 2026
  566. 566WorldScore repository (MIT LICENSE)Repository · 23 Jul 2026 · checked 11 Oct 2026
  567. 567WorldScore dataset metadatadataset-card · 26 Mar 2025 · checked 11 Oct 2026
  568. 568Evaluating Newtonian Mechanics in Video Generative Models with Real Physical Systems (arXiv abstract page, v3)Paper · 3 Apr 2025 · checked 11 Oct 2026
  569. 569Morpheus: Benchmarking Physical Reasoning of Video Generative Models with Real Physical Experiments (arXiv v1 abstract)Paper · 3 Apr 2025 · checked 11 Oct 2026
  570. 570Morpheus (arXiv HTML v3)Paper · 29 Jun 2026 · checked 11 Oct 2026
  571. 571morpheus-real-world dataset metadatadataset-card · 4 Jul 2026 · checked 11 Oct 2026
  572. 572T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation (arXiv abstract page)Paper · 1 May 2025 · checked 11 Oct 2026
  573. 573T2VPhysBench (arXiv HTML v1)Paper · 1 May 2025 · checked 11 Oct 2026
  574. 574IntPhys 2Paper · 11 Jun 2025 · checked 10 Oct 2026
  575. 575A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video PairsPaper · 11 Jun 2025 · checked 10 Oct 2026
  576. 576CausalVQAPaper · 11 Jun 2025 · checked 10 Oct 2026
  577. 577WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning (arXiv abstract page)Paper · 4 Jun 2025 · checked 10 Oct 2026
  578. 578WorldPrediction (arXiv HTML v1)Paper · 4 Jun 2025 · checked 10 Oct 2026
  579. 579WorldPrediction repository README and LICENSERepository · 17 Dec 2025 · checked 11 Oct 2026
  580. 580PhyWorldBench: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models (arXiv abstract page)Paper · 17 Jul 2025 · checked 11 Oct 2026
  581. 581PhyWorldBench (arXiv HTML, latest version)Paper · 26 May 2026 · checked 11 Oct 2026
  582. 582PhyWorldBench in ICLR 2026 proceedingsPaper · checked 11 Oct 2026
  583. 583phyworldbench dataset metadatadataset-card · checked 11 Oct 2026
  584. 584WorldMark: A Unified Benchmark Suite for Interactive Video World Models (arXiv abs)Paper · 5 Aug 2026 · checked 11 Oct 2026
  585. 585WorldMark arXiv HTML full text v2 (Tables 3-4, Section 4.3, Appendix G)Paper · 5 Aug 2026 · checked 11 Oct 2026
  586. 586WorldMark arXiv HTML full text v1Paper · 23 Apr 2026 · checked 11 Oct 2026
  587. 587AlayaLab/WorldMark GitHub repositoryRepository · 5 Aug 2026 · checked 11 Oct 2026
  588. 588WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation (arXiv abs)Paper · 25 May 2026 · checked 11 Oct 2026
  589. 589WBench arXiv HTML full text v1 (Sections 5.1-5.4, Appendix D.2)Paper · 25 May 2026 · checked 11 Oct 2026
  590. 590meituan-longcat/WBench GitHub repository (README news, LICENSE)Repository · 8 Oct 2026 · checked 11 Oct 2026
  591. 591WBench dataset carddataset-card · 29 May 2026 · checked 11 Oct 2026
  592. 592Cosmos 3 report full text v1 (HWB text and numbers already present)Report · 1 Jun 2026 · checked 11 Oct 2026
  593. 593Physics-IQ Verified (arXiv abstract page)Paper · 17 Jun 2026 · checked 10 Oct 2026
  594. 594Physics-IQ Verified (arXiv HTML v1)Paper · 17 Jun 2026 · checked 10 Oct 2026
  595. 595Physics-IQ Verified (arXiv PDF v1, affiliation block)Paper · 17 Jun 2026 · checked 10 Oct 2026
  596. 596Physics-IQ benchmark repository README (Verified workflow, leaderboard, disclaimer)Repository · 8 Oct 2026 · checked 10 Oct 2026
  597. 597Physics-IQ Verified dataset metadata (Hugging Face API)dataset-card · 19 Jun 2026 · checked 10 Oct 2026
  598. 598Physics-IQ Verified leaderboard siteLeaderboard · checked 10 Oct 2026
  599. 599PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives (arXiv abs)Paper · 14 Aug 2026 · checked 11 Oct 2026
  600. 600PlayWorld arXiv HTML full text v2 (Tables 2, 3, 6; Figure 6)Paper · 14 Aug 2026 · checked 11 Oct 2026
  601. 601PlayWorld GitHub repository (redirects to hku-sail/PlayWorld)Repository · 23 Aug 2026 · checked 11 Oct 2026
  602. 602PlayWorld benchmark dataset (Hugging Face metadata)dataset-card · checked 11 Oct 2026
  603. 603PlayWorld leaderboard (Hugging Face Space, metadata via API)Leaderboard · checked 11 Oct 2026
  604. 604DeepMind Control SuitePaper · 2 Jan 2018 · checked 10 Oct 2026
  605. 605dm_control LICENSERepository · checked 10 Oct 2026
  606. 606Mastering Diverse Domains through World Models (DreamerV3)Paper · 10 Jan 2023 · checked 10 Oct 2026
  607. 607Model Based Reinforcement Learning for Atari (SimPLe)Paper · 1 Mar 2019 · checked 10 Oct 2026
  608. 608Arcade Learning Environment LICENSERepository · checked 10 Oct 2026
  609. 609Benchmarking the Spectrum of Agent Capabilities (Crafter)Paper · 14 Sep 2021 · checked 10 Oct 2026
  610. 610crafter LICENSERepository · checked 10 Oct 2026
  611. 611DriveArena: A Closed-loop Generative Simulation Platform for Autonomous Driving (arXiv abs)Paper · 1 Aug 2024 · checked 10 Oct 2026
  612. 612DriveArena arXiv PDF v1 (Tables 1 to 3)Paper · 1 Aug 2024 · checked 10 Oct 2026
  613. 613PJLab-ADG/DriveArena GitHub repository (README leaderboard, LICENSE)Repository · 17 Sep 2025 · checked 10 Oct 2026
  614. 614DriveArena World Dreamer weights (Hugging Face model card metadata)dataset-card · checked 10 Oct 2026
  615. 615DriveArena project pageproject-page · checked 10 Oct 2026
  616. 616ACT-Bench: Towards Action Controllable World Models for Autonomous Driving (arXiv abs)Paper · 6 Dec 2024 · checked 10 Oct 2026
  617. 617ACT-Bench arXiv HTML full text v1Paper · 6 Dec 2024 · checked 10 Oct 2026
  618. 618ACT-Bench project page (includes Terra v2 results)project-page · checked 10 Oct 2026
  619. 619turingmotors/ACT-Bench GitHub repositoryRepository · 23 Dec 2024 · checked 10 Oct 2026
  620. 620ACT-Bench dataset carddataset-card · 23 Dec 2024 · checked 10 Oct 2026
  621. 621Bench2Drive-R: Turning Real World Data into Reactive Closed-Loop Autonomous Driving Benchmark by Generative Model (arXiv abs)Paper · 11 Dec 2024 · checked 10 Oct 2026
  622. 622Bench2Drive-R arXiv HTML full text v1Paper · 11 Dec 2024 · checked 10 Oct 2026
  623. 623Bench2Drive-R arXiv PDF v1 (Tables 2, 3, 6)Paper · 11 Dec 2024 · checked 10 Oct 2026
  624. 624WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World (arXiv abs, v1 and v2)Paper · 1 Jun 2026 · checked 10 Oct 2026
  625. 625WorldLens arXiv HTML full text v2Paper · 1 Jun 2026 · checked 10 Oct 2026
  626. 626WorldLens arXiv PDF v2 (title page, Table 2, Section 14)Paper · 1 Jun 2026 · checked 10 Oct 2026
  627. 627WorldLens project pageproject-page · checked 10 Oct 2026
  628. 628worldbench/WorldLens GitHub repository (README, LICENSE, repo metadata)Repository · 18 Jan 2026 · checked 10 Oct 2026
  629. 629Hugging Face dataset worldbench/videogen (card metadata)dataset-card · 22 Dec 2025 · checked 10 Oct 2026
  630. 630WorldLens leaderboard (Hugging Face Space, metadata via API)Leaderboard · 10 Dec 2025 · checked 10 Oct 2026
  631. 631DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving (arXiv abs)Paper · 7 Mar 2026 · checked 11 Oct 2026
  632. 632DrivingGen arXiv HTML full text v2 (incl. Appendix B.8, B.9 and Figure 5)Paper · 7 Mar 2026 · checked 11 Oct 2026
  633. 633DrivingGen project pageproject-page · checked 11 Oct 2026
  634. 634youngzhou1999/DrivingGen GitHub repositoryRepository · 13 Mar 2026 · checked 11 Oct 2026
  635. 635DrivingGen dataset carddataset-card · 5 Jun 2026 · checked 11 Oct 2026
  636. 636Matrix-Game: Interactive World Foundation Model (arXiv abs, technical report)Paper · 23 Jun 2025 · checked 10 Oct 2026
  637. 637Matrix-Game arXiv HTML full text v1 (Section 5, Table 2, Figure 8)Paper · 23 Jun 2025 · checked 10 Oct 2026
  638. 638GameWorldScore folder in SkyworkAI/Matrix-Game (README, LICENSE, assets)Repository · 29 Sep 2026 · checked 10 Oct 2026
  639. 639WorldEval Figure 4: real vs WorldEval success rates per taskPaper · 25 May 2025 · checked 10 Oct 2026
  640. 640OpenVLA (arXiv abs, author list)Paper · checked 11 Oct 2026
  641. 641Veo report Figure 4: nominal real vs predicted success ratesReport · 6 Jan 2026 · checked 10 Oct 2026
  642. 642Runway figure: Real vs. GWM-1 Success RatesOfficial blog · 27 Feb 2026 · checked 10 Oct 2026
  643. 643RoboArena (arXiv abs, author list)Paper · 22 Jun 2025 · checked 11 Oct 2026
  644. 644PolaRiS Figure 7: Pearson r and MMRV by evaluation methodPaper · 30 Dec 2025 · checked 10 Oct 2026
  645. 645GigaWorld-1 Fig. 16: real vs generated success rate with fit linesPaper · 2 Jul 2026 · checked 10 Oct 2026
  646. 646GigaWorld-1 Fig. 17: Gen - Real success-rate difference per subtask and modelPaper · 2 Jul 2026 · checked 10 Oct 2026
  647. 647World-in-World Figure 5: SR vs generation quality and SR vs controllability in ARPaper · 16 Aug 2026 · checked 10 Oct 2026
  648. 648FailBench: How Reliable are VLMs at Judging Robot Task Success?Paper · 3 Sep 2026 · checked 10 Oct 2026
  649. 649PAI-Bench generation leaderboard data fileLeaderboard · checked 10 Oct 2026
  650. 650Veo (Robotics) Fig. 9: OOD axes, real vs predicted success (MMRV, Pearson)Paper · 6 Jan 2026 · checked 11 Oct 2026