1X World Model Challenge Entry on this site | How well models predict future first-person frames of 1X's EVE humanoid from past frames and actions.352353 | 1X Technologies (2025 phases with OpenDriveLab) | Jun 2024 |
| EVA-Bench (Embodied Video Anticipation Benchmark) | Embodied video anticipation by world models that combine a vision-language model and a video generator, across four meta-tasks: Action-Description, How-To, Finish-Thinking and Next-Step, on real robots, simulated robots and egocentric human activities, with in-domain and out-of-distribution samples.354355356 | Hong Kong University of Science and Technology, Peking University (State Key Laboratory of Multimedia Information Processing) | 20 Oct 2024 |
| WorldSimBench | Video generation models used as world simulators for embodied agents, in three scenarios: an open-ended embodied environment (Minecraft via MineRL), autonomous driving (CARLA) and robot manipulation (CALVIN). It checks perceived visual quality per embodied dimension and whether generated videos can be turned into correct control signals.357358359+3 | arXiv and project page: The Chinese University of Hong Kong, Shenzhen; Shanghai Artificial Intelligence Laboratory; Beihang University; The University of Hong Kong. PMLR version adds Sun Yat-sen University, University of Oxford and the Guangdong Key Laboratory of Big Data Analysis and Processing. | 23 Oct 2024 |
| WorldModelBench | Whether text-to-video and image-to-video generators behave as world models in seven application domains: autonomous driving, robotics, human activities, industrial, natural scenes, simulation gaming and animation. It checks instruction following, physics adherence and commonsense (frame-wise and temporal quality).363364365+4 | UC Berkeley; UC San Diego; NVIDIA; MIT | 28 Feb 2025 |
EWMBench Entry on this site | How closely robot-manipulation videos from a video generator match real AgiBot World episodes in scene, motion and task meaning.370 | AgiBot; Shanghai Jiao Tong University; MMLab-CUHK; Harbin Institute of Technology (arXiv header) | May 2025 |
AgiBot World Challenge, World Model track Entry on this site | How well models predict robot head-camera video from actions on held-out AgiBot World episodes.371 | AgiBot (2025 with OpenDriveLab) | May 2025 |
| EnerVerse-AC (EVAC) as policy evaluator | Whether success rates of a Go-1 policy evaluated inside the EVAC action-conditioned world model match real-robot success rates across four retrieval tasks and across three training steps.372373374+3 | AgiBot; Shanghai Jiao Tong University; MMLab, CUHK | 14 May 2025 |
| DreamGen Bench | How well image-to-video world models adapt to a target robot embodiment and generalise to unseen objects, behaviours and environments: whether generated robot videos follow the instruction and obey physics. Setups: RoboCasa (simulated Franka) and three real Fourier GR1 humanoid splits (Object, Behavior, Environment).54377378+4 | NVIDIA (GEAR lab, lead), University of Washington, KAIST, UCLA, UCSD, Caltech, NTU, University of Maryland, UT Austin | 19 May 2025 |
| WorldEval | Whether success rates of real-robot manipulation policies, obtained by rolling each policy out inside an action-conditioned video world model, track and rank the same policies' real-robot success rates.298383299+1 | Midea Group; East China Normal University | 25 May 2025 |
| WorldGym | Whether success rates of VLA policies, obtained by Monte Carlo rollouts in an autoregressive action-conditioned video world model started from real first frames, match real-robot success rates and keep policy rankings. Also used to test policies on edited out-of-distribution scenes and instructions.297385386+3 | Stanford University; NYU; Google DeepMind | 31 May 2025 |
| Real-robot goal-image planning test (V-JEPA 2-AC) | Whether an action-conditioned world model can drive a real robot by planning: the robot is given goal images and the world model is used to search for actions, with no task-specific training or reward.390 | FAIR at Meta (V-JEPA 2 paper) | Jun 2025 |
| 1X World Model evaluation | How often the world model's success/failure prediction matches real outcomes, and whether its scores pick the same checkpoints and architectures as real double-blind A/B evaluations on humanoid robots.38391 | 1X Technologies | 16 Jun 2025 |
| WM-ABench | Whether vision-language models (not video generators) behave as internal world models, using atomic tests of perception (visual, spatial, temporal, quantitative, motion) and prediction (mechanistic simulation, transitive inference, compositional inference), with controlled counterfactual simulations.392393394+1 | Maitrix.org, UC San Diego, Johns Hopkins University, Cornell Tech, EPFL, University of Michigan | 27 Jun 2025 |
| IRASim policy evaluation (LIBERO) | Whether success rates of diffusion-policy checkpoints judged in IRASim-generated rollouts match the success rates measured in the LIBERO MuJoCo simulator.396397398+1 | Hong Kong University of Science and Technology; ByteDance Seed | 29 Jul 2025 |
PAI-Bench (Physical AI Bench) Entry on this site | Video generation and video understanding on physical-AI clips in six domains: common sense, driving, robots, industry, people, physics.400 | Georgia Tech; Carnegie Mellon University (NVIDIA acknowledged for support) | Sep 2025 |
| Ctrl-World | Whether the instruction-following rate and the task success rate of generalist DROID policies, rolled out closed-loop inside a multi-view action-conditioned video world model, match the same policies' rates on a real robot in a new DROID setup.80401402+5 | Stanford University; Tsinghua University | 11 Oct 2025 |
| Cosmos-Surg-dVRK | Whether success rates of surgical robot policies, rolled out online in a fine-tuned Cosmos world foundation model, match success rates of the same policies on a real da Vinci Research Kit (dVRK Si).406407408 | NVIDIA; Johns Hopkins University; Stanford University | 17 Oct 2025 |
| World-in-World | Whether visual world models help an embodied agent succeed in closed-loop tasks: Active Recognition (AR), Image-Goal Navigation (ImageNav), Active Embodied Question Answering (A-EQA) and robotic manipulation. Each world model is plugged into the same proposal-simulation-revision planning loop through a unified action API (text prompt, camera trajectory or low-level actions).409410411+4 | Johns Hopkins University (lead; corresponding author Jieneng Chen), Peking University, Princeton University, MIT, Harvard University | 20 Oct 2025 |
| Scalable Policy Evaluation with Video World Models (NVIDIA) | Whether success rates predicted by action-conditioned video models (post-trained Cosmos-Predict2-2B) with a VLM success judge match simulator success rates on four RoboMimic tasks and real success rates of three generalist policies on four Bridge-setup tasks.416417418 | NVIDIA Research; University of Toronto; Vector Institute | 14 Nov 2025 |
| Veo (Robotics) policy evaluation | Whether success rates of Gemini Robotics On-Device policy checkpoints, predicted by closed-loop rollouts in a Veo-based action-conditioned multi-view video model, match real ALOHA 2 success rates in nominal scenes and in edited out-of-distribution scenes, and whether the model finds unsafe behaviour that also occurs on the real robot.87419420 | Google DeepMind (Gemini Robotics Team) | 11 Dec 2025 |
PolaRiS Entry on this site | Whether policy scores in simulated copies of real scenes, built from short video scans with 2D Gaussian splatting, match real-robot scores of generalist DROID policies.421422423+2 | University of Washington; Princeton University; UC Berkeley; Stanford University; Toyota Research Institute; University of Southern California; Cornell University; Physical Intelligence | 18 Dec 2025 |
RBench Entry on this site | Whether video generators produce correct and physically plausible robot task videos across five task types and four robot body types.426 | Peking University; ByteDance Seed | Jan 2026 |
| WoW-World-Eval (Wow, wo, val) | Image-plus-text-to-video generation for robot manipulation, framed as an 'Embodied Turing Test' over five abilities: perception, planning, prediction, generalisation and execution. Includes a human Turing test (can people tell generated from real video) and an inverse-dynamics-model (IDM) Turing test (can actions recovered from generated video be executed on a real robot).427428429 | Peking University (State Key Laboratory of Multimedia Information Processing, School of Computer Science), Beijing Innovation Center of Humanoid Robotics, The Hong Kong University of Science and Technology | 7 Jan 2026 |
| DreamDojo policy evaluation (AgiBot fruit packing) | Whether the success rates of policy checkpoints simulated inside the DreamDojo world model match the success rates of the same checkpoints on a real robot, on one long-horizon fruit-packing task.93430431+2 | NVIDIA (lead); co-author affiliations also list HKUST, UC Berkeley, KAIST, University of Toronto, UC San Diego, University of Washington, Stanford, UT Austin | 6 Feb 2026 |
| WorldArena | How well embodied world models (text- or action-conditioned robot video models) predict bimanual manipulation video, and whether they are useful for three downstream jobs: generating training data for a policy (data engine), standing in for a simulator when scoring policies (policy evaluator), and planning actions (action planner). Human ratings are collected as a third view.433434435+10 | Tsinghua University (lead; corresponding author Yong Li), Shanghai Jiao Tong University, The University of Hong Kong, Princeton University, Chinese Academy of Sciences, University of Science and Technology of China, Peking University, National University of Singapore | 9 Feb 2026 |
| GWM-Robotics policy evaluation | Whether progress scores of VLA policies rolled out inside Runway's GWM-Robotics world model rank the policies in the same order as their real-world RoboArena evaluations.290224 | Runway | 27 Feb 2026 |
| GigaBrain Challenge 2026 - World Model Track (CVPR 2026 workshop) | World models as evaluators of a VLA policy (GigaBrain) on 8 real-robot manipulation tasks: video quality when replaying teleoperation actions, and whether closed-loop rollouts driven by policy actions reach the same outcome as the real-robot reference video.446447448+4 | GigaAI (organizers Zheng Zhu, Xiaofeng Wang) with co-organizers from University of Hong Kong, Peking University, Shanghai Jiao Tong University, RoboChallenge/Dexmal and Horizon Robotics | Mar 2026 |
| WorldArena Challenge @ CVPR 2026 | Two tracks built on WorldArena: Track 1 video perception quality of embodied world models; Track 2 whether generated worlds work as data engines and as policy evaluators.453454437+4 | Organisers listed on the challenge page: Amap CV Lab (AMAP), Manifold AI, Tsinghua University, Princeton University, National University of Singapore, The University of Hong Kong | Mar 2026 |
| PlayWorld policy evaluation | Whether success rates predicted by a video world model trained on autonomous robot play match real-robot success rates for many manipulation policies on contact-rich tasks, and whether predicted failure modes match observed ones.456457458+4 | Princeton University | 9 Mar 2026 |
| Interactive World Simulator sim-to-real policy evaluation | Whether task scores of imitation policies (final and intermediate checkpoints of DP, ACT, pi0 and pi0.5) evaluated closed-loop inside the learned world model match their real-robot scores on four ALOHA tasks.463464465+4 | Columbia University; Toyota Research Institute; Amazon; University of Illinois Urbana-Champaign | 9 Mar 2026 |
| PersistWorld world-model-to-real policy evaluation | Whether the task progress of robot policies rolled out inside an action-conditioned video world model matches their real-robot task progress, comparing PersistWorld (Ctrl-World post-trained with reinforcement learning on its own rollouts) with the base Ctrl-World.470471472+3 | Czech Institute of Informatics, Robotics and Cybernetics, Czech Technical University in Prague | 26 Mar 2026 |
| Cosmos-H-Surgical-Simulator evaluation on Open-H-Embodiment | How closely an action-conditioned surgical world model's generated video follows the recorded kinematic actions of held-out real surgical-robot episodes (open-loop replay), as groundwork for in-silico policy evaluation.476477478+2 | NVIDIA with the Open-H-Embodiment Consortium (50+ institutions in the paper's author list); model card by NVIDIA | Apr 2026 |
| RoboWM-Bench | Whether manipulation behaviour in videos generated by world models can be executed: generated human-hand or robot-arm videos are converted to robot actions and run in simulated scenes, including real-to-sim reconstructions of real tabletop scenes.481482483+2 | Peking University (lead; corresponding author Ruihai Wu), Tsinghua University, Lightwheel | 21 Apr 2026 |
| dWorldEval policy evaluation proxy | Whether success rates estimated inside a discrete-diffusion world model (with a progress token that marks task completion) match ground-truth success for pi0 checkpoints on LIBERO, heterogeneous policies on RoboTwin, and three policies on five real bimanual AgileX tasks; it also compares three video-diffusion world models on LIBERO.486487488+2 | Current Robotics; University of Toronto | 24 Apr 2026 |
Cosmos-HumanEval (Cosmos HUE) Entry on this site | Human judgement of physical laws, visual integrity, semantic alignment and geometry in generated videos.138491 | NVIDIA | May 2026 |
| WorldArena 2.0 | Extends WorldArena along three axes: visuotactile prediction (tactile plus video) for contact-rich tasks, world models used as interactive reinforcement-learning environments for policy improvement, and evaluation across two simulators and a real robot (RoboTwin 2.0, LIBERO, AgileX split-type ALOHA).492493494+4 | Tsinghua University (lead; corresponding author Yong Li), Shanghai Jiao Tong University, Zhejiang University, Stanford University, The University of Hong Kong, Princeton University, Chinese Academy of Sciences, University of Science and Technology of China, Peking University, National University of Singapore | 18 May 2026 |
| OSCAR policy evaluation on RoboArena | Whether OSCAR-generated replays of RoboArena episodes, judged by GPT-5, reproduce the real per-policy success rates and ranking of seven open-source DROID policies.499500 | Peking University; University of Michigan; NVIDIA | 3 Jun 2026 |
| WEAVER policy evaluation on real hardware | Whether success judged on rollouts imagined by the world model matches the real success rates of two pi0.5-based policies on five real manipulation tasks, compared with Ctrl-World and with the WEAVER model before task fine-tuning.501502503+5 | Mila - Quebec AI Institute; Universite de Montreal; Carnegie Mellon University; McGill University | 11 Jun 2026 |
RoboWorld Entry on this site | Whether scores from closed-loop rollouts of DROID policies in an autoregressive video world model match the RoboArena real-world leaderboard.509510511 | KAIST; Config | 1 Jul 2026 |
| WMBench (GigaWorld-1) | How well video world models serve as surrogate evaluators of robot policies, by comparing generated rollouts with matched real-robot executions on eight manipulation tasks built from teleoperation data and GigaBrain policy rollouts.295452512+3 | GigaAI; Tsinghua University | 2 Jul 2026 |
| WorldArena 2.0 Challenge @ IROS 2026 | Three tracks: Track 1 video quality of embodied world models with harder tasks and new out-of-distribution scenes; Track 2 world models as RL environments for policy optimisation; Track 3 real-world manipulation by world action models (WAM), in tactile and vision-only settings.514498497+2 | WorldArena 2.0 team (organisers not named on the challenge page; contact worldarenav2@outlook.com); paper team led by Tsinghua University | 10 Jul 2026 |
| TriWorldBench | Whether an embodied world model's synchronized head, left-wrist and right-wrist videos describe one consistent manipulation event, plus task alignment, physical and 3D coherence, motion quality, temporal consistency and visual quality.515516517+4 | Peking University, Tsinghua University, Beihang University, Shanghai Jiao Tong University, University of Science and Technology of China, Shanghai AI Laboratory (arXiv affiliation block); the website lists OpenCompass in place of Shanghai AI Laboratory | 28 Jul 2026 |
| Pelican-Sim 1.0 policy evaluation and ranking (RoboTwin) | Whether the Pelican-Sim world model combined with a fine-tuned VLM judge reproduces RoboTwin simulator success rates and the ranking of five checkpoints from one VLA training run.522523524+1 | Beijing Innovation Center of Humanoid Robotics (X-Humanoid), WFM System Group | 10 Sep 2026 |
| DexTouch-WM world models as policy evaluators | Whether scores of three VLA policies rolled out closed-loop inside two task-adapted visuo-tactile world models match their scores on a real dexterous-hand robot on four tasks.526527 | HKUST (Guangzhou); Xspark AI; Peking University; Tsinghua University; University of Hong Kong | 17 Sep 2026 |
| VBench (including VBench++) | General video generation quality split into 16 dimensions. Video Quality: subject consistency, background consistency, temporal flickering, motion smoothness, dynamic degree, aesthetic quality, imaging quality. Video-Condition Consistency: object class, multiple objects, human action, color, spatial relationship, scene, appearance style, temporal style, overall consistency. VBench++ adds image-to-video evaluation with an Image Suite, long-video evaluation, and trustworthiness (culture fairness, human bias, safety). There is no physics dimension.528529530+4 | S-Lab, Nanyang Technological University; Shanghai Artificial Intelligence Laboratory; Nanjing University; The Chinese University of Hong Kong | 29 Nov 2023 |
| VideoPhy | Whether text-to-video outputs follow the caption (semantic adherence) and physical commonsense in everyday interactions between materials: solid-solid, solid-fluid and fluid-fluid, with captions marked easy or hard by graphics researchers.535536537+2 | University of California Los Angeles; Google Research | 5 Jun 2024 |
| PhyGenBench (with PhyGenEval) | Whether text-to-video outputs show the correct physical phenomenon for prompts that each target one physical law, in four domains: mechanics, optics, thermal and material properties. Semantic alignment with the prompt is scored separately.540541542+1 | Shanghai Jiao Tong University; OpenGVLab, Shanghai AI Laboratory; The University of Hong Kong; The Chinese University of Hong Kong | 7 Oct 2024 |
| Physics-IQ | Whether a video generator can predict how a real, filmed physical experiment continues. Scenarios cover solid mechanics, fluid dynamics, optics, thermodynamics and magnetism. Each 8-second real video is split into a 3-second conditioning part and a 5-second test part. Video-to-video models get the 3 seconds; image-to-video models get only the last conditioning frame (the 'switch frame'); both can also get a text description that does not reveal the outcome.544545546+2 | Google DeepMind; INSAIT, Sofia University (first author, work done while at Google DeepMind) | 14 Jan 2025 |
| VideoPhy-2 | Physical commonsense and prompt adherence of text-to-video outputs for real-world actions (sports and physical activities, object interactions), plus whether specific physical rules for each video are followed.549550551+3 | University of California Los Angeles; Google Research | 9 Mar 2025 |
| IPV-Bench (Impossible Videos) | Whether video generators can follow prompts for impossible events that defy physical, biological, geographical or social laws while keeping high visual quality; separately, whether Video-LLMs understand impossible videos.555556557+1 | Show Lab, National University of Singapore | 18 Mar 2025 |
| VBench-2.0 | 'Intrinsic faithfulness' of generated video in five dimensions with 18 capabilities: Human Fidelity (human anatomy, identity, clothes); Creativity (diversity, composition); Controllability (dynamic spatial relationship, dynamic attribute, motion order understanding, human interaction, complex landscape, complex plot, camera motion); Physics (State Change: mechanics, thermotics, material; Geometry: multi-view consistency); Commonsense (motion rationality, instance preservation).559560561+1 | Shanghai Artificial Intelligence Laboratory; S-Lab, Nanyang Technological University; Sun Yat-sen University; The Chinese University of Hong Kong | 27 Mar 2025 |
| WorldScore | World generation from an image plus a camera-trajectory layout, posed as a sequence of next-scene tasks, so that 3D scene generators, 4D generators and video generators can be compared. Three aspects: controllability, quality and dynamics, over static worlds (indoor and outdoor) and dynamic worlds (5 motion types).562563564+3 | Stanford University | 1 Apr 2025 |
| Morpheus | Whether video generators conditioned on frames of real, filmed Newtonian rigid-body experiments continue them in line with the governing equations of motion and with conservation of energy, momentum and period.568569570+1 | University of Amsterdam; University of Trento | 3 Apr 2025 |
| T2VPhysBench | Whether text-to-video systems obey 12 physical laws in three groups: Newton's three laws and gravitation; conservation of energy, mass, linear and angular momentum; and phenomenological laws (Hooke's law, Snell's law, law of reflection, Bernoulli's principle). Also tests prompts with added law-specific hints and counterfactual prompts that ask for physics-breaking videos.572573 | Guilin University of Electronic Technology; University of Arizona; University of Wisconsin-Madison; Simons Institute for the Theory of Computing, UC Berkeley; Arizona State University | 1 May 2025 |
IntPhys 2 Entry on this site | Whether video models can tell physically possible from impossible events in synthetic videos.574 | FAIR at Meta | Jun 2025 |
MVPBench (Minimal Video Pairs) Entry on this site | Physical understanding of video-language models, with near-identical video pairs that have opposite answers.575 | FAIR at Meta | Jun 2025 |
CausalVQA Entry on this site | Cause-and-effect reasoning of video-language models on real Ego-Exo4D clips: counterfactual, anticipation and planning questions.576 | FAIR at Meta | Jun 2025 |
| WorldPrediction | High-level world modeling and long-horizon procedural planning on human activity video. WorldPrediction-WM: given an initial and a final state image, pick the action clip that causes the change. WorldPrediction-PP: pick the correctly ordered sequence of 3-10 action clips. Candidate actions are shown in other scenes ('action equivalents') so background continuity cannot be used.577578579 | Meta FAIR Paris; The Hong Kong University of Science and Technology; ISIR Sorbonne Université | 4 Jun 2025 |
| PhyWorldBench | Physical realism of text-to-video output across 10 physics categories, from object motion and energy conservation to rigid-body interactions and human or animal motion, plus an Anti-Physics category whose prompts ask for physics-violating events.580581582+1 | University of California, Santa Cruz; NVIDIA Research; Northeastern University; University of California, Santa Barbara | 17 Jul 2025 |
| WorldMark | Interactive image-to-video world models driven by navigation commands: how motion reacts to each command (direction, purity, latency, stability, per translation and rotation axis), whether the generated world stays the same over time (world memory at three time scales), and visual quality, across first- and third-person views, real and stylized scenes, and three difficulty tiers.584585586+1 | Alaya Lab; The University of Tokyo; Shanghai Innovation Institute | 23 Apr 2026 |
| WBench | Interactive video world models in multi-turn use: video quality, adherence to a stated world setting, adherence to interactions (navigation, subject action, event editing, perspective switching), consistency, and physics compliance.588589590+1 | Fudan University; Meituan LongCat Team | 25 May 2026 |
| Human World Bench (HWB) | Egocentric image-to-video generation of human manipulation tasks under task-level instructions: whether the video shows the requested actions and objects, and whether motion, object dynamics, contact and hand anatomy are physically plausible.99138592 | NVIDIA (Cosmos 3 technical report) | 1 Jun 2026 |
| Physics-IQ Verified | The same task and the same 66 real-world scenarios as Physics-IQ, after an audit that corrected prompts and ground-truth motion maps, so that the score reflects physical prediction rather than prompt ambiguity or unrelated motion in the recordings.593594595+3 | Anates Labs; Technical University of Munich; University of Technology Nuremberg; Tuebingen AI Center, University of Tuebingen; Helmholtz AI, Munich; Google DeepMind (Jaini and Geirhos, described as advisory in the acknowledgements) | 17 Jun 2026 |
| PlayWorld | Long-horizon capability of interactive video world models when a player pursues a stated objective: geometry consistency, interaction fidelity, out-of-sight evolution (what happens to things while unseen) and insight evolution (processes that should change while watched), plus basic video quality and action controllability.599600601+2 | The Chinese University of Hong Kong; The University of Hong Kong; Zhejiang University; Kling Team, Kuaishou Technology (as listed in the arXiv HTML; the code repository now resolves to github.com/hku-sail/PlayWorld) | 13 Aug 2026 |
| DeepMind Control Suite | Continuous-control performance of reinforcement-learning agents on simulated MuJoCo bodies (cart-pole, cheetah, walker, humanoid, manipulator and others); model-based agents such as Dreamer report results on it.604605606 | DeepMind (paper title and repository owner google-deepmind) | Jan 2018 |
| Atari 100k | How well a reinforcement-learning agent plays Atari games when it may interact with each game only briefly; world-model agents (SimPLe, Dreamer and others) use it to show that learning inside a model saves real interaction.607606608 | Introduced with SimPLe by Google Brain, deepsense.ai, Institute of Mathematics of the Polish Academy of Sciences, University of Warsaw, UIUC and Stanford (paper affiliations) | Mar 2019 |
| Crafter | Breadth of abilities of an agent in a 2D open-world survival game: exploration, resource collection, crafting, survival.609610606 | Danijar Hafner (Google Research, Brain Team; University of Toronto) | Sep 2021 |
| DriveArena | Open-loop and closed-loop driving performance of camera-based driving agents inside a generative simulation platform. A Traffic Manager simulates traffic on nuScenes, CARLA or OpenStreetMap road networks, and World Dreamer (a layout-conditioned diffusion model trained on nuScenes) generates the six surround-view camera images the agent sees at each step.611612613+2 | Shanghai Artificial Intelligence Laboratory; Zhejiang University; Shanghai Jiao Tong University; East China Normal University; Technical University of Munich | 1 Aug 2024 |
| ACT-Bench | Action fidelity of driving world models: whether a model that generates front-camera driving video from context frames plus a commanded trajectory actually shows the commanded motion.616617618+2 | Turing Inc. | 6 Dec 2024 |
| Bench2Drive-R | Reactive closed-loop evaluation of camera-based end-to-end driving models using real-world data: a rule-based behaviour controller built on the nuPlan simulator moves the other road users, and a diffusion-based generative renderer produces the camera images for each new state. Sensor rendering and behaviour control are separate modules. The paper also measures the renderer itself (image quality, layout adherence, temporal consistency).621622623 | Shanghai Jiao Tong University (Dept. of CSE, School of AI, and MoE Key Lab of AI) | 11 Dec 2024 |
| WorldLens | How well generative driving world models (multi-camera video generators conditioned on nuScenes scene layouts) produce video that looks real, keeps consistent 3D geometry, follows physics, lets a pretrained driving planner drive, and supports perception models trained on real data. Five aspects (Generation, Reconstruction, Action-Following, Downstream Task, Human Preference) with 24 dimensions in total.624625626+4 | WorldBench Team (the only group name on the arXiv paper, PDF title page and project page; no institutional affiliation block is given) | 11 Dec 2025 |
| DrivingGen | Generative video world models for driving (image-to-video, optionally conditioned on an ego trajectory): how real the video looks, how plausible the ego motion implied by the video is, temporal and per-agent consistency, and how closely the video follows a commanded ego trajectory, across varied weather, time of day, world regions and maneuvers.631632633+2 | University of Toronto; CUHK MMLab | 4 Jan 2026 |
| GameWorld Score | Minecraft world models that generate video from an initial image plus keyboard and mouse actions: visual quality, temporal quality, action controllability and physical rule understanding, in 8 dimensions.636637638 | Skywork AI | 23 Jun 2025 |