Seeing is not following.

WorldEcho & WorldSync

Do robotic world models really follow actions?

Diagnosing and aligning action-conditioned generation for policy learning.

Research team

  • Sixiang Chen1
  • Jiaming Liu1,*
  • Jixian Wu2,3,*
  • Yichen Guo5,*
  • Tinghao Wang2,4,*
  • Siyuan Qian1
  • Hao Chen6
  • Jiajun Cao2,1
  • Jian Tang2,†
  • Shanghang Zhang1,†
  1. 1 State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
  2. 2 Beijing Innovation Center of Humanoid Robotics
  3. 3 New York University
  4. 4 University of Electronic Science and Technology of China
  5. 5 Nanyang Technological University
  6. 6 The Chinese University of Hong Kong

* Core contributors† Corresponding authors

arXiv:2608.24885 · August 2026

Abstract

Action-conditioned world models are increasingly used as learned simulators for policy evaluation and improvement, yet their effectiveness rests on an unverified assumption: generated futures faithfully reflect arbitrary valid actions. Existing benchmarks are typically confined to expert demonstrations, leaving off-expert action following inadequately evaluated. To address this gap, we introduce WorldEcho, which probes action following over a broader action distribution using visual integrity and SE(3) trajectory alignment. Our diagnosis shows that current world models reasonably execute expert actions but struggle with diverse off-expert trajectories, either ignoring the commanded actions or producing visually invalid rollouts. We further propose WorldSync, which strengthens action following along three complementary axes: distributional coverage, representational grounding, and intervention-effect alignment. It broadens the training distribution over action consequences, grounds intermediate video representations in action-induced robot dynamics through an Action-Forcing Expert, and aligns predicted changes under action interventions with the corresponding changes in ground-truth futures. Experiments on RoboTwin benchmarks and real-robot tasks show that WorldSync improves WorldEcho metrics and serves as a more reliable simulator for iterative policy improvement, enabling policies to achieve higher success rates.

arXiv:2608.24885

Part I / The action-following problem

When world models
fail to follow actions.

An action-conditioned world model should follow the commanded motion. Compare the simulator with two observed failure modes: visual collapse, where robot geometry breaks down, and action mismatch, where a plausible video moves the wrong way.

Input

Reference frame

Grab roller

First recorded simulator frame.

First recorded simulator frame for Grab roller

Scroll horizontally to compare all three videos.

GT

Simulator ground truth

The simulator reference for the recorded action query.

Model output 1

Visual collapse

Robot geometry breaks, blurs, or disappears under the query.

Model output 2

Action mismatch

The video looks plausible, but its end-effector motion disagrees.

Input shows the first recorded simulator frame; the clips do not include numerical action inputs. Model outputs 1 and 2 identify the two examples, not separate model identities.

WorldEcho Part I / Action-following benchmark

From observed failures
to measurable errors.

To quantify these failures, we introduce WorldEcho, a benchmark for action following beyond expert replay. It checks both whether a generated future is visually valid and whether the robot motion agrees with the consequences of the commanded actions.

Its five query types span demonstrated actions, cross-state replay, local perturbations, policy rollouts, and feasible-space sampling. For each query, WorldEcho compares model and simulator futures from the same initial state and action sequence, jointly assessing visual integrity and SE(3) end-effector alignment.

Figure 3 · WorldEcho benchmark
Paper Figure 3 WorldEcho jointly evaluates visual integrity and SE(3) end-effector alignment. Valid rollouts use NDTW; invalid rollouts receive the fixed penalty κ. Figure PDF ↗

Visual integrity gate

Can this rollout be evaluated at all?

Quality, temporal smoothness, end-effector visibility, and arm integrity must all pass.

Trajectory alignment

Did the robot move as that action should?

AnyPos recovers both end-effector trajectories; pose-aware NDTW compares them in SE(3).

Integrity-gated score

Invalid videos cannot hide behind motion error.

Valid rollouts use trajectory error; visually invalid rollouts receive a fixed penalty.

WorldSync Part II / World model training

Train a world model
that follows actions.

Guided by the WorldEcho diagnosis, WorldSync trains the world model to improve action following through expanded action coverage, auxiliary Action-Forcing Expert supervision, and intervention-effect alignment.

Paper Figure 4 WorldSync jointly optimizes expanded action coverage, the Action-Forcing Expert, and intervention-effect supervision for faithful action following. Figure PDF ↗
  1. 01

    Coverage

    Action coverage expansion

    Mix diverse simulated expert and off-expert trajectories with target-domain real expert data in a shared relative Cartesian action space.

  2. 02

    Grounding

    Action-Forcing Expert

    Decode future end-effector trajectories from intermediate video features so the representation encodes action-induced robot dynamics.

  3. 03

    Sensitivity

    Intervention-effect supervision

    Pair the same observation with different actions and match the change in predicted futures to the change in ground-truth futures.

WorldEcho / Leaderboard

How well do models
follow actions?

Compare action following and visual integrity across 50 RoboTwin tasks. Ranked by integrity-gated error, with raw trajectory error and visual pass rate alongside.

50 tasks 7 models

Protocol & results ↗

Scroll horizontally to compare all metrics.

Published WorldEcho main comparison: 7 models across 50 RoboTwin tasks, ranked by integrity-gated error.
RankModelTraining dataGated error ↓Raw NDTW ↓Visual pass (%) ↑
Rank 1Peking UniversityWorldSync Ours Expanded60k updates 0.06610.022384.51
Rank 2Stanford UniversityCtrlWorld Expert20k updates 0.07160.026683.89
Rank 3NVIDIADreamDojo Expert20k updates 0.08050.021078.97
Rank 4NVIDIACosmos-Predict2.5 Expert20k updates 0.08940.019075.09
Rank 5Tsinghua UniversityMotus Expert20k updates 0.11160.054875.09
Rank 6Ant GroupLingBotVA Expert20k updates 0.11480.047371.83
Rank 7NVIDIACosmos3 Expert20k updates 0.14320.057263.94

Logos identify the lead institution, developer, or parent organization. Model links open the full project credits.

Published task-macro averages over 50 RoboTwin tasks. Baselines use expert demonstrations (20k updates); WorldSync uses expanded action coverage (60k updates). Training data and budgets differ. Expanded baseline results remain in the full paper.

WorldEcho / Evaluation coverage

Evaluating beyond expert demonstrations.

WorldEcho broadens evaluation with four complementary off-expert query types, covering state–action combinations beyond expert demonstrations.

Paper Figure 5c: Expanded evaluation coverage beyond expert demonstrations, illustrated by a PCA projection of robot states and action sequences for the Adjust bottle task. Four off-expert query categories extend beyond the expert distribution; dashed contours show 95% highest-density regions.
Paper Figure 5c · Expanded evaluation coverageView original figure ↗

WorldSync / Interactive preview

Build an action sequence.

Choose a scene and draft commands for either arm. Explore the controls here before live world-model inference is connected.

Interface preview
Reference imageRecorded simulator frame
Recorded simulator reference frame for Grab roller

A frame from the existing experiment recording.

Model rolloutNot connected
Your generated future

A live model will generate a rollout from your action sequence here.

No inference is running in this preview.

Recorded examples do not respond to your draft.

Action sequence 0 / 48

No commands yet. Choose an arm and use the controls to start.

Live model not connected

Action draft

Left arm
Translation
Rotation
Gripper

Drafted change UI axes

Position · cm0.0 / 0.0 / 0.0
Rotation · °0 / 0 / 0
GripperUnchanged

Focus these controls to use the keys. Each press adds one command. X / Y / Z are preview axes; no robot motion is executed.

Ready to draft. No model is connected.

Rollout comparisons

Real-world experiments

Robot demo 1

Video forthcoming

Robot demo 2

Video forthcoming

Full-size image ↗ Figure PDF ↗