Weiheng Zhao1 · Haoyi Jiang1 · Xin Shi2 · Liu Liu3 · Zhizhong Su3 · Wei Sui2 · Fan Huang4 · Xinggang Wang1
Huazhong University of Science and Technology1 · D-Robotics2 · Horizon Robotics3 · Xiamen University4
The key insight behind Faster-WAM is that future representations are not merely an auxiliary training signal, but essential inference-time context for robust action prediction under distribution shifts. Guided by this principle, Faster-WAM computes future representations once and selectively reuses them during action denoising, reducing redundant video-action interaction. It achieves state-of-the-art in-distribution performance and robust OOD generalization across simulated and real-world manipulation, while substantially reducing inference latency.
- Release Progress
- File Structure
- Environment Setup
- Model Preparation
- Dataset Download
- Training
- Released Checkpoints
- Evaluation
- Latency
- Acknowledgments
- Citation
- Training and inference code. [✔]
- LIBERO, LIBERO-Plus, and RoboTwin evaluation code. [✔]
- Model checkpoints. [✔]
FasterWAM/
├── configs/
│ ├── data/ # LIBERO and RoboTwin dataset configs
│ ├── model/ # FastWAM, JointWAM, and FasterWAM models
│ ├── task/ # Benchmark-specific training configs
│ ├── sim_libero.yaml # LIBERO evaluation defaults
│ ├── sim_libero_plus.yaml # LIBERO-Plus evaluation defaults
│ └── sim_robotwin.yaml # RoboTwin evaluation defaults
├── environments/ # Independent benchmark uv projects and locks
├── scripts/
│ ├── train.py
│ ├── train_zero1.sh
│ ├── preprocess_sparse_action_dit_backbone.py
│ ├── precompute_text_embeds.py
│ └── eval_fasterwam_*.sh
├── experiments/
│ ├── libero/ # LIBERO evaluation manager and worker
│ └── robotwin/ # RoboTwin evaluation manager and policy adapter
├── src/fasterwam/ # Core model, dataset, and training code
├── third_party/RoboTwin/ # RoboTwin evaluation integration
├── checkpoints/ # Wan components and model checkpoints
├── data/ # Preprocessed training datasets
├── runs/ # Training outputs
└── evaluate_results/ # Evaluation outputs
The final system is FasterWAM. The repository also keeps the FastWAM and
JointWAM baselines for controlled comparisons.
Run all commands below from the repository root.
Install uv first. The core training environment
uses Python 3.10 and the PyTorch 2.7.1 CUDA 12.8 wheels locked in uv.lock:
bash scripts/setup/install_core.sh
source .venv/bin/activateFasterWAM uses Wan2.2-TI2V-5B. By default, missing components are downloaded
from Hugging Face and stored under ./checkpoints. Set the directory explicitly
before model preparation, training, or evaluation:
mkdir -p checkpoints
export DIFFSYNTH_MODEL_BASE_PATH="$(pwd)/checkpoints"To use ModelScope instead, additionally set:
export DIFFSYNTH_DOWNLOAD_SOURCE=modelscopeBefore training FasterWAM from scratch, generate its SparseActionDiT initialization from the Wan2.2 video DiT:
python scripts/preprocess_sparse_action_dit_backbone.py \
--model-config configs/model/fasterwam.yaml \
--output checkpoints/SparseActionDiT_cond_0_4_8_12_16_20_24_28_Wan22_alphascale_1024hdim.pt \
--device cuda \
--dtype bfloat16FastWAM and JointWAM use the dense ActionDiT initialization:
python scripts/preprocess_action_dit_backbone.py \
--model-config configs/model/fastwam.yaml \
--output checkpoints/ActionDiT_linear_interp_Wan22_alphascale_1024hdim.pt \
--device cuda \
--dtype bfloat16FasterWAM uses the same preprocessed MuJoCo 3.3.2 LIBERO dataset as FastWAM:
Download the four archives and extract them under data/libero_mujoco3.3.2:
mkdir -p data/libero_mujoco3.3.2
huggingface-cli download yuanty/LIBERO-fastwam \
--repo-type dataset \
--local-dir data/libero_mujoco3.3.2
cd data/libero_mujoco3.3.2
for f in *.tar.gz; do tar -xzf "$f"; done
cd ../..The resulting layout must be:
data/libero_mujoco3.3.2/
├── libero_10_no_noops_lerobot/
├── libero_goal_no_noops_lerobot/
├── libero_object_no_noops_lerobot/
└── libero_spatial_no_noops_lerobot/
The preprocessed RoboTwin dataset is available from:
Download all split archives, concatenate them, and extract them as described in the FastWAM release:
mkdir -p data/robotwin2.0
huggingface-cli download yuanty/robotwin2.0-fastwam \
--repo-type dataset \
--local-dir data/robotwin2.0
cd data/robotwin2.0
cat robotwin2.0.tar.gz.part-* | tar -xzf -
cd ../..The expected layout is:
data/robotwin2.0/
├── dataset_stats.json
└── robotwin2.0/
├── data/
├── meta/
└── videos/
Training reads cached T5 instruction embeddings. Generate them once after the dataset has been extracted:
# LIBERO
python scripts/precompute_text_embeds.py task=libero_fasterwam_2cam224_1e-4
# RoboTwin
python scripts/precompute_text_embeds.py task=robotwin_fasterwam_3cam_384_1e-4For multi-GPU preprocessing:
torchrun --standalone --nproc_per_node=8 \
scripts/precompute_text_embeds.py \
task=libero_fasterwam_2cam224_1e-4The caches are written to data/text_embeds_cache_fasterwam/libero and
data/text_embeds_cache_fasterwam/robotwin.
NPROC_PER_NODE=8 bash scripts/train_fasterwam_libero.sh
NPROC_PER_NODE=8 bash scripts/train_fasterwam_robotwin.shBoth wrappers accept additional Hydra overrides. For example:
NPROC_PER_NODE=8 bash scripts/train_fasterwam_libero.sh \
batch_size=8 \
num_epochs=1 \
wandb.enabled=trueThe released FasterWAM checkpoints and their corresponding dataset statistics are available on Hugging Face.
mkdir -p checkpoints/fasterwam_release
huggingface-cli download hustvl/FasterWAM \
--local-dir checkpoints/fasterwam_releaseAfter downloading, the checkpoint directory should have the following layout:
checkpoints/fasterwam_release/
├── libero/
│ ├── step_021700.pt
│ └── dataset_stats.json
└── robotwin/
├── step_029355.pt
└── dataset_stats.json
LIBERO, LIBERO-Plus, and RoboTwin are managed in three separate uv environments to isolate their simulator dependencies. Run the corresponding setup command once before evaluation; each evaluation launcher automatically uses the matching environment.
bash scripts/setup/install_libero.sh
TASK_NAME=libero_fasterwam_2cam224_1e-4 \
CKPT_PATH=checkpoints/fasterwam_release/libero/step_021700.pt \
DATASET_STATS_PATH=checkpoints/fasterwam_release/libero/dataset_stats.json \
NUM_GPUS=8 \
bash scripts/eval_fasterwam_libero.shbash scripts/setup/install_libero_plus.sh
TASK_NAME=libero_fasterwam_2cam224_1e-4 \
CKPT_PATH=checkpoints/fasterwam_release/libero/step_021700.pt \
DATASET_STATS_PATH=checkpoints/fasterwam_release/libero/dataset_stats.json \
NUM_GPUS=8 \
bash scripts/eval_fasterwam_libero_plus.shbash scripts/setup/install_robotwin.sh
TASK_NAME=robotwin_fasterwam_3cam_384_1e-4 \
CKPT_PATH=checkpoints/fasterwam_release/robotwin/step_029355.pt \
DATASET_STATS_PATH=checkpoints/fasterwam_release/robotwin/dataset_stats.json \
NUM_GPUS=8 \
bash scripts/eval_fasterwam_robotwin.shTo measure the Inference Latency:
# All models
bash scripts/measure_latency.sh
# Selected models: jointwam, fastwam, fasterwam
bash scripts/measure_latency.sh --models jointwam fasterwamOur codebase is built upon:
- FastWAM: https://github.com/yuantianyuan01/FastWAM
- Wan2.2: https://github.com/Wan-Video/Wan2.2
- LIBERO: https://github.com/Lifelong-Robot-Learning/LIBERO
- LIBERO-Plus: https://github.com/sylvestf/LIBERO-plus
- RoboTwin: https://github.com/RoboTwin-Platform/RoboTwin
We thank these teams for contributing their impressive code and models to the community.
If you find this repository helpful for your research, please consider citing our paper:
@article{zhao2026faster,
title = {Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models},
author = {Zhao, Weiheng and Jiang, Haoyi and Shi, Xin and Liu, Liu and Huang, Fan and Su, Zhizhong and Sui, Wei and Wang, Xinggang},
journal = {arXiv preprint arXiv:2608.04404},
year = {2026}
}