sudo apt-get update
sudo apt-get install -y git git-lfs tmux xvfb xauth libegl1 libgl1-mesa-glx
git lfs installcd /path/to/lerobot
conda create -n wepvla python=3.10 -y
conda activate wepvla
pip install -e ".[smolvla,libero]"For the bundled RLBench implementation:
pip install -e benchmarks/RLBench
python -c "import pyrep, rlbench; print('RLBench import ok')"Verify the runtime before starting an experiment:
python -c "import torch; print(torch.__version__, torch.cuda.is_available(), torch.cuda.device_count())"
nvidia-smiThe repository also includes dependency records in benchmarks/requirements.txt and
benchmarks/environment-rlbench.yml. CUDA extensions must be rebuilt for the target machine's
PyTorch and CUDA versions; exported local paths are not portable installation instructions.
The experiments use two local resources:
SmolVLM2-500M-Video-Instruct: VLM architecture, tokenizer, and processor;smolvla_base: pretrained SmolVLA weights.
Download them with:
bash benchmarks/RLBench/scripts/download_vlm_models.shThe expected layout is:
benchmarks/vlm_model/
├── SmolVLM2-500M-Video-Instruct/
└── smolvla_base/
For offline runs, pass the model paths directly to the command that needs them, for example:
VLM_MODEL_NAME=/path/to/SmolVLM2-500M-Video-Instruct \
VLM_WEIGHTS_PATH=/path/to/smolvla_base \
PYTHON=python \
bash benchmarks/RLBench/scripts/collect_data.sh \
--dataset-root /path/to/rlbench_datasetAll commands below are run from the repository root.
RLBench requires a compatible CoppeliaSim installation. Download it from https://www.coppeliarobotics.com/downloads and pass its path to each command:
COPPELIASIM_ROOT=/path/to/CoppeliaSim \
LD_LIBRARY_PATH="/path/to/CoppeliaSim:${LD_LIBRARY_PATH:-}" \
QT_QPA_PLATFORM=xcb \
QT_QPA_PLATFORM_PLUGIN_PATH=/path/to/CoppeliaSim \
QT_PLUGIN_PATH="" \
bash benchmarks/RLBench/scripts/evaluate.sh --helpOn a headless server:
Xvfb :99 -screen 0 1280x1024x24 -nolisten tcp >/tmp/rlbench-xvfb.log 2>&1 &PYTHON=python \
DATASET_ROOT=/path/to/rlbench_dataset \
COPPELIASIM_ROOT=/path/to/CoppeliaSim \
LD_LIBRARY_PATH="/path/to/CoppeliaSim:${LD_LIBRARY_PATH:-}" \
QT_QPA_PLATFORM=xcb \
QT_QPA_PLATFORM_PLUGIN_PATH=/path/to/CoppeliaSim \
QT_PLUGIN_PATH="" \
bash benchmarks/RLBench/scripts/collect_data.sh \
--dataset-root /path/to/rlbench_datasetThe collection flow writes the LeRobot dataset and the PointSeg cache. To rebuild only the cache,
use the cache utility under benchmarks/RLBench/scripts/tools/.
DATASET_ROOT=/path/to/rlbench_dataset \
OUTPUT_ROOT=/path/to/rlbench_output \
GPU_IDS=0 \
bash benchmarks/RLBench/scripts/train.shEVAL_POLICY_PATH=/path/to/checkpoint/pretrained_model \
EVAL_ROOT=/path/to/rlbench_eval \
EVAL_SAVE_VIDEO=0 \
EVAL_SAVE_ACTION_RECORDS=0 \
EVAL_SAVE_ACTION_CHUNKS=0 \
DISPLAY=:99 \
bash benchmarks/RLBench/scripts/evaluate.sh \
--tasks close_box close_fridge close_laptop_lid phone_on_base stack_wine \
sweep_to_dustpan take_frame_off_hanger \
take_umbrella_out_of_umbrella_stand toilet_seat_down water_plants \
--episodes 100Use tmux for long-running evaluations:
tmux new -d -s rlbench_eval \
"DISPLAY=:99 bash benchmarks/RLBench/scripts/evaluate.sh --episodes 100"The four LIBERO entry points are under benchmarks/song_real_libero/.
PYTHON_BIN=/path/to/python \
DEMO_ROOT=/path/to/libero_demos \
DATASET_ROOT=/path/to/libero_dataset \
bash benchmarks/song_real_libero/prepare_dataset.shPYTHON_BIN=/path/to/python \
DATASET_ROOT=/path/to/libero_dataset \
CACHE_ROOT=/path/to/libero_cache \
GPU_IDS=0 NPROC=1 \
bash benchmarks/song_real_libero/build_cache.shPYTHON_BIN=/path/to/python \
DATASET_ROOT=/path/to/libero_dataset \
CACHE_ROOT=/path/to/libero_cache \
BASE_POLICY=/path/to/base_policy/pretrained_model \
OUTPUT_ROOT=/path/to/libero_output \
GPU_IDS=0 \
bash benchmarks/song_real_libero/train.shPYTHON_BIN=/path/to/python \
POLICY_PATH=/path/to/checkpoint/pretrained_model \
OUTPUT_DIR="benchmarks/song_real_libero/outputs/eval_$(date +%Y%m%d_%H%M%S)" \
CUDA_DEVICE=0 EPISODES=50 \
bash benchmarks/song_real_libero/evaluate.shThe default LIBERO evaluator uses two task workers, one episode shard per worker, and
inference-batch-size=2. Every run requires a new output directory and refuses to overwrite an
existing result.