Skip to content
This repository was archived by the owner on Sep 21, 2026. It is now read-only.
IRLLPublic archive

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

Multi-Stew Game

This is Truong's human-AI teaming branch. This repo is static and out of date from the original multistew repository. The purpose of this repository is to preserve code — it is kept for display and reference, not for active development.

Credits

The Colyseus game server (backend/node/) and the React frontend (front/) were built in large part by Azadeh Mostaghel (@amostaghel) and Seti Rahmani (@Seti17) of Aggregate Intellect — together they account for the majority of the commits behind those two services. This branch layers the RL and human-AI teaming research on top of it.

The project was supervised by Amir Feizpour, CEO of Aggregate Intellect.

Architecture

  React (front)  <-->  Colyseus game server (node)  <--/predict_all-->  RL server (rl-multistew)
     :5173                      :2567                                        :8800
  • Frontend — React + Vite (front/)
  • Game server — Colyseus / Node + TypeScript (backend/node/), owns the kitchen state and levels
  • RL server — Python + FastAPI (backend/rl-multistew/), receives the full game state on POST /predict_all and returns one action per agent
  • Sherpa / LLM service — Python + FastAPI (backend/llm_service/)
  • Additional services: backend/rl-vt-collab/ (:8008), backend/human-eval/ (:8808)

Running the game

There are two ways to bring the stack up. Option B is recommended: the RL server pulls in big packages (PyTorch, OpenCV), so you do not want Docker building it every time.

Option A — everything in Docker

cp .env.example .env    # then fill in the variables you need
docker compose up --build

Endpoints:

Production variant:

export NODE_ENV=production
docker compose -f docker-compose.yml -f docker-compose.prod.yml up --build
# frontend on http://localhost:80

Option B — Docker for front + node, RL server on the host (recommended)

1. Point the game server at the host's RL server.

In docker-compose.yml, the node service sets RL_SERVICE_URL=http://rl-multistew:8800, which only resolves inside the compose network. Since the RL server now runs outside Docker, change it to:

- RL_SERVICE_URL=http://172.17.0.1:8800    # 172.17.0.1 maps to the host's localhost (Linux)
- RL_SERVICE_URL=http://host.docker.internal:8800   # use this on macOS

2. Build and start only the frontend and game server.

./runnode.sh
# which is just: sudo docker compose up --build --no-deps front node

--no-deps is what skips rl-multistew, sherpa, and the other dependencies declared in docker-compose.yml.

3. Set up a virtualenv and run the RL server as a plain Python script.

python3 -m venv env          # python >= 3.10
source env/bin/activate
pip3 install -r backend/rl-multistew/requirements.txt

python3 backend/rl-multistew/server.py -h   # lists every command-line argument

Run with image observations (CNN agent). This is the checkpoint shipped in backend/rl-multistew/CNN_models/ — the flags below are the exact ones it was trained with, so they must match or the weights will not load:

python3 backend/rl-multistew/server.py \
    --endpoint inference \
    --model_path "backend/rl-multistew/CNN_models/2_agents_overcooked_counter_circuit_v0_seed_2_image_RS_4_agents_vision_[100, 100]_modelT_CNNAgent_hiddenD_256_numH_1_FS_1_TS_16.pth" \
    --feature image \
    --model_type CNNAgent \
    --layout overcooked_counter_circuit_v0 \
    --num_agents 2 \
    --num_envs 1 \
    --num_steps 256 \
    --seed 2 \
    --region_size 4 \
    --agents_RS 100 100 \
    --hidden_dim 256 \
    --num_hidden_layers 1 \
    --grid_width 10 \
    --grid_height 7 \
    --frame_stack 1 \
    --tile_size 16 \
    --lr 1e-4 \
    --num_mini_batches 4 \
    --ppo_epochs 5 \
    --clip_param 0.2 \
    --value_loss_coef 0.5 \
    --entropy_coef 0.01 \
    --max_grad_norm 0.5 \
    --gamma 0.99 \
    --lam 0.95 \
    --normalize_advantages \
    --data_path path-to-data-folder

Quote the --model_path — the filename contains spaces and square brackets.

The checkpoint's filename encodes how it was trained (see save_model_checkpoint in rl_multistew/utils.py): 2 agents, overcooked_counter_circuit_v0, seed 2, image feature, region size 4, CNNAgent with hidden dim 256 / 1 hidden layer, frame stack 1, tile size 16. Because its layout is counter circuit, pick Level 6 - Counter Circuit in the frontend's level dropdown — it is one of the two levels left uncommented in App.jsx.

To train instead of serve, swap --endpoint inference for --endpoint ppo and add --max-steps, --log, and --checkpoint_dir.

Run with handcrafted feature observations (linear agent):

python3 backend/rl-multistew/server.py \
    --max-steps 10000 --num_agents 2 --num_envs 1 --num_steps 256 --seed 1 \
    --layout overcooked_counter_circuit_v0 --num_mini_batches 4 --ppo_epochs 5 --clip_param 0.2 \
    --value_loss_coef 0.5 --entropy_coef 0.01 --max_grad_norm 0.5 --gamma 0.99 --lam 0.95 \
    --data_path path-to-data-folder \
    --feature global \
    --model_path path-to-model \
    --region_size 100 \
    --endpoint inference \
    --model_type Default

Notes:

  • --model_path is the checkpoint loaded in inference mode.
  • --endpoint selects the request handler: inference, ppo, saliency, or hightlights.
  • --rl_port (default 8800) changes the FastAPI port.

4. In the browser, always tick "Synchronize Agents" in the Game Setup panel before starting a game. Without it the game server and the RL server step out of sync.


Reference: server.py options

Defined in command_line_parser.py.

--feature (observation type) — global, local, localv2, minimal_spatial_other_agent_aware, region_size_invariant, image, image_v2, image_v3.

--model_type (network) — Default, LinearAgent, CNNAgent, LayerNormCNNAgent. Any model_type containing CNNAgent activates the CNN path and uses --grid_width, --grid_height, --frame_stack, --tile_size; those four are ignored otherwise.

--region_size is only meaningful for the local features; 100 effectively means "unused".


Writing your own RL server endpoint

The game server sends its internal state to POST /predict_all, caught by predict_all(game_state) in server.py. That function handles one step of interaction and branches on the --endpoint flag:

@app.post("/predict_all", response_model=ActionResponse)
def predict_all(game_state: FullGameState):
    if mappo.args.endpoint == "hightlights":
        ...
    elif mappo.args.endpoint == "saliency":
        ...
    elif mappo.args.endpoint == "inference":
        ...
    else:
        assert mappo.args.endpoint == "ppo", f"Invalid endpoint specified: {mappo.args.endpoint}"

To add your own: create backend/rl-multistew/endpoints/myendpoint.py alongside the existing endpoints/, then add a branch for it in that if chain.


Adding a new level / map

Outdated. This is the procedure as it stood in this archived branch, and it is what the code here still expects — registering a level by hand in three separate overcookedLevels sets is exactly the kind of thing that has since changed upstream. Follow it to understand or run this snapshot; do not assume it still applies to the active multistew repository.

  1. Create the level file in backend/node/src/levels/, e.g. levels/yourlevel.txt. Follow an existing level for inspiration — the file specifies how many AI and human players can join.

  2. Add the yourlevel string to the overcookedLevels set in all three of:

    private overcookedLevels: Set<string> = new Set([
        'level0', 'level0_1', 'level0_2', 'level0_3',
        'level4', 'level5', 'level6', 'level7',
        'level10',   // ToM
        'level11',   // special visibility map
    ]);
  3. Add an <option> for it in the level dropdown in App.jsx so it shows up in the UI:

    <option value="yourlevel">level12 - my level</option>

    (Most levels in this branch are commented out — only level_demo and level6 are exposed.)


Training on Compute Canada (SLURM + Apptainer)

Training runs headless: the Python RL backend, the node game server, and a headless client (test-backend.js) are all started inside one SLURM job and wired together over dynamically allocated ports.

Venv setup

module load python/3.12
module load gcc opencv/4.11.0

virtualenv env
source env/bin/activate
pip3 install -r backend/rl-multistew/requirements.txt

Build the Apptainer containers

module load apptainer/1.3.5
cd backend/node
ls   # should contain node.def and headles.def

apptainer build ../node.sif    node.def      # game server; rebuild after changes under src/
apptainer build ../<map>.sif   headles.def   # headless client

Build into the parent directory (../). node.def copies the whole of node/ into the image, so leaving a .sif inside backend/node/ makes every later build enormous.

In this branch the headless client hardcodes its level in test-backend.js:49 (level: 'level0'), so there is one client container per map: edit that line, then build the matching .sif (cramped.sif, forced.sif, ring.sif, counter.sif, headless_asym.sif). Rebuild the client only when test-backend.js changes — rebuilding node.sif is the common case.

The two scripts

  1. CC/headless_0407.sh — the sbatch script. It allocates two free ports (node server + RL FastAPI) so several jobs can share one node, boots server.py, node.sif and <map>.sif, and runs one training job. You normally don't edit this — you pass it positional arguments. It always adds --normalize_advantages and --log, and injects --rl_port.
  2. CC/train_CNN_agentV2.sh — the launcher. It defines the sweep (seeds, region sizes, maps, learning rates) and calls sbatch headless_0407.sh <args...> once per configuration. This is the file you edit to configure a run.

Positional arguments to headless_0407.sh

The order is fixed; args 1–28 map onto server.py flags, arg 29 names the client container.

# server.py flag Meaning Example
1 --max-steps Total training steps 15000000
2 --num_agents Number of agents 2
3 --num_envs Parallel environments 1
4 --num_steps Rollout length per env 256
5 --seed Random seed 1
6 --layout Overcooked env name (see map table) overcooked_counter_circuit_v0
7 --num_mini_batches PPO mini-batches 4
8 --ppo_epochs PPO epochs 5
9 --clip_param PPO clip 0.2
10 --value_loss_coef Value loss coefficient 0.5
11 --entropy_coef Entropy coefficient 0.01
12 --max_grad_norm Gradient clipping 0.5
13 --gamma Discount factor 0.99
14 --lam GAE lambda 0.95
15 --data_path Folder for logged training data data_test
16 --feature Observation type — selects the agent image_v3 / global / local
17 --region_size Visibility region (only used by local) 4 (local) / 100 (global)
18 --checkpoint_dir Folder where checkpoints are saved models_test
19 --agents_RS (a) Per-agent region size, agent 0 100
20 --agents_RS (b) Per-agent region size, agent 1 100
21 --model_type Network architecture LayerNormCNNAgent / LinearAgent
22 --hidden_dim Hidden layer width 256
23 --num_hidden_layers Hidden layers 1 (CNN) / 2 (Linear)
24 --grid_width Grid width — CNN only 10
25 --grid_height Grid height — CNN only 7
26 --frame_stack Frames stacked — CNN only 1
27 --tile_size Pixels per tile — CNN only 16
28 --lr Learning rate 1e-4
29 (last arg) Client container — runs ../backend/<arg29>.sif counter

How the agent is selected

By the combination of --feature (16) and --model_type (21):

  • ImageAgent (CNN) — --feature image_v3 (image / image_v2 are older variants) plus --model_type LayerNormCNNAgent (or CNNAgent). Uses args 24–27.
  • LinearAgent (MLP) — --feature global (whole kitchen) or --feature local (local window; set arg 17) plus --model_type LinearAgent. Args 24–27 are ignored.

Container name → level → env name

Arg 29 must be paired with the matching --layout (arg 6), since --layout is what names checkpoints and CSVs on the RL side.

Container (arg 29) level in test-backend.js --layout (arg 6)
cramped level0 overcooked_cramped_room_v0
forced level4 overcooked_forced_coordination_v0
ring level5 overcooked_coordination_ring_v0
counter level6 overcooked_counter_circuit_v0
headless_asym level7 overcooked_asymmetric_advantages_v0

Single job (sanity check)

Train an ImageAgent (CNN):

cd CC
sbatch headless_0407.sh 15000000 2 1 256 1 overcooked_counter_circuit_v0 \
  4 5 0.2 0.5 0.01 0.5 0.99 0.95 \
  data_test image_v3 4 models_test 100 100 LayerNormCNNAgent 256 1 10 7 1 16 1e-4 counter

Train a LinearAgent (MLP):

cd CC
sbatch headless_0407.sh 15000000 2 1 256 1 overcooked_counter_circuit_v0 \
  4 5 0.2 0.5 0.01 0.5 0.99 0.95 \
  data_test_linear global 100 models_test_linear 100 100 LinearAgent 256 2 10 7 1 16 1e-4 counter

For a local LinearAgent, set arg 16 to local and arg 17 to the window size (e.g. 4).

Full sweep

Edit the arrays and shared settings in train_CNN_agentV2.sh, then:

cd CC
./train_CNN_agentV2.sh

Checkpoint and CSV filenames embed feature, model_type, seed, layout, and region_size (see save_model_checkpoint in rl_multistew/utils.py), so different agents never collide even in a shared folder — separate data_* / models_* folders are for your own sanity, not correctness.

Because headless_0407.sh allocates its own ports, submit_job just calls sbatch directly — the old --exclude=$NODES trick from param_tune.sh is no longer needed.


Docker housekeeping

docker compose up front                 # one service at a time
docker compose build <service_name>     # rebuild one service
docker compose logs -f [service_name]   # follow logs
docker compose down                     # stop everything
docker compose down --rmi all --volumes --remove-orphans   # full cleanup

Podman equivalents live in podman-compose.yml, run-with-podman.sh, and cleanup-podman.sh.

Troubleshooting

node_modules issues:

cd front/         && rm -rf node_modules && npm install
cd backend/node/  && rm -rf node_modules && npm install

Hot reloading not working — make sure CHOKIDAR_USEPOLLING=true is set, then:

docker compose restart front

The game server can't reach the RL server — you are almost certainly running Option B with RL_SERVICE_URL still pointing at http://rl-multistew:8800. See step 1 of Option B.

Further reading

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages