This is Truong's human-AI teaming branch. This repo is static and out of date from the original multistew repository. The purpose of this repository is to preserve code — it is kept for display and reference, not for active development.
The Colyseus game server (backend/node/) and the React frontend (front/) were built in
large part by Azadeh Mostaghel (@amostaghel) and
Seti Rahmani (@Seti17) of Aggregate Intellect — together they
account for the majority of the commits behind those two services. This branch layers the RL and
human-AI teaming research on top of it.
The project was supervised by Amir Feizpour, CEO of Aggregate Intellect.
React (front) <--> Colyseus game server (node) <--/predict_all--> RL server (rl-multistew)
:5173 :2567 :8800
- Frontend — React + Vite (
front/) - Game server — Colyseus / Node + TypeScript (
backend/node/), owns the kitchen state and levels - RL server — Python + FastAPI (
backend/rl-multistew/), receives the full game state onPOST /predict_alland returns one action per agent - Sherpa / LLM service — Python + FastAPI (
backend/llm_service/) - Additional services:
backend/rl-vt-collab/(:8008),backend/human-eval/(:8808)
There are two ways to bring the stack up. Option B is recommended: the RL server pulls in big packages (PyTorch, OpenCV), so you do not want Docker building it every time.
cp .env.example .env # then fill in the variables you need
docker compose up --buildEndpoints:
- Frontend — http://localhost:5173
- Colyseus server — ws://localhost:2567
- RL server — http://localhost:8800
- Sherpa service — http://localhost:8000
Production variant:
export NODE_ENV=production
docker compose -f docker-compose.yml -f docker-compose.prod.yml up --build
# frontend on http://localhost:801. Point the game server at the host's RL server.
In docker-compose.yml, the node service sets
RL_SERVICE_URL=http://rl-multistew:8800, which only resolves inside the compose network.
Since the RL server now runs outside Docker, change it to:
- RL_SERVICE_URL=http://172.17.0.1:8800 # 172.17.0.1 maps to the host's localhost (Linux)
- RL_SERVICE_URL=http://host.docker.internal:8800 # use this on macOS2. Build and start only the frontend and game server.
./runnode.sh
# which is just: sudo docker compose up --build --no-deps front node--no-deps is what skips rl-multistew, sherpa, and the other dependencies declared in
docker-compose.yml.
3. Set up a virtualenv and run the RL server as a plain Python script.
python3 -m venv env # python >= 3.10
source env/bin/activate
pip3 install -r backend/rl-multistew/requirements.txt
python3 backend/rl-multistew/server.py -h # lists every command-line argumentRun with image observations (CNN agent). This is the checkpoint shipped in
backend/rl-multistew/CNN_models/ — the flags below are the exact ones it was trained with,
so they must match or the weights will not load:
python3 backend/rl-multistew/server.py \
--endpoint inference \
--model_path "backend/rl-multistew/CNN_models/2_agents_overcooked_counter_circuit_v0_seed_2_image_RS_4_agents_vision_[100, 100]_modelT_CNNAgent_hiddenD_256_numH_1_FS_1_TS_16.pth" \
--feature image \
--model_type CNNAgent \
--layout overcooked_counter_circuit_v0 \
--num_agents 2 \
--num_envs 1 \
--num_steps 256 \
--seed 2 \
--region_size 4 \
--agents_RS 100 100 \
--hidden_dim 256 \
--num_hidden_layers 1 \
--grid_width 10 \
--grid_height 7 \
--frame_stack 1 \
--tile_size 16 \
--lr 1e-4 \
--num_mini_batches 4 \
--ppo_epochs 5 \
--clip_param 0.2 \
--value_loss_coef 0.5 \
--entropy_coef 0.01 \
--max_grad_norm 0.5 \
--gamma 0.99 \
--lam 0.95 \
--normalize_advantages \
--data_path path-to-data-folderQuote the --model_path — the filename contains spaces and square brackets.
The checkpoint's filename encodes how it was trained (see save_model_checkpoint in
rl_multistew/utils.py): 2 agents, overcooked_counter_circuit_v0, seed 2, image feature,
region size 4, CNNAgent with hidden dim 256 / 1 hidden layer, frame stack 1, tile size 16.
Because its layout is counter circuit, pick Level 6 - Counter Circuit in the frontend's level
dropdown — it is one of the two levels left uncommented in App.jsx.
To train instead of serve, swap --endpoint inference for --endpoint ppo and add
--max-steps, --log, and --checkpoint_dir.
Run with handcrafted feature observations (linear agent):
python3 backend/rl-multistew/server.py \
--max-steps 10000 --num_agents 2 --num_envs 1 --num_steps 256 --seed 1 \
--layout overcooked_counter_circuit_v0 --num_mini_batches 4 --ppo_epochs 5 --clip_param 0.2 \
--value_loss_coef 0.5 --entropy_coef 0.01 --max_grad_norm 0.5 --gamma 0.99 --lam 0.95 \
--data_path path-to-data-folder \
--feature global \
--model_path path-to-model \
--region_size 100 \
--endpoint inference \
--model_type DefaultNotes:
--model_pathis the checkpoint loaded in inference mode.--endpointselects the request handler:inference,ppo,saliency, orhightlights.--rl_port(default8800) changes the FastAPI port.
4. In the browser, always tick "Synchronize Agents" in the Game Setup panel before starting a game. Without it the game server and the RL server step out of sync.
Defined in command_line_parser.py.
--feature (observation type) — global, local, localv2,
minimal_spatial_other_agent_aware, region_size_invariant, image, image_v2, image_v3.
--model_type (network) — Default, LinearAgent, CNNAgent, LayerNormCNNAgent.
Any model_type containing CNNAgent activates the CNN path and uses --grid_width,
--grid_height, --frame_stack, --tile_size; those four are ignored otherwise.
--region_size is only meaningful for the local features; 100 effectively means "unused".
The game server sends its internal state to POST /predict_all, caught by predict_all(game_state)
in server.py. That function handles one step of interaction and
branches on the --endpoint flag:
@app.post("/predict_all", response_model=ActionResponse)
def predict_all(game_state: FullGameState):
if mappo.args.endpoint == "hightlights":
...
elif mappo.args.endpoint == "saliency":
...
elif mappo.args.endpoint == "inference":
...
else:
assert mappo.args.endpoint == "ppo", f"Invalid endpoint specified: {mappo.args.endpoint}"To add your own: create backend/rl-multistew/endpoints/myendpoint.py alongside the existing
endpoints/, then add a branch for it in that if chain.
Outdated. This is the procedure as it stood in this archived branch, and it is what the code here still expects — registering a level by hand in three separate
overcookedLevelssets is exactly the kind of thing that has since changed upstream. Follow it to understand or run this snapshot; do not assume it still applies to the active multistew repository.
-
Create the level file in backend/node/src/levels/, e.g.
levels/yourlevel.txt. Follow an existing level for inspiration — the file specifies how many AI and human players can join. -
Add the
yourlevelstring to theovercookedLevelsset in all three of:private overcookedLevels: Set<string> = new Set([ 'level0', 'level0_1', 'level0_2', 'level0_3', 'level4', 'level5', 'level6', 'level7', 'level10', // ToM 'level11', // special visibility map ]);
-
Add an
<option>for it in the level dropdown in App.jsx so it shows up in the UI:<option value="yourlevel">level12 - my level</option>
(Most levels in this branch are commented out — only
level_demoandlevel6are exposed.)
Training runs headless: the Python RL backend, the node game server, and a headless client
(test-backend.js) are all started inside one SLURM job and wired together over dynamically
allocated ports.
module load python/3.12
module load gcc opencv/4.11.0
virtualenv env
source env/bin/activate
pip3 install -r backend/rl-multistew/requirements.txtmodule load apptainer/1.3.5
cd backend/node
ls # should contain node.def and headles.def
apptainer build ../node.sif node.def # game server; rebuild after changes under src/
apptainer build ../<map>.sif headles.def # headless clientBuild into the parent directory (../). node.def copies the whole of node/ into the
image, so leaving a .sif inside backend/node/ makes every later build enormous.
In this branch the headless client hardcodes its level in
test-backend.js:49 (level: 'level0'), so there is one client
container per map: edit that line, then build the matching .sif (cramped.sif, forced.sif,
ring.sif, counter.sif, headless_asym.sif). Rebuild the client only when test-backend.js
changes — rebuilding node.sif is the common case.
- CC/headless_0407.sh — the
sbatchscript. It allocates two free ports (node server + RL FastAPI) so several jobs can share one node, bootsserver.py,node.sifand<map>.sif, and runs one training job. You normally don't edit this — you pass it positional arguments. It always adds--normalize_advantagesand--log, and injects--rl_port. - CC/train_CNN_agentV2.sh — the launcher. It defines the sweep
(seeds, region sizes, maps, learning rates) and calls
sbatch headless_0407.sh <args...>once per configuration. This is the file you edit to configure a run.
The order is fixed; args 1–28 map onto server.py flags, arg 29 names the client container.
| # | server.py flag |
Meaning | Example |
|---|---|---|---|
| 1 | --max-steps |
Total training steps | 15000000 |
| 2 | --num_agents |
Number of agents | 2 |
| 3 | --num_envs |
Parallel environments | 1 |
| 4 | --num_steps |
Rollout length per env | 256 |
| 5 | --seed |
Random seed | 1 |
| 6 | --layout |
Overcooked env name (see map table) | overcooked_counter_circuit_v0 |
| 7 | --num_mini_batches |
PPO mini-batches | 4 |
| 8 | --ppo_epochs |
PPO epochs | 5 |
| 9 | --clip_param |
PPO clip | 0.2 |
| 10 | --value_loss_coef |
Value loss coefficient | 0.5 |
| 11 | --entropy_coef |
Entropy coefficient | 0.01 |
| 12 | --max_grad_norm |
Gradient clipping | 0.5 |
| 13 | --gamma |
Discount factor | 0.99 |
| 14 | --lam |
GAE lambda | 0.95 |
| 15 | --data_path |
Folder for logged training data | data_test |
| 16 | --feature |
Observation type — selects the agent | image_v3 / global / local |
| 17 | --region_size |
Visibility region (only used by local) |
4 (local) / 100 (global) |
| 18 | --checkpoint_dir |
Folder where checkpoints are saved | models_test |
| 19 | --agents_RS (a) |
Per-agent region size, agent 0 | 100 |
| 20 | --agents_RS (b) |
Per-agent region size, agent 1 | 100 |
| 21 | --model_type |
Network architecture | LayerNormCNNAgent / LinearAgent |
| 22 | --hidden_dim |
Hidden layer width | 256 |
| 23 | --num_hidden_layers |
Hidden layers | 1 (CNN) / 2 (Linear) |
| 24 | --grid_width |
Grid width — CNN only | 10 |
| 25 | --grid_height |
Grid height — CNN only | 7 |
| 26 | --frame_stack |
Frames stacked — CNN only | 1 |
| 27 | --tile_size |
Pixels per tile — CNN only | 16 |
| 28 | --lr |
Learning rate | 1e-4 |
| 29 | (last arg) | Client container — runs ../backend/<arg29>.sif |
counter |
By the combination of --feature (16) and --model_type (21):
- ImageAgent (CNN) —
--feature image_v3(image/image_v2are older variants) plus--model_type LayerNormCNNAgent(orCNNAgent). Uses args 24–27. - LinearAgent (MLP) —
--feature global(whole kitchen) or--feature local(local window; set arg 17) plus--model_type LinearAgent. Args 24–27 are ignored.
Arg 29 must be paired with the matching --layout (arg 6), since --layout is what names
checkpoints and CSVs on the RL side.
| Container (arg 29) | level in test-backend.js |
--layout (arg 6) |
|---|---|---|
cramped |
level0 |
overcooked_cramped_room_v0 |
forced |
level4 |
overcooked_forced_coordination_v0 |
ring |
level5 |
overcooked_coordination_ring_v0 |
counter |
level6 |
overcooked_counter_circuit_v0 |
headless_asym |
level7 |
overcooked_asymmetric_advantages_v0 |
Train an ImageAgent (CNN):
cd CC
sbatch headless_0407.sh 15000000 2 1 256 1 overcooked_counter_circuit_v0 \
4 5 0.2 0.5 0.01 0.5 0.99 0.95 \
data_test image_v3 4 models_test 100 100 LayerNormCNNAgent 256 1 10 7 1 16 1e-4 counterTrain a LinearAgent (MLP):
cd CC
sbatch headless_0407.sh 15000000 2 1 256 1 overcooked_counter_circuit_v0 \
4 5 0.2 0.5 0.01 0.5 0.99 0.95 \
data_test_linear global 100 models_test_linear 100 100 LinearAgent 256 2 10 7 1 16 1e-4 counterFor a local LinearAgent, set arg 16 to local and arg 17 to the window size (e.g. 4).
Edit the arrays and shared settings in train_CNN_agentV2.sh, then:
cd CC
./train_CNN_agentV2.shCheckpoint and CSV filenames embed feature, model_type, seed, layout, and region_size
(see save_model_checkpoint in rl_multistew/utils.py), so different agents never collide even in
a shared folder — separate data_* / models_* folders are for your own sanity, not correctness.
Because headless_0407.sh allocates its own ports, submit_job just calls sbatch directly — the
old --exclude=$NODES trick from param_tune.sh is no longer needed.
docker compose up front # one service at a time
docker compose build <service_name> # rebuild one service
docker compose logs -f [service_name] # follow logs
docker compose down # stop everything
docker compose down --rmi all --volumes --remove-orphans # full cleanupPodman equivalents live in podman-compose.yml, run-with-podman.sh, and cleanup-podman.sh.
node_modules issues:
cd front/ && rm -rf node_modules && npm install
cd backend/node/ && rm -rf node_modules && npm installHot reloading not working — make sure CHOKIDAR_USEPOLLING=true is set, then:
docker compose restart frontThe game server can't reach the RL server — you are almost certainly running Option B with
RL_SERVICE_URL still pointing at http://rl-multistew:8800. See step 1 of Option B.
- ROLE_SYSTEM_README.md — the role system
- backend/rl-multistew/README.md — RL service API details
- Multistew-setup-locally.pdf — the original local-setup walkthrough