Can an LLM actually play the game?
Natural language → movement script → executed in a real Unity simulator → scored.
25 models, 342 runs, one scale.
English · 한국어
An LLM receives a natural-language instruction, writes a movement script, and the script is executed in a headless Unity simulator. Every model is scored on the same maps with the same instructions, so the comparison is apples to apples.
| Models | 25 — Gemini, Claude (Bedrock), Nova, GPT / o-series (Azure), Grok, DeepSeek |
| Runs | 342 — 288 multi-agent + 54 single agent |
| MAS grid | 3 trials × 3 instructions × 3 maps × 32 model combinations |
| Scale | about 1.72M tokens, roughly 8 hours per full sweep |
| Pipeline | Plan → Validate → Script (multi-agent), then executed in Unity |
User (Unity WebGL / Nginx)
↓
api-server (FastAPI, 8000) ← main server: LLM calls, episode management, scoring
↓ ↓
SoonSoon API test-server (Unity build, 8001)
(LLM proxy) ← game simulator: runs the script and returns the result
↓
Gemini / AWS Bedrock / Azure OpenAI
| Service | Port | Description |
|---|---|---|
| api-server | 8000 | Main API server (episodes, LLM orchestration) |
| test-server | 8001 | Unity game simulator |
| postgres-server | 5432 | PostgreSQL database |
| nginx | 80 | Reverse proxy + WebGL static files |
llm-Game-interface-Benchmark/
├── api-server/ # Main API server (FastAPI)
│ ├── episode/ # Episode generation, benchmarking, scoring
│ ├── llm/ # LLM calls, MAS (Multi-Agent System)
│ ├── rag/ # RAG service (experimental)
│ ├── common/ # DB setup, shared schemas
│ └── alembic/ # DB migrations
├── test-server/ # Unity game simulator (build binary)
├── episode-generator/ # CLI tool for bulk episode generation
├── nginx/ # Nginx config + Unity WebGL build
├── docs/ # Design documents
└── .github/workflows/ # CI/CD (GitHub Actions → Azure VM)
- Docker & Docker Compose
- A
.envfile (see below)
cp .env.example .envSOONSOON_API_KEY=<LLM proxy API key>
SOONSOON_API_ADDRESS=<LLM proxy address>Ask the project maintainer for
SOONSOON_API_KEYandSOONSOON_API_ADDRESS.
# build and start everything
docker-compose up --build
# detached
docker-compose up --build -d
# stop
docker-compose down| URL | Description |
|---|---|
| http://localhost | Unity WebGL game (Nginx) |
| http://localhost:8000/docs | API server Swagger |
| http://localhost:8001/docs | Test server Swagger |
To work on api-server without Docker:
cd api-server
pip install -r requirements.txt
# DB migration
alembic upgrade head
# run
uvicorn main:app --reload --port 8000PostgreSQL must be running locally. Set
DATABASE_URLin.env:DATABASE_URL=postgresql://user:password@localhost:5432/fastapi_db
# single-agent benchmark
POST /episodes/auto-benchmark-single
# MAS (Multi-Agent System) benchmark
POST /episodes/auto-benchmark# single-agent game script generation
POST /llm/game-script
# MAS game script generation (Plan → Validate → Script)
POST /llm/mas-gen
# list available models
GET /llm/modelsFull API reference is at http://localhost:8000/docs once the server is running.
cd episode-generator
# generate episodes from prompts.csv
python episode-generator.py --action generate \
--llm-provider gemini-text \
--llm-model gemini-2.0-flash \
--prompts-file prompts.csv \
--output-file episodes.csv
# fetch existing episodes into CSV
python episode-generator.py --action getpython reset_db.py| Area | Tech |
|---|---|
| Backend | FastAPI, Python, SQLAlchemy, Alembic |
| Database | PostgreSQL |
| LLM orchestration | LangGraph (MAS workflow) |
| Game simulator | Unity (Linux headless build) |
| Proxy | Nginx |
| Container | Docker, Docker Compose |
| CI/CD | GitHub Actions → Azure VM |