Skip to content

Repository files navigation

LLM Game Interface & Benchmark

Can an LLM actually play the game?
Natural language → movement script → executed in a real Unity simulator → scored.
25 models, 342 runs, one scale.

Python FastAPI LangGraph PostgreSQL Unity Docker Nginx GitHub Actions

English · 한국어


Overview

An LLM receives a natural-language instruction, writes a movement script, and the script is executed in a headless Unity simulator. Every model is scored on the same maps with the same instructions, so the comparison is apples to apples.

Models 25 — Gemini, Claude (Bedrock), Nova, GPT / o-series (Azure), Grok, DeepSeek
Runs 342 — 288 multi-agent + 54 single agent
MAS grid 3 trials × 3 instructions × 3 maps × 32 model combinations
Scale about 1.72M tokens, roughly 8 hours per full sweep
Pipeline Plan → Validate → Script (multi-agent), then executed in Unity

Screenshots

Benchmark results Benchmark simulator

Architecture

User (Unity WebGL / Nginx)
    ↓
api-server (FastAPI, 8000)   ← main server: LLM calls, episode management, scoring
    ↓                ↓
SoonSoon API      test-server (Unity build, 8001)
(LLM proxy)        ← game simulator: runs the script and returns the result
    ↓
Gemini / AWS Bedrock / Azure OpenAI

Services

Service Port Description
api-server 8000 Main API server (episodes, LLM orchestration)
test-server 8001 Unity game simulator
postgres-server 5432 PostgreSQL database
nginx 80 Reverse proxy + WebGL static files

Layout

llm-Game-interface-Benchmark/
├── api-server/             # Main API server (FastAPI)
│   ├── episode/            # Episode generation, benchmarking, scoring
│   ├── llm/                # LLM calls, MAS (Multi-Agent System)
│   ├── rag/                # RAG service (experimental)
│   ├── common/             # DB setup, shared schemas
│   └── alembic/            # DB migrations
├── test-server/            # Unity game simulator (build binary)
├── episode-generator/      # CLI tool for bulk episode generation
├── nginx/                  # Nginx config + Unity WebGL build
├── docs/                   # Design documents
└── .github/workflows/      # CI/CD (GitHub Actions → Azure VM)

Getting started

Requirements

  • Docker & Docker Compose
  • A .env file (see below)

Environment

cp .env.example .env
SOONSOON_API_KEY=<LLM proxy API key>
SOONSOON_API_ADDRESS=<LLM proxy address>

Ask the project maintainer for SOONSOON_API_KEY and SOONSOON_API_ADDRESS.

Run

# build and start everything
docker-compose up --build

# detached
docker-compose up --build -d

# stop
docker-compose down

Endpoints

URL Description
http://localhost Unity WebGL game (Nginx)
http://localhost:8000/docs API server Swagger
http://localhost:8001/docs Test server Swagger

Local development

To work on api-server without Docker:

cd api-server
pip install -r requirements.txt

# DB migration
alembic upgrade head

# run
uvicorn main:app --reload --port 8000

PostgreSQL must be running locally. Set DATABASE_URL in .env:

DATABASE_URL=postgresql://user:password@localhost:5432/fastapi_db

Main API

Benchmark

# single-agent benchmark
POST /episodes/auto-benchmark-single

# MAS (Multi-Agent System) benchmark
POST /episodes/auto-benchmark

LLM

# single-agent game script generation
POST /llm/game-script

# MAS game script generation (Plan → Validate → Script)
POST /llm/mas-gen

# list available models
GET /llm/models

Full API reference is at http://localhost:8000/docs once the server is running.


Bulk episode generation (CLI)

cd episode-generator

# generate episodes from prompts.csv
python episode-generator.py --action generate \
  --llm-provider gemini-text \
  --llm-model gemini-2.0-flash \
  --prompts-file prompts.csv \
  --output-file episodes.csv

# fetch existing episodes into CSV
python episode-generator.py --action get

Reset the database

python reset_db.py

Stack

Area Tech
Backend FastAPI, Python, SQLAlchemy, Alembic
Database PostgreSQL
LLM orchestration LangGraph (MAS workflow)
Game simulator Unity (Linux headless build)
Proxy Nginx
Container Docker, Docker Compose
CI/CD GitHub Actions → Azure VM

About

A benchmark harness that turns natural language into game scripts, runs them in a real Unity simulator, and ranks 25 LLMs on the same scale.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages