A high-performance HTTP request proxy that sits in front of a number of llama.cpp servers, acting as a queue, multiplexor, and router.
Multiple clients, already configured to use llama.cpp LLM models via HTTP, are pointed at llm-queue instead, and llm-queue forwards their requests to the actual llama.cpp servers.
It can act as a simple queue — buffering incoming requests so they reach a single llama.cpp server one at a time, in arrival order, while guarding that requests complete. It can also act as a router, with per-client configuration deciding which llama.cpp server each request goes to: all clients use the same connection info (IP, port), and llm-queue routes by user agent, origin IP address, an HTTP header, or a query parameter. This makes it possible to send, say, all requests from a specific client to a llama-server running on a weaker NVIDIA Jetson Nano with a Qwen model optimized for one-shot request/response (--cache-ram 0).
The primary focus is performance and durability: fast llama.cpp request/response pass-through (including SSE streaming), and strong guarantees that llama.cpp servers are protected from overload — especially when llama-server runs with --threads=1 or --parallel=1.
client ──▶ llm-queue ──▶ rule matcher ──▶ backend queue (FIFO, bounded) ──▶ llama-server
503 if full 504 on wait timeout
- Each backend has a concurrency limit (
max_concurrent, default 1 — match your--parallel) and a bounded FIFO queue (queue_capacity). - A backend slot is held until the response is fully streamed to the client (or the client disconnects), so a
--parallel=1llama-server never sees a second concurrent request. - When the queue is full, requests are rejected immediately with
503+Retry-Afterinstead of piling up; requests that wait longer thanqueue_wait_timeout_secsget504. - Request bodies are buffered (bounded by
max_body_bytes, 413 above it), which allows safe retry on connection failures. Response bodies are never buffered — pure streaming pass-through, sostream: trueworks. - Routing rules are evaluated in order, first match wins; all matchers within a rule must match. Unmatched requests go to
default_backend.
cargo build --release
./target/release/llm-queue --config llm-queue.tomlSee llm-queue.example.toml for a fully commented configuration example. CLI flags: --config <path> (default llm-queue.toml), --listen <addr> (overrides the config), --log-level <level> (default info; RUST_LOG also works).
Every path and method is proxied transparently to the routed backend (/completion, /v1/chat/completions, /health, …), except:
GET /llm-queue/status— JSON queue state per backend: in-flight, waiting, totals.
llm-queue generates JSON errors in llama.cpp's {"error": {"code", "message"}} shape:
| Status | Meaning |
|---|---|
| 503 | Backend queue full (Retry-After header set) |
| 504 | Timed out waiting for a queue slot |
| 502 | Backend unreachable (after connect retries) |
| 413 | Request body larger than max_body_bytes |
cargo test # unit + integration tests
cargo clippy -- -D warnings
cargo fmt