Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

llm-queue

A high-performance HTTP request proxy that sits in front of a number of llama.cpp servers, acting as a queue, multiplexor, and router.

Multiple clients, already configured to use llama.cpp LLM models via HTTP, are pointed at llm-queue instead, and llm-queue forwards their requests to the actual llama.cpp servers.

It can act as a simple queue — buffering incoming requests so they reach a single llama.cpp server one at a time, in arrival order, while guarding that requests complete. It can also act as a router, with per-client configuration deciding which llama.cpp server each request goes to: all clients use the same connection info (IP, port), and llm-queue routes by user agent, origin IP address, an HTTP header, or a query parameter. This makes it possible to send, say, all requests from a specific client to a llama-server running on a weaker NVIDIA Jetson Nano with a Qwen model optimized for one-shot request/response (--cache-ram 0).

The primary focus is performance and durability: fast llama.cpp request/response pass-through (including SSE streaming), and strong guarantees that llama.cpp servers are protected from overload — especially when llama-server runs with --threads=1 or --parallel=1.

How it works

client ──▶ llm-queue ──▶ rule matcher ──▶ backend queue (FIFO, bounded) ──▶ llama-server
                                           503 if full     504 on wait timeout
  • Each backend has a concurrency limit (max_concurrent, default 1 — match your --parallel) and a bounded FIFO queue (queue_capacity).
  • A backend slot is held until the response is fully streamed to the client (or the client disconnects), so a --parallel=1 llama-server never sees a second concurrent request.
  • When the queue is full, requests are rejected immediately with 503 + Retry-After instead of piling up; requests that wait longer than queue_wait_timeout_secs get 504.
  • Request bodies are buffered (bounded by max_body_bytes, 413 above it), which allows safe retry on connection failures. Response bodies are never buffered — pure streaming pass-through, so stream: true works.
  • Routing rules are evaluated in order, first match wins; all matchers within a rule must match. Unmatched requests go to default_backend.

Usage

cargo build --release
./target/release/llm-queue --config llm-queue.toml

See llm-queue.example.toml for a fully commented configuration example. CLI flags: --config <path> (default llm-queue.toml), --listen <addr> (overrides the config), --log-level <level> (default info; RUST_LOG also works).

Endpoints

Every path and method is proxied transparently to the routed backend (/completion, /v1/chat/completions, /health, …), except:

  • GET /llm-queue/status — JSON queue state per backend: in-flight, waiting, totals.

Error responses

llm-queue generates JSON errors in llama.cpp's {"error": {"code", "message"}} shape:

Status Meaning
503 Backend queue full (Retry-After header set)
504 Timed out waiting for a queue slot
502 Backend unreachable (after connect retries)
413 Request body larger than max_body_bytes

Development

cargo test                     # unit + integration tests
cargo clippy -- -D warnings
cargo fmt

About

A high-performance HTTP request proxy that sits in front of a number of llama.cpp servers, acting as a queue, multiplexor, and router.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages