Skip to content

Latest commit

 

History

History
687 lines (641 loc) · 48.7 KB

File metadata and controls

687 lines (641 loc) · 48.7 KB

Server & Web UI

dw.serve runs the workflow engine as a local HTTP server with a full web UI: browse and run workflows, build them in a form-based editor with introspection-driven autocomplete, watch jobs stream live progress, review past generations in a gallery, and manage the models on disk.

python -m dw.serve                       # http://127.0.0.1:8765
python -m dw.serve --port 8000 --workflow-dir ./workflows --output-dir ./outputs --prompt-dir ./prompts

# or point it at a workspace, which supplies all four directories
python -m dw.serve --workspace ~/studio

# your own workflows, with a checkout's examples alongside them read-only
python -m dw.serve --workspace ~/studio --examples-dir ~/src/diffusers-workflow/workflows
python -m dw.serve --host 0.0.0.0 --token "some-long-random-string"   # reachable off this machine
python -m dw.serve --host 0.0.0.0 --token "..." --mcp   # ...and drivable by an agent on another machine
python -m dw.serve --trust-workflows      # only if nothing untrusted can reach POST /api/jobs - see Security model

Installed as a package, the same server is dw-serve. Interactive API docs (OpenAPI) are at /docs.

The server keeps a persistent GPU worker underneath: models stay loaded between runs, so re-running a workflow with a new prompt skips the load entirely.

The pages

The UI is organised by workspace. A sidebar lists every workspace on the server; the selected one opens into Overview, Workflows, Jobs, Gallery, Assets, Editor, and the hash carries the workspace (#/ws/studio/gallery), so a link names where it points and an old #/gallery bookmark redirects to the last workspace you were in. Below the workspaces, Shared holds what every workspace sees — the prompt library, the common asset library and the read-only example workflows — and Server holds Models, Schema and Status.

  • Overview — a glance at the workspace: its recent outputs, its recent jobs, its workflows (ones with a proof first, plus a filled New workflow button), its asset count, and disk usage. Each panel loads and fails independently, so a slow gallery does not hold the jobs list back and a panel's own error shows in place of its content rather than reading as an empty workspace. Manage on this page is where a workspace is deleted (disabled for default); a workspace is created from the sidebar.
  • Workflows — every workflow on the workspace's own search path (--workflow-dir), as cards with descriptions, output kinds, and variable counts; the read-only ones from an --examples-dir are under Shared → Examples instead, and opening one's Editor there saves a copy into the current workspace. Folders one level deep become sections. Click through to a run form generated from the workflow's variables, with the raw JSON alongside.
  • Prompts — under Shared: the prompt library under --prompt-dir (default: discovered the way a CLI run discovers it, then pinned for every job, so the page and prompt: resolution always agree), plus the read-only prompts/ beside each --examples-dir, so an example's prompt: references resolve: stored prompts as cards with descriptions, intended-model badges, and tags, foldered the same way workflows are. Each opens in an editor with form, split, and schema-aware JSON views, and an Enhance with AI panel that expands an idea into a full prompt with a local language model (a preset per target model family; the model runs as an ordinary queued job). A workflow argument written as prompt:name loads the stored text at run time, and deleting a prompt warns which workflows reference it.
  • Jobs — the queue and full run history (persisted in ~/.diffusers_helper/jobs.sqlite). A workspace's Jobs page lists its own; Server → Status lists every workspace's with a filter. A running job streams step-by-step progress, per-step denoising ticks, what each step is doing when it is not denoising (loading a model, decoding, saving), and its result files as they land. A run whose steps chose a result.subfolder shows its results under final/ and intermediate/ headings (or whatever the step named), the deliverable first; one that chose none shows them as before. Jobs can be cancelled mid-denoise and re-run with one click. A finished job's Export button gathers the run into the workspace's exports/ on the server (POST /api/jobs/{id}/export) and downloads it as one zip - the same bundle MCP's export_job makes; an export that already exists asks before it is replaced.
  • Editor — build or modify workflows without writing JSON by hand. Forms are generated from the live pipeline signatures (see introspection), references (variable: / previous_result:) autocomplete from the workflow itself, and three views — form, split, and raw JSON — edit the same definition. The split view puts the form beside the JSON with both sides editable; changes apply when a side loses focus. Validate, save, and run from the same screen; a valid verdict is followed by the run's plan - the step and list counts, the minutes from the workflow's cost block with its basis, how many steps the step cache would serve, and the weights this server would download first (describePlan, ui/src/lib/plan.ts, reading POST /api/validate's plan). A Monaco editor with the workflow JSON schema backs the JSON views. A fourth view, flow, renders the workflow's data-flow graph read-only: one box per step, arrows for each previous_result reference labeled with the argument it feeds, entry-point steps marked apart from steps that depend on earlier ones, and fan-in points - steps combining more than one upstream producer - flagged with the cartesian-product multiplier where it's known statically (e.g. a literal num_images_per_prompt on both producers). It's a diagram of the JSON, not a second way to edit it; clicking a step jumps to it in the form view.
  • Gallery — everything in the selected workspace's output directory, which the engine lays out as <workflow>/<run id>/. The folder filter groups a workflow's runs together rather than listing each run separately, and a subfolder pick beside the text filter - offered once any output landed in one - narrows the grid to final/, intermediate/ or whatever a step's result.subfolder named; each run directory also holds a manifest.json describing what produced it (see Workspaces). Images generated with embed_metadata carry their full workflow definition and seed; open as workflow loads that definition into the editor with the seed pinned, so any image can be reproduced or riffed on. Each tile carries a checkbox (shift-click extends a range, Select all takes whatever the filter leaves showing); a selection can be downloaded as one zip or deleted in bulk, which is how a directory that fills up over a few hundred runs gets cleared out. Anything that fails to delete stays selected. Keep as asset promotes one generated file into the workspace's asset library under a stable name, so a later workflow can reference it as asset:<name> instead of a run id that pruning would break.
  • Models — the Hugging Face hub cache: every cached repo with sizes, revisions, and last-used dates, plus free disk space. Download a repo by id with live progress (cancellable; partial files resume on retry), and delete to free disk (refused while a job is running; the next workflow that needs the model downloads it again). The page also shows the installed diffusers version (with its commit for a git install) and can upgrade it to GitHub HEAD - new model pipelines usually land there before a PyPI release. The idle worker restarts on success so the next job imports the new version; the upgrade is refused while a job runs.
  • Schema — the workflow JSON schema the running server validates against, as a browsable tree: the document root plus every definition, with types, required markers, defaults, enums, and descriptions. $ref labels jump to their definition; a filter narrows the list.
  • Server → Status — what this server is and how to reach it: device, version, bind address and LAN addresses, whether a token is required, whether /mcp is mounted (with the claude mcp add line to connect to it), and the directories in use — and, below, the queue across every workspace. Workspaces are created from the sidebar and deleted from their Overview.

Workspaces

One server can hold several workspaces — each with its own workflows/, assets/ and outputs/, all sharing the root's one prompt library. The root's own folders are the workspace named default.

Every scoped route takes an optional ?workspace=<name>; omitting it means default, so nothing written against a single-workspace server changes meaning. POST /api/jobs also accepts "workspace" in the body, and a job holds onto the directories it was submitted with — through the run, a rerun, and when history serves its files back. GET /api/jobs spans every workspace unless one is named.

A workspace is a namespace, not a security boundary: the API token is all-or-nothing. See Workspaces.

The sidebar lists every workspace, and + new there creates one; a workspace's own Overview page is where it is deleted. Server → Status shows the resolved directories and the claude mcp add line for connecting an agent from another machine, plus the queue across every workspace:

The Server page before the sidebar: address picker, generated claude mcp add line, resolved directories — the workspace list it shows now lives in the sidebar

Jobs API

Route What it does
POST /api/jobs Queue a run: {"workflow_path": ...} or an inline {"workflow": {...}, "base_dir": ...}, plus arguments for variable overrides. workflow_path accepts a stored workflow name as listed by /api/workflows (with or without .json, nested names included), or a relative/absolute path that still resolves under --workflow-dir - confined the same way the /api/workflows CRUD routes are; a path that names a real file outside that directory is rejected with 400, not opened. Answers with argument warnings from signature checking. Takes an optional acknowledged_cost: true is recorded as acknowledged: boolean; the object {fingerprint, minutes, downloads} from a validate answer's plan is bound - the server re-plans the run for the arguments given and answers 409 when the fingerprint differs or a repo in downloads_required is not in downloads (a download that has since vanished is not a refusal); the body is {"detail": {message, reason: "fingerprint" | "downloads" | "unplannable", acknowledged, plan}} with the current plan, so the caller re-quotes from it. minutes is recorded, never compared. Nothing is required: the web UI and every caller that sends nothing are acknowledged: none, and every job answer and history row carries acknowledged (and acknowledged_cost when bound). POST /api/jobs/{id}/rerun takes the same field and checks against the stored spec; a fresh seed does not change a fingerprint.
GET /api/jobs?workspace=&status=&limit= Queue + history summaries, oldest first, with total beside them. status narrows to one state or a comma-separated set (queued, running, succeeded, failed, cancelled; anything else is a 400); limit keeps the newest N, and total still reports how many matched, so a bounded answer cannot be mistaken for a complete one. No parameters means every job, which is what the web UI polls
GET /api/jobs/{id} Full detail: spec, events, manifest, error. A manifest entry for a step served from the step cache carries reused: true. Every entry carries subfolder - the in-run subfolder the step's result.subfolder chose, '' for none. A for_each step appears in the manifest as its members (shot@wide_open, shot@closeup), because the manifest records what ran; the run's workflow.json keeps the for_each form, because it records what was asked
GET /api/jobs/{id}/workflow The workflow the job ran: {id, definition, realized, seed_variable}. seed_variable names the variable a new_seed rerun would draw into (null when the workflow has none), read from the workflow as written rather than the realized copy, whose seed is pinned. realized: true is the copy the run itself wrote (workflow.json in its run directory), with arguments, seed, prompts and output:latest pinned; false falls back to the submitted definition, which is what a job from before run tracking has. 404 means neither is readable - the job itself still is. The equivalent MCP tool is get_job_workflow (see MCP.md)
POST /api/jobs/{id}/export?workspace=&overwrite= Gather one finished job into <workspace>/exports/<job id>/: workflow.json, manifest.json, job.json, README.md, assets/, inputs/, outputs/. 201 with the file list, total bytes, anything it could not find, a zip_url, and the three JSON files inline. 404 unknown job, 409 for a job still running or an existing export without overwrite
GET /exports/{id}.zip?workspace= The same tree as one archive, built on request rather than kept as a second copy. Entries are named <job id>/<relative path>. Ungated exactly as /outputs is
GET /api/jobs/{id}/events Server-sent events stream; ?after=N / Last-Event-ID replay missed events, so reconnects are lossless. Every event carries seq and at - seconds since the job started (since it was queued, for the events before that) - so a step's cost is a subtraction: step_start to generating is what a reused pipeline still pays before it runs, generating to the first pipeline_step is the prompt and reference encoding
GET /api/jobs/{id}/event-log?after=-1&limit=200 The same events as the SSE stream, as one JSON page: {id, status, events, last_seq, truncated, note}. after is exclusive; page by passing back the previous last_seq. A job restored from history serves the bounded event tail persisted with it; a job that finished before events were retained returns an empty list and a note saying so.
POST /api/jobs/{id}/cancel Cooperative cancel (takes effect at the next step boundary or denoise step)
POST /api/jobs/{id}/rerun Re-queue a finished job's spec. Body {"new_seed": true} draws a fresh seed into the workflow's seed variable instead of repeating the original arguments; 400 when the workflow pins its seed to a literal or names none. A plain rerun of a seeded workflow is served whole from the step cache — the earlier run's files, republished with reused: true, generating nothing
POST /api/jobs/{id}/move Reorder a queued job: {"direction": "up"|"down"|"front"|"back"}. Job listings carry each waiting job's queue_position.

One job runs at a time (it is one GPU); submissions queue in order, and the waiting portion of the queue can be reordered.

Progress events

Every event in the stream carries a seq and an event name:

event when payload
job_status queued/running/terminal transitions status
log worker output lines; each top-level block of a ModularPipeline as it starts (MiniMaxAI/MiniMax-H3: vae_encoder) - the lead-in before the denoise loop is where a reference encode's minutes go, and the block name is what says which one it is in; and each file the saving phase writes, named as it starts (writing shot.mp4 (121 frames)) and costed as it finishes (wrote shot.mp4 in 1.3s (1.4 MB)), which is the other stretch a step spends with its denoise counter frozen message, and for a file file plus seconds on the closing one
memory device memory after a run info
run_start the run directory is chosen, before the first step run_id, identity, run_dir
workflow_start the run begins workflow, total_steps, steps, seed
step_start / step_end each step step, index, total_steps; files and subfolder at the end. A step served from the step cache adds reused: true to step_end, and its files are the earlier run's files rather than newly written ones
iteration_start each argument combination in a step step, iteration, total_iterations
pipeline_step each denoise step step, total_steps. Emitted for a pipeline that takes a callback_on_step_end, and for a ModularPipeline (H3, LTX-2, Qwen-Image), which takes none - there the denoise block's own progress bar is what reports
phase the step changes what it is doing phase, detail
pipeline_released a step with release_pipeline drops its pipeline step, index, gpu_memory_allocated_mb and gpu_memory_allocated_before_mb (both null where the backend cannot say). Emitted between the step's generation and its files being written, which is where the release happens - so the ordering is readable off the event stream rather than by trying to poll memory through a sub-second write
warning a step finds something wrong with what it is about to write message, plus a kind and the figures behind it (level_spread: spread_db, measure, command; fps_mismatch: declared_fps, source_fps; audio_no_headroom: file, peak_dbfs - a deliverable at or above -0.5 dBFS, which an mp3 or AAC encode decodes over full scale; step_elided: step, overridden_by when a supplied argument is what made it unreferenced). Also appended to the job's warnings, prefixed with the step it fired in - the event keeps the moment, warnings keeps it where a caller polling the finished job will look, since a warning about the artifact outlives the run that noticed it
workflow_end the run finishes manifest

A step spends most of its wall clock outside the denoise loop, and pipeline_step cannot see any of it. phase is what fills that silence: loading (with the model or component in detail), cached (the same pipeline as a previous run - milliseconds, not minutes), generating (the denoise loop, or a chain's segment N/M - which is why the counter restarts), decoding (latents, after the last denoise step), saving (writing files, including video encode - it names each file on the log stream rather than running silent, since the denoise counter is frozen at its last step throughout) and task (a task step, named in detail). Emits are a handful per step, not per denoise tick.

Progress on a running job

GET /api/jobs/{id} carries a progress block while a job is running (null before it starts and once it is terminal, where the manifest is the better answer). It is the same information the event log holds, kept as the events arrive so a caller does not have to page back through a trimmed log to learn where a long render is:

field
step, step_index, total_steps the workflow step being run
phase, phase_detail the latest phase and what it named
seconds_in_phase how long it has been in it
seconds_since_event how long since anything at all happened - the number that separates a slow run from a stuck one
denoise_step, denoise_total_steps the denoise loop's counter, null until it starts

A null denoise_step under generating is the pipeline's lead-in - encoding the prompt and every reference - which emits nothing and runs well over a minute on a large video model. How long it runs follows what it has to encode: on MiniMax H3, ~90 s for a prompt with an image or audio reference, but ~10 min once a video reference is among them - a measured run encoding one 5 s 960x544 clip on an RTX 3090 sat silent from 94 s to 723 s. The keys are always present so that lead-in can be told from a loop that has stopped advancing: seconds_since_event is a stall signal once denoise_step is a number, or in any phase other than generating.

The lead-in is no longer silent, though: each of a modular pipeline's top-level blocks emits a log naming it as it starts (before_encode, text_encoder, vae_encoder, denoise, decode on H3), so the last event says which one the run is inside. Only a SequentialPipelineBlocks is narrated this way - a conditional container picks one branch rather than running them all, and walking its sub-blocks would be a wrong answer bought with a progress message.

It is a coarse one even then. A step's cost is not uniform when the pipeline configures a transformer block cache ("cache": {"type": "first_block"}): most steps are served from it in seconds and every few steps one is computed in full, so the same healthy run emits four pipeline_step events in 20 s and then nothing for 133 s. Liveness is denoise_step having moved between polls minutes apart, not silence measured against a fixed threshold - on H3 that threshold would have to exceed ~140 s to mean anything.

Introspection API

The editor's forms come from these; they are just as usable from scripts:

  • GET /api/pipelines, GET /api/pipelines/{name} — diffusers pipeline classes and their call signatures

  • GET /api/classes?kind=..., GET /api/classes/{name}?target=call|init|load — any allowed class (diffusers + registered extension modules), described for calling, constructing, or from_pretrained loading

  • GET /api/tasks — the task commands and processors

  • GET /api/tasks/{command} — a task's argument schema, read from its registered implementation's real signature

  • GET /api/schema — the workflow JSON schema. ?section= answers one part of it - steps, pipelines, tasks, result, variables or configuration - as {section, sections, elsewhere, schema}, where elsewhere names the section holding each definition the fragment still $refs; the no-argument call is the whole schema, unchanged

  • GET /api/guides — the documentation that bears on choosing a capability: each guide's name, what it covers, and its section headings

  • GET /api/guides/{name}?section= — one section of a guide; section names match loosely. Without a section the answer is the guide's index - its opening, its first section, and sections/withheld naming the rest - rather than the whole file, which for WORKFLOW_GUIDE.md is ~19.6k tokens in one call (#101). Served by the engine so an MCP client at another version reads the guides for the server it is driving, not its own. A checkout serves the repo's docs/; an install the copy build_dist.sh puts under dw/docs/

  • POST /api/validate — schema validation plus signature-level argument warnings for pipeline and task steps (catches the typo before the model loads); warnings also names an entry key of a list-driven variable that no step reads, at the entry's path (variables.shots[0] or, when the caller's own arguments supplied the list, arguments.shots[0]). Accepts workflow_path (same resolution and confinement as /api/jobs, above) as an alternative to inline workflow - exactly one of the two, or a 400. Every schema violation is returned in errors ([{path, message}], sorted by path, capped at 25), and joined one per line in error. It also takes the arguments a caller is about to run with, and checks them the way the run would: a name the workflow does not declare, a value that will not coerce to the declared type, and an asset:, prompt: or output: reference that names nothing this workspace can reach - each reported at arguments.<name>. POST /api/jobs makes the same check and answers 400 rather than queuing a job that would fail on its first step; checked_arguments on a valid answer names what was covered, since without arguments the verdict is about the stored defaults only. A reference set a pipeline would refuse is an error here too - too many images, videos or audio clips for the family, or, for MiniMax-H3, audio as the only reference - because the pipeline enforces those only once its checkpoint is loaded, minutes into an acknowledged run (dw/reference_limits.py, which reads each limit off the diffusers block that enforces it rather than restating it). A LoRA loaded onto the wrong checkpoint partition is an error for the opposite reason - the pipeline accepts it: MiniMax-H3's ref2va holds transformer_ref alone, so an FL2VA-trained adapter loads onto it, the run succeeds and only the identity retention is worse (dw/adapter_compatibility.py, #155). A weight_name carrying neither ref2v nor fl2v cannot be placed, so it is a warnings entry naming the rule rather than a refusal - a reference-trained checkpoint nobody has named yet still gets through.

    A valid answer also carries plan, what the run will execute for those arguments: fingerprint (sha256:… over the realized, expanded definition with the seed and the documentation keys removed and output:…/latest/… left unpinned - the same work hashes the same, a longer list or an edited stored prompt does not); steps, the expanded member count; list_entries, {variable: length} for each for_each over a list variable; cached_steps, how many of those steps the worker's step cache would serve (0 for an unseeded workflow, null when the worker is busy or did not answer - a workflow with no seed also gets a warning saying so, since 0 alone does not distinguish a disabled cache from an empty one); downloads_required, each model_name the hub cache does not hold as {repo, gb, gated, access_blocked} (gb from the hub, null when it could not be asked - ?sizes=false skips the hub, and then gated and access_blocked are null too); gated is the hub's own field for the repo (false, "auto" or "manual") or null when the lookup itself failed for a reason other than the gate; access_blocked is true when this box's Hugging Face token specifically has not been granted access to a gated repo, false when the repo isn't gated or the token is accepted, and null when it could not be determined either way. A gated repo's own metadata is served by the hub regardless of this token's access, so model_info succeeding proves nothing about the gate; a gated entry gets a second, real check - a HEAD request against one of the repo's own files - and it is that request's GatedRepoError that sets access_blocked: true (the pre-flight signal for what would otherwise be a 403 partway into a run, #186). access_blocked is null when there is no file to probe or the probe itself fails for an unrelated reason (e.g. offline) - "unknown" is not "not blocked". validate_workflow's warnings carries one line per entry with access_blocked: true. Each from_single_file URL is {repo: null, url, gb: null, gated: null, access_blocked: null}, since a direct file URL is never gated; and estimate, {minutes, basis, device, measured_on, partial, runs} from this box's own history when it has one and otherwise from the workflow's cost block - basis is observed (the cold median of this server's own finished runs of this shape, with runs saying how many; preferred over a curated figure, and quoted only for the bucket the caller's arguments fall in), catalog (the stored total, for a run whose lists are the ones it was measured with), per_entry (re-priced from a measured per-entry rate, when the entry carries per_entry), derived (the stored total extrapolated linearly over a list whose length the caller changed - an estimate, not a measurement), other_device (no entry for the serving backend; the first entry's figure, which is a warning rather than a quote) or unknown (no cost block, or more than one list changed so there is nothing honest to extrapolate along); a composed child's cost is added to a curated figure and partial is true when a child has none - an observed figure already measured the whole run, children included, so nothing is added to it and partial is false. plan is null when it could not be built; an invalid answer carries no plan key.

Files and models

  • GET /api/workflows — the stored workflow names, plus a details entry per workflow: description, kinds (the output content types' top-level halves), steps, variables (a count) and variable_names, and prompt_refs naming the stored prompts it leans on, and configures - for a workflow under models/, the templates/ name it is a tuned configuration of, empty when it is a template itself or when the name does not resolve (then configures_missing carries what was written). Four more say what the workflow makes, read off its definition (a top-level shape, traits or summary in the file overrides): shape, one of image, image-set, image-edit, shot, sequence, audio, text, utility; traits, a sorted subset of has-audio, chained, image-conditioned, identity-referenced, needs-input-media, composes-workflows; summary, the first sentence of the description, capped at 120 characters; and cost, the maintainer-measured {device, name, vram_gb, minutes} runs, or null when nobody has measured it - a list-driven workflow's cost entry may also carry a measured per_entry ({variable, minutes, entries}), the cost of one entry of the list it was measured against. The response's cost_basis says what that is - curated: figures a maintainer measured once and wrote into the workflow, never derived from this server's own job history, so null means nobody wrote one down rather than "this box has never run it". Beside it, observed is the derived figure the same listing is allowed to carry (#93): what this box's own finished runs of that workflow took, as cold_minutes/cold_runs (model load included, so comparable to a curated cost) and warm_minutes/warm_runs (model already resident), with the drivers the figure is for, since, and unclassified_runs when a run's persisted events were trimmed past its loading phase. Runs are bucketed by the workflow's declared cost_drivers - the variables that move its cost - so a 345-frame run never informs a 124-frame figure; a list driver buckets on its length. A workflow declaring no drivers falls back to runs that overrode nothing at all, and a run whose every step was a step-cache hit is excluded. The compact view carries only observed_minutes (cold) and observed_runs; GET /api/workflows/{name}/variables carries the whole block beside the defaults. Derived from the job rows in one query - so the figures outlive a pruned run directory - and cached against the jobs table's high-water mark rather than a file mtime, because a job landing changes every figure and changes no file. observed never replaces cost: a maintainer's claim on a named card and this machine's last week are different things. A models/ entry takes its shape and traits from the template it configures and keeps its own cost. A list-driven workflow (one with a for_each step) also carries lists: per list variable, the fields an entry takes, the steps run over it and the default's length. Enough to choose a workflow and know what to pass it without reading each one; the variable defaults are deliberately left out, being an order of magnitude more payload on a listing the UI reloads. Cached by file mtime

    Optional query params narrow and shrink it: ?shape=&traits=&configures=&include_models=&view=compact. shape keeps entries of that shape and traits (comma-separated) those carrying all of them - an unknown value in either is a 400 whose detail lists the vocabulary. configures=<template> keeps that template's model configs. view=compact is the agent's projection: it drops description, origin, writable, prompt_refs, steps and variables, keeps summary, shape, traits, cost, kinds, variable_names and lists (carried only when the workflow has a list-driven step, like configures), and lists templates only unless include_models=true or a configures asks otherwise. With no params the response is what it always was, plus the new fields

  • GET/PUT/DELETE /api/workflows/{name} — read, save, delete workflow files (confined to --workflow-dir)

  • GET /api/workflows/{name:path}/download — download a workflow file as JSON

  • GET /api/workflows/{name:path}/variables — a workflow's variables and what they default to, without the definition around them. Long string defaults are cut to 200 characters and named in truncated, including strings inside a list default, named like shots[0].prompt; full=true returns them whole. Each entry also carries origin and writable, and the body carries libraries, shadowed and the workspace it lists. Saving one whose validation gate itself crashes answers 500; an invalid one is 400

  • GET /api/prompts, GET/PUT/DELETE /api/prompts/{name} — the prompt library (confined to --prompt-dir, names held to what a prompt: reference can load); the listing carries libraries, shadowed and an origin/writable per entry in details; deleting a read-only entry is a 403; saves are validated against the prompt schema, served at GET /api/prompt-schema

  • GET /api/prompts/{name:path}/download — download a prompt file as text

  • GET /api/enhancers, POST /api/enhance — prompt-enhancement presets, and {"idea": ..., "preset": ..., "model_name": ..., "device": ...} to queue an enhancement as an ordinary job whose saved text file is the result

  • GET /api/gallery, GET /api/gallery/{name}/metadata, DELETE /api/gallery/{name} — outputs and their embedded metadata. Each gallery entry carries folder (the workflow identity, the run id dropped) and subfolder (what followed the run id - the final/intermediate a step's result.subfolder chose, '' when it chose none); ?folder= and ?subfolder= filter independently (?version= too - with ?folder=, the one run the gallery labels v4), and the reply's folders and subfolders list every distinct value over the whole tree, '' always a member of each so root-level files stay selectable

  • GET /api/gallery/{name:path}/download — download an output file

  • POST /api/gallery/archive — {"names": [...]} (1-1000) bundles a multi-file selection into one zip, named by each file's gallery-relative path so output subfolders survive. A browser cannot zip on its own and throttles a burst of single downloads, so the gallery's bulk download goes through here; an unknown or out-of-directory name 404s the whole request rather than yielding a partial archive

  • GET /api/workspaces, POST /api/workspaces ({"name": ...}), DELETE /api/workspaces/{name}?acknowledged=true — the workspaces on this server. The workspace root's own workflows/assets/outputs are the default workspace and a named one is a subdirectory beside them, sharing the root's one prompt library. Delete answers with what it would remove and refuses until acknowledged, refuses the default, and refuses a workspace with jobs still queued. Each listed workspace carries a usage ({files, bytes}) — roughly how much disk its own folders hold, walked at most once a minute per workspace and deliberately approximate; the shared prompt library counts against the default workspace alone rather than once per workspace. A workspace is a namespace, not a security boundary: the API token is all-or-nothing

    exports/ sits beside the workspace's own folders, holding one directory per exported job. It is a reserved name: no workspace can be called exports, and the folder is never listed as one.

  • GET /api/assets — the asset library: input media, each with the asset: reference a workflow carries rather than a path, since a path only means something on the server's own machine. Empty rather than an error when no library is configured. libraries lists the roots searched, in order, each {origin, root, writable} — which of them an upload or delete can actually reach is the one with writable: true. shadowed lists the entries a nearer library hides: same shape as an assets entry but without url (that URL would serve the shadowing file, not this one), plus shadowed_by naming the origin that won. The workflow and prompt listings carry the same libraries and shadowed ([{name, origin, shadowed_by}]); a filter narrows shadowed with the entries. Upload, keep and delete answer 409 This workspace has no asset library when there is none

  • POST /api/assets/keep ({"name": ..., "asset_name": ..., "overwrite": false, "shared": false}) — keep a generated file as an input asset under a stable name, returning its asset: reference. A run's files are named by the run that made them, which is the wrong thing for a later workflow to depend on: latest moves and a pinned run id breaks when outputs are pruned. The copy happens inside the workspace and is a hard link where the filesystem allows one, so keeping one frame of a large render costs no second copy of it. Refuses an existing name unless overwrite. asset_name may name a folder and takes the kept file's extension when it carries none (a contradicting one is a 400) — the same rule the upload route follows, and what keeps a kept asset from landing under an extensionless name the library listing never shows. "shared": true keeps it in <root>/common/assets instead — the library every workspace under the root shares, which is where a recurring cast belongs

  • DELETE /api/assets/{name} — remove one file from the asset library, deleting from whichever library on the search path holds it (the workspace's own before the shared one, the order asset: resolves in). An asset from a read-only examples library answers 403, the same as a read-only prompt or workflow; a name nothing holds answers 404

  • POST /api/assets/archive — {"names": [...]} (1-1000) bundles a multi-file asset selection into one zip, named by each file's library-relative path, which is the name its asset: reference carries. The gallery archive's counterpart on the input side; it resolves down the same search path a run does, so a selection spanning this workspace's library, the shared one and an examples tree downloads as one archive, and an unknown or out-of-library name 404s the whole request rather than yielding a partial one. A duplicate name (repeated in the selection, or differing only by leading/trailing whitespace) collapses onto the one zip entry. Media (image/video/audio) stores rather than deflates, unless it's a raw format that still compresses (.bmp, .wav) - everything else the libraries hold is an already-compressed container, and the response does not start until the archive is complete, so deflating it is latency the caller waits through for nothing. Everything else - .json, .md, .txt, an unrecognized extension - deflates; so does the export zip's text files (workflow.json, manifest.json, job.json, the README)

  • POST /api/uploads?filename=... — the raw bytes of one image, video or audio file (200MB ceiling, checked from Content-Length before a byte is read, and again on the body; extension held to the allowed image/video list), saved into the asset library's uploads/ subfolder - the shared library at <root>/common/assets when shared=true, this workspace's own otherwise - under a generated name, or under asset_name when one is given (cast/priya-voice.wav, folders allowed, the uploaded file's extension assumed, confined to the library the way keep_output's name is). Answers 201 with path - asset:uploads/<name>, the reference a saved workflow can carry and still resolve on a later run - and url, the same file under the /inputs mount, for the editor's preview. A server started without an asset library falls back to the output directory's uploads/ and an absolute path. This is how the UI's file pickers get a local file onto the machine that will run the workflow. The body is the file itself, so no multipart parser is needed for a single-file upload

  • GET /api/models, DELETE /api/models?repo={repo_id} — hub cache inventory and deletion

  • POST /api/models/download ({"repo_id": ...}), GET /api/models/downloads, POST /api/models/downloads/{id}/cancel — background snapshot downloads with byte-level progress

  • GET /api/system/diffusers, POST /api/system/diffusers/update — installed diffusers version/commit, and a background diffusers install/ update (refused while a job is running or queued). The POST body is optional JSON, {"commit": ..., "revert": ...}: with neither, it pip install --upgrades from GitHub HEAD; commit (7-40 hex characters, validated before it reaches the command line) pins the git install to that commit instead of HEAD; revert: true pins back to the known-good published release instead of installing from git - the diffusers floor version read from pyproject.toml (pip install diffusers==<floor>). commit and revert are mutually exclusive. The status response includes before (the version/commit that was installed when the update started) alongside the live version/commit, so a revert has a concrete before/after to compare

  • GET /api/memory, GET /api/health — worker VRAM/RAM stats and liveness; memory answers live (whether info was measured by this call), stale, reason (job_running, worker_stopped, worker_busy, worker_unreachable) and age_seconds, so a cached reading is never mistaken for the worker's memory now - info: null means nothing has been measured because nothing is resident. health also reports hostname, device and whether mcp is mounted, so a remote client can tell which machine answered

  • POST /api/memory/clear (#221) — drops every loaded pipeline and the step cache (MCP clear_memory), and returns the memory reading taken right after. Refused with 409 while a job is running or queued - the queue is FIFO, so the caller retries once it finishes rather than this call blocking until it does

  • GET /api/server — connection details for the Server page: hostname, version, device, the bind_host/port/wildcard_bind the server was started with, auth_required (whether a token is configured - never the token itself), mcp (mounted plus its path), the directories in use, and runtime (#222) - Python version, torch version and the CUDA version torch was built against, the NVIDIA driver version (via nvidia-smi, when it's on PATH), and the installed versions of diffusers, transformers, accelerate, bitsandbytes, peft, safetensors and sentencepiece (null for one not installed) - for diagnosing an environment mismatch between boxes without shelling in; the machine's non-loopback addresses; a client composes its URLs from an address, the port and the MCP path

Security model

The server is built to serve your own GPU to your own browser, not the network:

  • Binds to 127.0.0.1 by default. --host 0.0.0.0 (or any other non-loopback address) is possible; without a token configured (see Authentication, below) the server logs a startup warning, since anything that can reach that address can queue jobs and browse/delete files.
  • Requests carrying an Origin header are rejected (403) unless its hostname is a loopback name, the configured --host, or the hostname the request itself was addressed to (Host). The last clause lets a browser on another machine use a --host 0.0.0.0 server by LAN IP or hostname; it still blocks cross-site pages and DNS rebinding, where the attacker's page carries its own Origin while Host is whatever resolved. Scheme and port are ignored, so a TLS-terminating proxy that forwards Host unchanged needs no configuration. An Origin that cannot be parsed is refused the same way (403), not answered with a 500.
  • Every response carries X-Content-Type-Options: nosniff and X-Frame-Options: DENY: a browser renders nothing as a type the server did not declare, and no page elsewhere can frame the UI. The UI itself carries no Content-Security-Policy yet.
  • /outputs and /inputs share the UI's origin, where the API token lives in localStorage, so a file served as an active document type (text/html, application/xhtml+xml, text/xml, application/xml, image/svg+xml) carries Content-Security-Policy: sandbox: it opens under an opaque origin with no script. Range and ETag answers are unchanged. The engine does not write text/html or text/xml results at all (below), so such a file is one planted on disk.
  • Requests carrying a Host header that names neither a loopback address nor the configured --host are rejected (400). A wildcard bind (--host 0.0.0.0 or ::) skips this check - clients reach such a server by the machine's LAN IP or hostname, never by the bind address, so there is no allowlist to build from it. This is defense-in-depth, not the DNS-rebinding fix by itself - the Origin check above already covers browser requests, since a browser's Origin reflects the real requesting origin regardless of what DNS name resolved to this address. The Host check closes the remaining gap: a non-browser client (curl, a script, the MCP client) that never sends Origin at all.
  • Every path from HTTP input goes through dw/security.py validation; workflow files (both the /api/workflows CRUD routes and a workflow_path given to /api/jobs or /api/validate) are confined to the workflow directory, prompt files to the prompt directory, outputs to the output directory, and traversal (../) is blocked throughout.
  • Inline workflow definitions are schema-validated before queueing, and their base_dir is validated like any other path input.
  • A workflow JSON file can execute arbitrary Python (pre_load_modules, dotted *_type/config_type values - see Trust model). dw-serve refuses that surface by default for every job it runs, inline or from a file, MCP-submitted or not; --trust-workflows lifts the refusal for the whole server and should only be passed when nothing untrusted can reach POST /api/jobs.
  • --mcp mounts the MCP tool surface at /mcp (Streamable HTTP) behind the same token as /api, for an agent on another machine with no local install. Both /mcp and /mcp/ are answered, and the token is accepted only as an Authorization: Bearer header there - never as ?token=. It is refused on a non-loopback bind without a token. See REMOTE.md.

Authentication

There is no authentication by default - the checks above assume a trusted local machine or LAN. An optional static bearer token closes that gap:

python -m dw.serve --token "some-long-random-string"
# or
export DW_API_TOKEN="some-long-random-string"
python -m dw.serve

When a token is configured, every /api/* request must carry Authorization: Bearer <token> or gets a 401. The UI's own static files and /outputs (generated media) stay reachable without it - the page has to load far enough for a user to enter the token, and an <img>/<script> tag cannot attach a header anyway. That is why an active document served from /outputs or /inputs is sandboxed (Security model, above). A few GET API routes additionally accept the token as a ?token=... query parameter, because the browser loads them without being able to set headers: the SSE stream, GET /api/jobs/{id}/events (EventSource), and the gallery grid's GET /api/gallery/{name}/thumbnail (an <img> tag). The three /download routes (gallery output, workflow, prompt) accept ?token=... the same way, since a download button is a plain <a href download> navigation that cannot set a header either. That is a deliberate, narrower trade-off (a token that can leak into logs or browser history for those URLs) rather than a general alternative to the header - every other route accepts the header only.

The web UI has a one-time token field (next to the theme toggle) that stores the token in localStorage and attaches it to every API call, including the two query-parameter routes above. The MCP server reads the same DW_API_TOKEN variable (or dw-mcp --token), so one export configures both ends - see MCP.md. It is a convenience, not a credential vault - anyone with access to the browser profile can read it back out of localStorage.

A token configured this way is a single shared static secret, not a login system: there is one token, checked with a constant-time comparison, and no notion of separate users or sessions. It raises the bar for exposing the server on a LAN or beyond; it is not a substitute for a real network boundary (a firewall, a VPN, or simply binding to 127.0.0.1) for anything more exposed than that.

Running on another machine: REMOTE.md is the end-to-end recipe

  • token, firewall, systemd unit, the browser, --mcp, and what to do beyond the LAN.