This document describes the repeatable process for updating the
kohya-ss/sd-scripts version used by this Docker image and verifying that the
resulting image still starts and imports the expected training dependencies.
Use this procedure for routine sd-scripts updates, PyTorch or CUDA refreshes, dependency lockfile refreshes, and Docker image validation.
The update normally touches these files:
Dockerfilepyproject.tomluv.lockREADME.md
The repository uses pyproject.toml as the local dependency declaration, but
the source of truth for sd-scripts runtime dependencies is the upstream
requirements.txt at the selected sd-scripts commit.
This image also carries a few local bundled dependencies that may not be active
in upstream requirements.txt:
- BLIP captioning dependencies used by
finetune/make_captions.py. - WD14 ONNX captioning dependencies used by
finetune/tag_images_by_wd14_tagger.py --onnx. open-clip-torchfor SDXL-related workflows.lycoris-loraas an additional bundled network package.
Keep those local extras unless the README and supported image behavior are updated at the same time.
Install or make available:
- Git
- Docker Engine with BuildKit
uvhadolint
The validation commands below assume the image name
sd-scripts:update-test. Use another temporary tag if that tag is already
meaningful in your local environment.
Start from the latest main branch and create a focused update worktree:
git fetch origin main
git worktree add -b update-sd-scripts-vX.Y.Z .agents/worktrees/update-sd-scripts-vX.Y.Z origin/main
cd .agents/worktrees/update-sd-scripts-vX.Y.ZIf the branch or worktree already exists, inspect it before reusing it:
git status --short
git log --oneline --decorate -5Keep the update PR focused on the sd-scripts, dependency, Dockerfile, and
documentation changes needed for the selected version. Do not update VERSION
in the same PR unless the PR is intentionally publishing a Docker image
release.
List upstream version tags:
git ls-remote --tags https://github.com/kohya-ss/sd-scripts.git 'refs/tags/v*'For a candidate tag, fetch the exact commit and release date:
git clone --filter=blob:none --no-checkout https://github.com/kohya-ss/sd-scripts.git /tmp/sd-scripts-upstream
git -C /tmp/sd-scripts-upstream fetch --tags origin
git -C /tmp/sd-scripts-upstream show --no-patch --format='%H %ci %s' vX.Y.ZOnly adopt an sd-scripts release after the repository cooldown window has passed. The current baseline is a minimum of 7 days after the upstream release timestamp or tag commit timestamp.
Record the selected commit SHA. Prefer the tag commit for released versions instead of a branch name.
Print the upstream requirements and relevant installation notes for the selected version:
git -C /tmp/sd-scripts-upstream show vX.Y.Z:requirements.txt
git -C /tmp/sd-scripts-upstream show vX.Y.Z:README.md | sed -n '/About requirements.txt and PyTorch/,+20p'Update pyproject.toml so its [project].dependencies match the active
dependencies in upstream requirements.txt, translated to valid PEP 508
strings.
Keep the comment above the first dependency pointed at the exact upstream commit, for example:
# https://github.com/kohya-ss/sd-scripts/blob/<commit-sha>/requirements.txtWhen translating requirements:
- Keep pinned versions exactly pinned.
- Keep lower bounds and compatible-release specifiers as upstream writes them.
- Do not copy
-e .; the Dockerfile installs sd-scripts itself after the source checkout. - Preserve local extras that this image documents or intentionally bundles.
- If an upstream dependency is commented out, include it only when this image still needs it for documented behavior.
- If a local extra conflicts with the refreshed dependency graph, prefer a compatible newer version of the local extra over dropping the feature.
- Preserve
[tool.uv] exclude-newer = "P7D".
sd-scripts does not list PyTorch in requirements.txt because the correct
wheel depends on the CUDA target. Read the selected upstream README and choose
the PyTorch stack that matches this image's supported GPU generation.
For RTX 50 series support, use the upstream-recommended PyTorch and CUDA line when available. For example, PyTorch 2.13.0 with CUDA 13.0 uses:
"torch==2.13.0+cu130",
"torchvision==0.28.0+cu130",
"xformers==0.0.35",and:
[[tool.uv.index]]
name = "pytorch-cu130"
url = "https://download.pytorch.org/whl/cu130"
explicit = true
[tool.uv.sources]
torch = { index = "pytorch-cu130" }
torchvision = { index = "pytorch-cu130" }
xformers = { index = "pytorch-cu130" }Before locking, confirm that the selected wheels exist for Python 3.10 and Linux x86_64:
curl -fsSL https://download.pytorch.org/whl/cu130/torch/ | \
rg 'torch-2\.13\.0\+cu130-cp310-cp310-.*x86_64'
curl -fsSL https://download.pytorch.org/whl/cu130/torchvision/ | \
rg 'torchvision-0\.28\.0\+cu130-cp310-cp310-.*x86_64'
curl -fsSL https://download.pytorch.org/whl/cu130/xformers/ | \
rg 'xformers-0\.0\.35-.*x86_64'If the upstream README recommends a different CUDA line, replace cu130,
package versions, and the pytorch-cu130 index name consistently.
Do not assume that the CUDA runtime base image must use the same CUDA release as the PyTorch wheel index. PyTorch wheels bundle their CUDA user-space libraries, while the PyPI ONNX Runtime GPU wheel can require an older CUDA major version from the base image. Verify both stacks on a GPU before changing the runtime base. A mixed setup is acceptable when PyTorch, xformers, and ONNX Runtime all execute successfully and the reason is recorded in the changelog and pull request.
Set SD_SCRIPTS_VERSION to the selected upstream commit:
ARG SD_SCRIPTS_VERSION=<commit-sha>Keep base images pinned by digest:
ARG CUDA_RUNTIME_IMAGE=nvidia/cuda:<tag>@sha256:<digest>
ARG UV_IMAGE=ghcr.io/astral-sh/uv:<tag>@sha256:<digest>To refresh a base image digest, inspect the image:
docker buildx imagetools inspect nvidia/cuda:<tag>
docker buildx imagetools inspect ghcr.io/astral-sh/uv:<tag>Use the multi-architecture index digest unless the image is intentionally limited to one platform.
Keep apt packages version-pinned so hadolint can verify the Dockerfile. To
find apt candidate versions for the pinned CUDA runtime image:
docker run --rm 'nvidia/cuda:<tag>@sha256:<digest>' bash -lc \
'set -euo pipefail; apt-get update >/dev/null; apt-cache policy ca-certificates git libgl1 libglib2.0-0t64 tk'Package names can change between Ubuntu releases. For Ubuntu 24.04, use
libgl1 and libglib2.0-0t64 instead of the older Ubuntu 22.04 package names
libgl1-mesa-glx and libglib2.0-0.
If the CUDA major or minor version changes, verify the libnvrtc.so workaround
path before editing it:
docker run --rm 'nvidia/cuda:<tag>@sha256:<digest>' bash -lc \
'ls -l /usr/local/cuda-*/targets/x86_64-linux/lib/libnvrtc.so*'Then update both the source and destination paths in the Dockerfile.
Run lockfile resolution from the repository root or the update worktree:
uv lock --upgradeAfter the lockfile update, verify it is consistent:
uv lock --checkReview the resolver output and uv.lock diff for unexpected package upgrades.
If a package appears newer than the cooldown policy should allow, stop and
investigate before building the image.
Pay special attention to:
- PyTorch, torchvision, xformers, and NVIDIA CUDA wheel versions.
protobufconstraints, because ONNX and open-clip packages often constrain it differently.numpymajor-version changes, because older captioning and ONNX packages may not support the selected version.- Local extras that disappeared because upstream commented them out.
If README links point at upstream sd-scripts documentation by commit SHA, update those links to the selected commit. For example:
https://github.com/kohya-ss/sd-scripts/blob/<commit-sha>/docs/train_README-ja.mdIf the update changes the image's supported GPU, CUDA, Python, or Docker requirements, update the README requirements section in the same PR.
Run Dockerfile linting:
hadolint DockerfileRun Docker's build checks:
docker build --check .Both commands should complete without warnings for routine updates.
If hadolint reports DL3008, pin the affected apt package versions. If it
reports DL3003 or SC2164 around a cd, prefer git -C or WORKDIR
instead of changing directories inside a shell block.
Build the image:
docker build --progress=plain -t sd-scripts:update-test .The first build after a dependency update can take several minutes because the CUDA, PyTorch, xformers, ONNX Runtime, and NVIDIA library wheels are large.
If the build fails while pulling ghcr.io/astral-sh/uv, check local Docker
authentication state first. The image may still be publicly readable even when
an expired local credential causes Docker to fail.
Check that core Python packages import and report the expected versions:
docker run --rm --entrypoint python sd-scripts:update-test -c \
'import torch, torchvision, xformers, onnx, onnxruntime, open_clip, timm, fairscale; import accelerate, transformers, diffusers; print(torch.__version__); print(torch.version.cuda); print(torchvision.__version__); print(xformers.__version__); print(onnx.__version__); print(onnxruntime.__version__); print(open_clip.__version__); print(timm.__version__); print(accelerate.__version__); print(transformers.__version__); print(diffusers.__version__)'Check that a representative training entrypoint can render help:
docker run --rm sd-scripts:update-test \
--help
docker run --rm sd-scripts:update-test \
sdxl_train_network.py --helpWhen an NVIDIA GPU is available, also run a GPU visibility smoke test:
docker run --rm --gpus all --entrypoint python sd-scripts:update-test -c \
'import torch; print(torch.cuda.is_available()); print(torch.cuda.get_device_name(0) if torch.cuda.is_available() else "no cuda")'For PyTorch-stack updates, exercise CUDA kernels instead of relying on imports
or torch.cuda.is_available() alone. At minimum, run a PyTorch tensor
operation, xformers memory-efficient attention, and a bitsandbytes operation.
For ONNX Runtime updates, create an inference session with
CUDAExecutionProvider, run a small model, and confirm that the session did
not fall back to CPUExecutionProvider.
Also run the WD14 ONNX captioning path with an actual image and model. This loads the full accelerate, timm, ONNX, and ONNX Runtime path that previously passed import-only checks but failed at startup:
docker run --rm --gpus all -v /path/to/smoke-data:/work \
sd-scripts:update-test \
finetune/tag_images_by_wd14_tagger.py --onnx --batch_size 1 \
--model_dir /work/model /work/imagesFor a release that changes PyTorch, CUDA, xformers, or sd-scripts training behavior, run at least one small project-specific training or captioning workflow before merging when practical.
The Docker image includes the selected upstream sd-scripts checkout under
/opt/sd-scripts, including the upstream tests/ directory. Install the
uv-managed dev test dependencies temporarily in a disposable container and run
the same lightweight upstream pytest set that CI uses:
scripts/run-sd-scripts-release-tests.sh --image sd-scripts:update-testBefore relying on that set, compare every path named in
scripts/run-sd-scripts-release-tests.sh with the selected upstream checkout.
Upstream releases can remove or rename tests; a stale explicit pytest path
causes the CI release gate to fail before any tests run. Update the selected
path list in the same pull request as the sd-scripts version change, while
retaining the documented exclusions below unless the image support contract
also changes.
This set intentionally excludes upstream tests and scripts that need model
checkpoints, local datasets, or dependencies this image does not bundle, such
as tests/test_optimizer.py requiring dadaptation.
When an NVIDIA GPU and compatible test checkpoint are available, also run at
least one upstream inpainting training smoke test. The scripts generate
synthetic image data automatically when --data is omitted:
docker run --rm --gpus all --entrypoint bash \
-v /path/to/models:/models:ro \
sd-scripts:update-test \
tests/run_sd15_inpainting_test.sh --mode ft --model /models/sd15-inpainting.safetensors --steps 1
docker run --rm --gpus all --entrypoint bash \
-v /path/to/models:/models:ro \
sd-scripts:update-test \
tests/run_sdxl_inpainting_test.sh --mode ft --model /models/sdxl-base-or-inpainting.safetensors --steps 1If GPU or checkpoint coverage is skipped, record the exact reason in the pull request.
Review the diff:
git diff --stat
git diff --check
git diff -- Dockerfile pyproject.toml uv.lock README.md docs/update-sd-scripts.mdConfirm:
SD_SCRIPTS_VERSIONis an exact commit SHA from the selected tag.- The
pyproject.tomlupstream requirements comment points at the same commit. - README upstream documentation links point at the same commit.
- The CUDA base image digest matches the intended CUDA and Ubuntu tag.
- Apt package versions were checked inside the selected CUDA image.
uv.lockresolves the selected PyTorch CUDA index.- Local extras are intentionally kept, updated, or removed.
- Upstream sd-scripts pytest release tests passed in the built image.
- PyTorch, xformers, bitsandbytes, and ONNX Runtime executed GPU operations, and the WD14 ONNX captioning path completed without provider fallback.
- Every explicit pytest path in
scripts/run-sd-scripts-release-tests.shexists in the selected upstream checkout. - Upstream inpainting shell tests were run, or GPU/checkpoint coverage was explicitly skipped.
- Verification commands and any skipped GPU tests are documented in the PR.
Create a focused commit, push the branch, and open a pull request.
The PR body should include:
- The selected sd-scripts version and commit SHA.
- Any CUDA, Ubuntu, PyTorch, xformers, or local-extra dependency changes.
- The checks run locally.
- Whether upstream pytest release tests passed in the built image.
- Whether runtime GPU and upstream inpainting training tests were tested or skipped.
Do not update VERSION as part of the sd-scripts update unless the same PR is
also intentionally publishing a Docker image release. Follow the README release
procedure separately after the update PR is merged.