Follow-up to #64, which added linux/arm64 to the CD image build so images run natively on the arm64 (Oracle Ampere) deploy host. It works, but the arm64 half is built under QEMU emulation on an amd64 runner, which is slow enough to be worth fixing.
Measured cost
So the arm64 half costs roughly +12 min on a cold cache.
Root cause
It is not the UniDic download. That layer (Dockerfile:43, ~700 MB) is network-bound, so emulation barely affects it.
The dominant cost is uv sync at Dockerfile:24. As the comment at Dockerfile:9-13 notes, pyopenjtalk ships source-only on PyPI with no wheels for any arch, so it compiles its bundled OpenJTalk/HTS-engine C++ from sdist during the build. That is pure CPU-bound C++ compilation, and under QEMU it runs roughly an order of magnitude slower than native.
Why this is not urgent
Layer caching is already correct — every expensive step sits before COPY . . (Dockerfile:45) and is gated by uv.lock / download_unidic.sh, so an app-code-only commit reuses all of it. GHA cache usage is currently 3.64 GB / 10 GB across 41 entries, so LRU eviction is not yet a factor either.
The thing that will actually bite: GitHub evicts caches not accessed for 7 days. CD only runs on pushes to main, so any gap longer than a week means the next run pays the full cold-cache emulation cost again.
Proposal
This repo is public, so GitHub's native arm64 runners (ubuntu-24.04-arm) are free. Build each arch natively and merge the manifest, eliminating QEMU entirely.
Sketch:
- Matrix job over
{ubuntu-latest → linux/amd64, ubuntu-24.04-arm → linux/arm64}; each builds and pushes by digest (outputs: type=image,push-by-digest=true,name-canonical=true) with its own cache scope.
- A
merge job runs docker buildx imagetools create over the per-arch digests to produce the tagged multi-arch index, applying the docker/metadata-action tags.
Complication worth planning for
The current smoke test depends on the single-platform shape: it builds amd64 with load: true (buildx cannot load a multi-platform result), runs the container on the runner, then pushes both arches in a second build that reuses the cache. A matrix split changes that flow — but it also improves it, since each arch can then smoke-test its own image natively instead of only amd64 being exercised. The merge job should only publish tags once both arch smoke tests pass, so a broken arm64 image can never take the dev/stable tag.
Also note provenance: true / sbom: true currently attach attestations (the two unknown/unknown entries in the index); the merge step needs to preserve these.
Not blocking
#64 is deployed and correct — ghcr.io/sessatakuma/api-tools:dev is now an OCI image index with both linux/amd64 and linux/arm64. This issue is purely a CI wall-clock optimization.
Follow-up to #64, which added
linux/arm64to the CD image build so images run natively on the arm64 (Oracle Ampere) deploy host. It works, but the arm64 half is built under QEMU emulation on an amd64 runner, which is slow enough to be worth fixing.Measured cost
So the arm64 half costs roughly +12 min on a cold cache.
Root cause
It is not the UniDic download. That layer (
Dockerfile:43, ~700 MB) is network-bound, so emulation barely affects it.The dominant cost is
uv syncatDockerfile:24. As the comment atDockerfile:9-13notes, pyopenjtalk ships source-only on PyPI with no wheels for any arch, so it compiles its bundled OpenJTalk/HTS-engine C++ from sdist during the build. That is pure CPU-bound C++ compilation, and under QEMU it runs roughly an order of magnitude slower than native.Why this is not urgent
Layer caching is already correct — every expensive step sits before
COPY . .(Dockerfile:45) and is gated byuv.lock/download_unidic.sh, so an app-code-only commit reuses all of it. GHA cache usage is currently 3.64 GB / 10 GB across 41 entries, so LRU eviction is not yet a factor either.The thing that will actually bite: GitHub evicts caches not accessed for 7 days. CD only runs on pushes to
main, so any gap longer than a week means the next run pays the full cold-cache emulation cost again.Proposal
This repo is public, so GitHub's native arm64 runners (
ubuntu-24.04-arm) are free. Build each arch natively and merge the manifest, eliminating QEMU entirely.Sketch:
{ubuntu-latest → linux/amd64, ubuntu-24.04-arm → linux/arm64}; each builds and pushes by digest (outputs: type=image,push-by-digest=true,name-canonical=true) with its own cache scope.mergejob runsdocker buildx imagetools createover the per-arch digests to produce the tagged multi-arch index, applying thedocker/metadata-actiontags.Complication worth planning for
The current smoke test depends on the single-platform shape: it builds amd64 with
load: true(buildx cannotloada multi-platform result), runs the container on the runner, then pushes both arches in a second build that reuses the cache. A matrix split changes that flow — but it also improves it, since each arch can then smoke-test its own image natively instead of only amd64 being exercised. The merge job should only publish tags once both arch smoke tests pass, so a broken arm64 image can never take thedev/stabletag.Also note
provenance: true/sbom: truecurrently attach attestations (the twounknown/unknownentries in the index); the merge step needs to preserve these.Not blocking
#64 is deployed and correct —
ghcr.io/sessatakuma/api-tools:devis now an OCI image index with bothlinux/amd64andlinux/arm64. This issue is purely a CI wall-clock optimization.