Skip to content

Resize in YUV while converting to RGB (one swscale pass); bump to 0.4.2 - #17

Closed
MilkClouds wants to merge 2 commits into
mainfrom
fused-resize
Closed

MilkClouds wants to merge 2 commits into
mainfrom
fused-resize

Conversation

@MilkClouds

@MilkClouds MilkClouds commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Resize used to be slower than decoding at native size. It now resizes in YUV while converting to RGB, in one swscale pass at the output size, so a resized decode costs the same as a native-size one or slightly less. This is the only Resize behavior; there is no flag.

TorchCodec converts to RGB first and then resizes. Our outputs therefore differ from TorchCodec's by about 0.9 levels on average, mostly because of TorchCodec's darker bias, and they are closer to a float64 reference. Parity with TorchCodec still holds for decoding without a resize and for crops: both are bit-identical to 0.4.1.

What the cost actually was

Full-size YUV→RGB conversion is not the bottleneck. On an i7-14700K with conda-forge FFmpeg 7.1.1, swscale converts a 640×480 yuv420p frame to rgb24 in 0.07 ms on its SIMD unscaled path. The scaler flags make no difference there, and slice threads only add overhead (0.08–0.12 ms).

The old RGB→RGB resize pass cost 0.56 ms at 224² and 0.38 ms at 128². The single pass writes planar GBRP with SIMD and costs 0.12 ms and 0.07 ms, plus interleaving. Writing packed rgb24 at full chroma would take 0.30 / 0.13 ms, because swscale's packed full-chroma writer is C only. The no-resize path has no cheap win to take.

Design

There is one resize mechanism: Decoder::resize. It runs one bilinear swscale pass from a source region straight to RGB at the output size, with flags BILINEAR | FULL_CHR_H_INT | FULL_CHR_H_INP, writing GBRP/GBRP16LE and then interleaving. Pipeline decides what that source region is:

  • First resize. It reads the decoded planes. Crops before it become plane offsets.
    • An offset inside a chroma pair (odd, for 4:2:0) starts the chroma plane at that pair and sets src_h_chr_pos/src_v_chr_pos to −128, so chroma stays sited correctly.
    • A display rotation moves after this resize, onto the small RGB frame. The result is bit-identical to rot90 of the unrotated result.
  • Later resizes (after a resize and a crop) read the RGB image with the same pass. FULL_CHR_H_INP keeps it from decimating chroma; the mean absolute error against a float resize of the same RGB drops from 0.20–0.98 to 0.05–0.19.
  • Fallback. After a crop, pixel formats that plane offsets cannot address (paletted, bitstream, hardware, or packed with subsampled chroma) are converted at full size first, and the resize then reads RGB.
  • Same-size resizes are dropped as no-ops.
  • Removed: the old RGB→RGB Resizer (fixed SWS_BILINEAR, packed in and out) and the fused flag.

Known swscale limit (documented). For a region of odd width or height, swscale treats the subsampled chroma as if it covered the region exactly, which stretches chroma by up to half a pixel at the far edge. Even-sized regions are unaffected.

Bit-identical to 0.4.1 when nothing is resized: 54 arrays checked, covering yuv420p, yuv420p10le, yuvj420p, bgr0, gray, yuv422p, yuv444p16le, a rotated clip and AV1; no transform and crops; uint8 and uint16.

Pixel differences

Measured on 640×480 GOP-2 clips, 20 frames each. "real" is a 640×480 crop of the Sintel 480p trailer; "syn" is testsrc2. The float64 reference is exact limited-range YUV→RGB with bilinear, centered chroma upsampling, followed by the antialiased bilinear filter used by TorchVision v2 and PIL.

clip output vs TorchCodec: max / mean / bias / >1 level / p99.9 vs float64 ref, mean abs (bias): TorchCodec / TensorCodec
syn AV1 224² 42 / 0.80 / +0.60 / 14.6% / 18 1.27 (−0.65) / 1.03 (−0.05)
syn AV1 128² 25 / 0.76 / +0.60 / 15.0% / 12 0.84 (−0.65) / 0.60 (−0.04)
syn AV1 crop 400² → 224² 59 / 1.07 / +0.72 / 22.3% / 33 1.43 (−0.76) / 1.15 (−0.04)
real AV1 224² 17 / 0.90 / +0.87 / 21.8% / 5 0.91 (−0.90) / 0.25 (−0.03)
real AV1 128² 11 / 0.89 / +0.87 / 20.8% / 4 0.91 (−0.90) / 0.23 (−0.04)
real AV1 crop 400² → 224² 18 / 1.08 / +1.04 / 26.5% / 4 1.09 (−1.08) / 0.29 (−0.03)
syn H.264 224² 49 / 0.81 / +0.59 / 15.2% / 19 1.27 (−0.65) / 1.05 (−0.06)
syn H.264 128² 27 / 0.78 / +0.60 / 15.7% / 12 0.83 (−0.64) / 0.61 (−0.04)
syn H.264 crop 400² → 224² 58 / 1.08 / +0.72 / 22.6% / 33 1.43 (−0.76) / 1.16 (−0.04)
real H.264 224² 26 / 0.90 / +0.86 / 21.9% / 5 0.90 (−0.89) / 0.25 (−0.03)
real H.264 128² 16 / 0.89 / +0.86 / 20.9% / 4 0.90 (−0.89) / 0.22 (−0.04)
real H.264 crop 400² → 224² 27 / 1.08 / +1.04 / 26.6% / 5 1.08 (−1.07) / 0.29 (−0.04)

Where the bias comes from. I measured native-size output against exact arithmetic: swscale's fast full-size conversion is 0.6–1.0 levels darker. The old RGB resize itself is unbiased, so it simply carries that darkness through. Native-size decoding uses the same conversion in TensorCodec and TorchCodec, so the bias is shared there.

The largest differences are at sharp, saturated color edges. There, nearest-neighbour chroma followed by an RGB resize disagrees with interpolating chroma directly.

Tests

  • Resizes are checked against the float64 reference on tests/resources/nasa_13013.mp4, which is natural content (320×180, BT.709, limited range). The reference uses NumPy only, and its filter matches torch.nn.functional.interpolate(antialias=True) to 3e-16.
  • Tolerances:
    • Mean error within 0.25 levels of zero, and mean absolute error below 0.75 levels.
    • Measured on even-sized regions: |bias| ≤ 0.16 and mean absolute error 0.19–0.63.
    • TorchCodec's path, with its 0.6–1.1 level bias, would fail the bias bound.
  • Sizes covered: uneven, a 1.6× downscale, 2×, a 2× upscale, uint16 output, and crops at even and odd offsets.
  • Siting check: the output must be at least 0.15 levels closer to the reference than to a reference with chroma shifted by half a pixel in either direction (measured 0.62 against 0.86–0.90).
  • Other formats: yuv422p, bgr0 and gray, with crops.
  • Rotation: the rotated, cropped and resized output must equal rot90 of the unrotated result.
  • Pipeline behavior: ordering, a second resize, and same-size resizes as no-ops.
  • TorchCodec oracle test: crop parity is unchanged; resize cases now assert shape only. All other TorchCodec parity tests are untouched.

Benchmarks (ms per returned frame)

Command: python -m benchmarks.transform_bench. Each value is the mean of two runs per version; each run is the median of 50 alternating calls. Setup: one decoder thread, one core pinned, i7-14700K under WSL2 with other load present, FFmpeg 7.1.1. The 1-frame case is the middle frame; the 20-frame case is every 15th frame, so each frame needs a seek.

clip frames native 0.4.1 native 0.4.2 224² 0.4.1 224² 0.4.2 128² 0.4.1 128² 0.4.2
syn AV1 1 1.89 1.93 2.34 1.91 2.17 1.83
syn AV1 20 1.62 1.62 2.09 1.60 1.89 1.52
real AV1 1 2.09 1.92 2.52 1.89 2.37 1.81
real AV1 20 1.58 1.52 2.05 1.48 1.87 1.40
syn H.264 1 2.24 2.30 2.78 2.30 2.52 2.18
syn H.264 20 1.75 1.78 2.23 1.75 2.05 1.67
real H.264 1 3.10 3.06 3.58 3.08 3.37 2.96
real H.264 20 2.03 2.02 2.49 1.99 2.31 1.91

Changes

  • av/src/ffmpeg.rs: Pipeline, region_planes, the unified Resizer and Decoder::resize, plus configure_colors, scale_context, ensure_frame and interleave. The full-size conversion is unchanged and bit-identical.
  • src/tensorcodec/transforms.py: the Resize docstring states the behavior. There is no API change.
  • tests/test_transforms.py: float64-reference resize tests replace the FFmpeg RGB-resize comparisons; TorchCodec resize cases are shape-only.
  • docs/compatibility.md gets a new "Resizing" section with the table above; README wording is updated.
  • benchmarks/transform_bench.py is new and linted in CI; benchmarks/README.md shows 0.4.1 vs 0.4.2.
  • A separate commit bumps the version to 0.4.2.

Verification

  • pytest --compare: 472 passed.
  • Contract tests with --backend torchcodec: 67 passed.
  • test_transforms.py with an FFmpeg 6.1 CLI: passed.
  • ruff, cargo fmt --check, cargo clippy -D warnings: clean.
  • Codex review of the branch against origin/main: no actionable findings. The first round, on the earlier opt-in version, found a benchmark crash on short clips, which is fixed.

🤖 Generated with Claude Code

https://claude.ai/code/session_01YAijSnAE4aAjuS1EcCTwqS

MilkClouds and others added 2 commits October 5, 2026 13:41
Resize used to convert each frame to RGB at full size, then resize it in
RGB with a second swscale pass, as TorchCodec 0.17 does. The conversion
is cheap (~0.07 ms for 640x480 yuv420p on swscale's SIMD unscaled path);
the RGB pass cost 0.4-0.6 ms per frame, so a resized decode was slower
than a native-size one.

Every resize is now one bilinear swscale pass from its source region
straight to RGB at the output size: planar output (SIMD), full-resolution
chroma, then interleaved. The first resize reads the decoded planes:
crops before it become plane offsets, with an offset inside a chroma
pair re-sited via src_h/v_chr_pos, and a display rotation moves after
it. Later resizes read the RGB image. Formats that plane offsets cannot
address are converted at full size after a crop. Resizes to the current
size are dropped.

Decoding without a resize and crops stay bit-identical to 0.4.1 and keep
TorchCodec parity. Resized frames differ from TorchCodec by ~0.9 levels
on average, mostly TorchCodec's darker bias from swscale's full-size
conversion, and are closer to a float64 reference; tests now check them
against that reference. Deltas and timings are in docs/compatibility.md
and benchmarks/README.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YAijSnAE4aAjuS1EcCTwqS
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YAijSnAE4aAjuS1EcCTwqS
@MilkClouds MilkClouds changed the title Add Resize(fused=True): scale during color conversion; bump to 0.4.2 Resize in YUV while converting to RGB (one swscale pass); bump to 0.4.2 Oct 5, 2026
@MilkClouds

Copy link
Copy Markdown
Collaborator Author

Closing in favor of #18: not needed now, and several cases (upscaling, high-res and high-bit-depth sources, full-range input) are unverified.

@MilkClouds MilkClouds closed this Oct 5, 2026
@MilkClouds
MilkClouds deleted the fused-resize branch October 5, 2026 05:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant