Repository navigation
Resize in YUV while converting to RGB (one swscale pass); bump to 0.4.2 - #17
Closed
MilkClouds wants to merge 2 commits into
Closed
MilkClouds wants to merge 2 commits into
MilkClouds wants to merge 2 commits into
Conversation
MilkClouds
force-pushed
the
fused-resize
branch
from
October 5, 2026 03:52
927513d to
f4d4173
Compare
Resize used to convert each frame to RGB at full size, then resize it in RGB with a second swscale pass, as TorchCodec 0.17 does. The conversion is cheap (~0.07 ms for 640x480 yuv420p on swscale's SIMD unscaled path); the RGB pass cost 0.4-0.6 ms per frame, so a resized decode was slower than a native-size one. Every resize is now one bilinear swscale pass from its source region straight to RGB at the output size: planar output (SIMD), full-resolution chroma, then interleaved. The first resize reads the decoded planes: crops before it become plane offsets, with an offset inside a chroma pair re-sited via src_h/v_chr_pos, and a display rotation moves after it. Later resizes read the RGB image. Formats that plane offsets cannot address are converted at full size after a crop. Resizes to the current size are dropped. Decoding without a resize and crops stay bit-identical to 0.4.1 and keep TorchCodec parity. Resized frames differ from TorchCodec by ~0.9 levels on average, mostly TorchCodec's darker bias from swscale's full-size conversion, and are closer to a float64 reference; tests now check them against that reference. Deltas and timings are in docs/compatibility.md and benchmarks/README.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YAijSnAE4aAjuS1EcCTwqS
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YAijSnAE4aAjuS1EcCTwqS
MilkClouds
force-pushed
the
fused-resize
branch
from
October 5, 2026 04:44
f4d4173 to
9e0bead
Compare
Collaborator
Author
|
Closing in favor of #18: not needed now, and several cases (upscaling, high-res and high-bit-depth sources, full-range input) are unverified. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Resizeused to be slower than decoding at native size. It now resizes in YUV while converting to RGB, in one swscale pass at the output size, so a resized decode costs the same as a native-size one or slightly less. This is the onlyResizebehavior; there is no flag.TorchCodec converts to RGB first and then resizes. Our outputs therefore differ from TorchCodec's by about 0.9 levels on average, mostly because of TorchCodec's darker bias, and they are closer to a float64 reference. Parity with TorchCodec still holds for decoding without a resize and for crops: both are bit-identical to 0.4.1.
What the cost actually was
Full-size YUV→RGB conversion is not the bottleneck. On an i7-14700K with conda-forge FFmpeg 7.1.1, swscale converts a 640×480 yuv420p frame to rgb24 in 0.07 ms on its SIMD unscaled path. The scaler flags make no difference there, and slice threads only add overhead (0.08–0.12 ms).
The old RGB→RGB resize pass cost 0.56 ms at 224² and 0.38 ms at 128². The single pass writes planar GBRP with SIMD and costs 0.12 ms and 0.07 ms, plus interleaving. Writing packed rgb24 at full chroma would take 0.30 / 0.13 ms, because swscale's packed full-chroma writer is C only. The no-resize path has no cheap win to take.
Design
There is one resize mechanism:
Decoder::resize. It runs one bilinear swscale pass from a source region straight to RGB at the output size, with flagsBILINEAR | FULL_CHR_H_INT | FULL_CHR_H_INP, writing GBRP/GBRP16LE and then interleaving.Pipelinedecides what that source region is:src_h_chr_pos/src_v_chr_posto −128, so chroma stays sited correctly.rot90of the unrotated result.FULL_CHR_H_INPkeeps it from decimating chroma; the mean absolute error against a float resize of the same RGB drops from 0.20–0.98 to 0.05–0.19.Resizer(fixedSWS_BILINEAR, packed in and out) and thefusedflag.Known swscale limit (documented). For a region of odd width or height, swscale treats the subsampled chroma as if it covered the region exactly, which stretches chroma by up to half a pixel at the far edge. Even-sized regions are unaffected.
Bit-identical to 0.4.1 when nothing is resized: 54 arrays checked, covering yuv420p, yuv420p10le, yuvj420p, bgr0, gray, yuv422p, yuv444p16le, a rotated clip and AV1; no transform and crops; uint8 and uint16.
Pixel differences
Measured on 640×480 GOP-2 clips, 20 frames each. "real" is a 640×480 crop of the Sintel 480p trailer; "syn" is
testsrc2. The float64 reference is exact limited-range YUV→RGB with bilinear, centered chroma upsampling, followed by the antialiased bilinear filter used by TorchVision v2 and PIL.Where the bias comes from. I measured native-size output against exact arithmetic: swscale's fast full-size conversion is 0.6–1.0 levels darker. The old RGB resize itself is unbiased, so it simply carries that darkness through. Native-size decoding uses the same conversion in TensorCodec and TorchCodec, so the bias is shared there.
The largest differences are at sharp, saturated color edges. There, nearest-neighbour chroma followed by an RGB resize disagrees with interpolating chroma directly.
Tests
tests/resources/nasa_13013.mp4, which is natural content (320×180, BT.709, limited range). The reference uses NumPy only, and its filter matchestorch.nn.functional.interpolate(antialias=True)to 3e-16.rot90of the unrotated result.Benchmarks (ms per returned frame)
Command:
python -m benchmarks.transform_bench. Each value is the mean of two runs per version; each run is the median of 50 alternating calls. Setup: one decoder thread, one core pinned, i7-14700K under WSL2 with other load present, FFmpeg 7.1.1. The 1-frame case is the middle frame; the 20-frame case is every 15th frame, so each frame needs a seek.Changes
av/src/ffmpeg.rs:Pipeline,region_planes, the unifiedResizerandDecoder::resize, plusconfigure_colors,scale_context,ensure_frameandinterleave. The full-size conversion is unchanged and bit-identical.src/tensorcodec/transforms.py: theResizedocstring states the behavior. There is no API change.tests/test_transforms.py: float64-reference resize tests replace the FFmpeg RGB-resize comparisons; TorchCodec resize cases are shape-only.docs/compatibility.mdgets a new "Resizing" section with the table above; README wording is updated.benchmarks/transform_bench.pyis new and linted in CI;benchmarks/README.mdshows 0.4.1 vs 0.4.2.Verification
pytest --compare: 472 passed.--backend torchcodec: 67 passed.test_transforms.pywith an FFmpeg 6.1 CLI: passed.cargo fmt --check,cargo clippy -D warnings: clean.origin/main: no actionable findings. The first round, on the earlier opt-in version, found a benchmark crash on short clips, which is fixed.🤖 Generated with Claude Code
https://claude.ai/code/session_01YAijSnAE4aAjuS1EcCTwqS