Problem
Resize decodes, converts YUV→RGB at full size, then runs a second RGB→RGB bilinear swscale pass. The second pass is most of the resize cost: on a 640×480 clip the full-size conversion is ~0.07 ms/frame (swscale's unscaled SIMD path) while the RGB resize is ~0.56 ms at 224² and ~0.38 ms at 128². A resized decode is therefore slower than a native one (real AV1, 20 frames: native 1.58 ms/frame, 224² 2.05, 128² 1.87).
Prototype
#17 (closed) did both in one swscale pass: source planes (crops as plane offsets, chroma siting corrected for odd offsets) → GBRP at the output size with SWS_BILINEAR | SWS_FULL_CHR_H_INT | SWS_FULL_CHR_H_INP, then interleave to RGB; display rotation moved after the resize; one resize mechanism, old RGB→RGB resizer removed.
- Speed (640×480, real AV1, 20 frames): 224² 2.05 → 1.48 ms/frame, 128² 1.87 → 1.40 (about 25%); resized decodes cost about the same as native.
- Pixels: resize is linear and the colour transform affine, so the order only matters through clipping, the intermediate 8-bit rounding, and the conversion arithmetic. Against a float64 reference the one-pass output has ~0.25 mean abs error vs ~0.9 for the current path (whose full-size fast converter is ~0.6–1.0 levels dark). Max differences (17–59) sit at saturated chroma edges. Differs from TorchCodec 0.17's convert-then-resize by ~0.9 levels on average.
Open before adopting
Not yet verified, each a possible regression:
- Speed: upscaling (output larger than input), 1080p/4K sources, large outputs (the planar→packed interleave scales with output size), 10/16-bit sources and outputs, 4:4:4 and RGB sources, chained resizes.
- Accuracy: full-range sources (yuvj420p/MJPEG), BT.601 and unspecified colour spaces, beyond the BT.709 limited-range clip tested.
- Odd-sized regions stretch subsampled chroma by up to half a pixel at the far edge.
- Resize parity with TorchCodec would no longer hold (decode without resize and crops stay identical).
A benchmark matrix ({360p, 480p, 1080p, 4K} × {down, up} × {8-bit 420, 10-bit 420, 444, full-range}) and accuracy tests on those sources should pass before this lands.
Problem
Resizedecodes, converts YUV→RGB at full size, then runs a second RGB→RGB bilinear swscale pass. The second pass is most of the resize cost: on a 640×480 clip the full-size conversion is ~0.07 ms/frame (swscale's unscaled SIMD path) while the RGB resize is ~0.56 ms at 224² and ~0.38 ms at 128². A resized decode is therefore slower than a native one (real AV1, 20 frames: native 1.58 ms/frame, 224² 2.05, 128² 1.87).Prototype
#17 (closed) did both in one swscale pass: source planes (crops as plane offsets, chroma siting corrected for odd offsets) → GBRP at the output size with
SWS_BILINEAR | SWS_FULL_CHR_H_INT | SWS_FULL_CHR_H_INP, then interleave to RGB; display rotation moved after the resize; one resize mechanism, old RGB→RGB resizer removed.Open before adopting
Not yet verified, each a possible regression:
A benchmark matrix ({360p, 480p, 1080p, 4K} × {down, up} × {8-bit 420, 10-bit 420, 444, full-range}) and accuracy tests on those sources should pass before this lands.