x86: SSE2/SSSE3 vec_dot for PTQ1_0 and PQ2_0 - #248
Conversation
|
Tested on an Intel laptop: Core Ultra X7 358H (Panther Lake, 16 cores, hybrid), Windows 11, MSYS2 UCRT64 GCC 16.2, CPU only (
|
| model | base pp64 | PR pp64 | base tg32 | PR tg32 |
|---|---|---|---|---|
| Ternary-Bonsai 1.7B PQ2_0 | 46.57 / 46.62 | 81.09 / 78.57 | 10.85 / 10.74 | 13.99 / 13.27 |
| 2B PTQ1_0 | 25.06 / 25.79 | 56.36 / 55.83 | 7.02 / 6.94 | 9.58 / 9.27 |
Native build (AVX2 + AVX-VNNI CPU): the SSSE3 PTQ1_0 kernel is also what an AVX2 build picks up:
| model | base pp64 | PR pp64 | base tg32 | PR tg32 |
|---|---|---|---|---|
| 2B PTQ1_0 | 27.77 / 27.11 | 59.46 / 61.72 | 7.40 / 7.15 | 9.25 / 9.43 |
| Bonsai 2 27B PTQ1_0 | — | 3.87 / 3.91 | — | 1.83 / 1.84 |
For the 27B, the generic kernel on current prism measures 1.70 / 1.67 pp64 and 1.05 / 1.04 tg32. For comparison on the same box, #250's AVX2+VNNI kernel does 4.65 / 2.00, so on AVX2 hardware #250 is faster, and this PR's value is the pre-AVX2 tiers.
Greedy 128-token output (2B PTQ1_0, temp 0) is byte-identical to its base in the native build.
Tested with Claude Code.
PTQ1_0 had no x86 SIMD path at all - arch-fallback.h aliased it to the
generic C implementation, so every x86 CPU decoded base-3 trits one
weight at a time. PQ2_0 had an x86 kernel, but it is gated on VNNI
(dpbusd), so everything below Ice Lake / Zen 4 fell through to the
scalar loop.
Add a 128-bit path for both, in two tiers:
SSE2 - trit decode by multiply-and-shift, sign-extended _mm_madd_epi16
for the dot. Works on any x86-64.
SSSE3 - _mm_maddubs_epi16 for the unsigned x signed multiply-add, which
halves the ALU work in the inner loop.
The ternary offset (codes are 0..2, values -1..1) is folded into the
accumulator as a separate maddubs against ones rather than subtracted
per element, which avoids negating -128 activations.
Also add GGML_SSSE3 as a build variant so the dispatcher can score it,
matching how the other x86 feature tiers are handled.
Measured on an i5-12600K with AVX disabled at compile time, 8 threads.
Isolated A/B: same binary, only libggml-cpu.so swapped, runs alternated
to cancel background load, two pairs per configuration.
Ternary Bonsai 1.7B PQ2_0 pp128 tg32
scalar 15.33/15.39 9.49/9.62
SSE2 42.91/43.23 25.20/25.35
SSSE3 59.49/58.98 32.07/31.42
Ternary Bonsai 2 27B PTQ1_0 pp64 tg16
scalar 0.71/0.72 0.60/0.62
SSSE3 2.66/2.66 2.11/2.10
3.3x tg on PQ2_0 and 3.5x on PTQ1_0 over scalar.
Also verified on a dual Xeon E5645 (Westmere, SSE4.2 ceiling, no AVX):
builds clean and runs both formats, which is the hardware tier this is
mainly for.
The existing test_vec_dot_q() compares against a tolerance, which is the right call for lossy formats but too loose to catch an unpack bug in a ternary kernel: a wrong trit ordering still lands inside the error bound. Add a check that is exact instead. Weights and activations are driven over 256 bit patterns per format, and every scale is a power of two, so dequantize-then-dot and the packed vec_dot must agree bit for bit. Any mismatch in trit order, plane split, or the -1 offset shows up as a hard failure rather than a slightly larger error. Covers both block counts (1 and 3) so the multi-block loop is exercised. Passes on SSE2, SSSE3 and AVX2 builds, and on a dual Xeon E5645 where neither AVX path is available.
PQ2_0 moved to Q8_K activations, so ggml_vec_dot_pq2_0_q8_0 is no longer
what the traits table dispatches and the SSE tier added earlier in this
branch was dead code for normal inference. Add the same 128-bit tier to
ggml_vec_dot_pq2_0_q8_K instead, below the AVX2 path.
Structure mirrors the AVX2 kernel at half the width: the four sub-block
dots of a 128-weight block accumulate in int32 and the block scale is
applied once, rather than per sub-block. The existing unpack already
emits byte b's bit-pair j at element 4b+j, which is the order the Q8_K
activations are in, so no repermutation is needed.
CPUs with AVX2 keep the upstream kernel; this only changes what runs
below it, where PQ2_0 was falling back to the scalar loop.
Also fix the ternary test to build the right activation type per format
(Q8_K for PQ2_0, Q8_0 for PTQ1_0) and expand Q8_K by hand, since it is
an activation-only type with no to_float. Verified the exact comparison
still catches errors: injecting an off-by-one in the accumulator fails
all 512 PQ2_0 cases.
i5-12600K, AVX disabled at compile time, 8 threads, isolated A/B
(same binary, only libggml-cpu.so swapped, alternated, two pairs):
Ternary Bonsai 1.7B PQ2_0 pp128 tg32
scalar 15.04/15.12 9.97/9.83
SSSE3 83.81/82.98 44.91/45.21
4.5x tg. Higher than the 3.3x the Q8_0 path gave, because one activation
scale per 256 removes the per-sub-block float chain.
test-quantize-fns passes on SSE2, SSSE3 and AVX2 builds.
55158e9 to
aed0a63
Compare
|
Thanks for the thorough review, and for catching this on real hardware with the ymm-instruction check - that is exactly the kind of verification I would not have thought to add myself. Q8_K activation typeAlready fixed, just before your review: The three-way overlap on
|
|
Re-tested
Performance, SSE-only builds (
That's lower than your 4.5x tg on the i5-12600K. This is a hybrid Panther Lake laptop part at 8 threads, so I wouldn't read much into the difference. Native AVX2 build: PQ2_0 takes the unchanged upstream AVX2 kernel (the diff only adds On landing order, your proposal (#250 for AVX2/VNNI, this PR for pre-AVX2) matches what I measured: #250's kernel is the fastest on AVX2 hardware under both GCC and MSVC. The call belongs to the authors and maintainers. Tested with Claude Code. |
|
Thank you for the re-test, and for going back to a second device rather than just re-confirming on the same box - that is a better check than I would have run myself. Good to have both findings closed: clean merge onto current And thank you for confirming the landing-order proposal with actual measurements instead of just going along with it - "#250 fastest on AVX2 under both GCC and MSVC" is a real answer to a question I could only guess at from here. That is genuinely useful to whoever ends up doing the three-way merge. Appreciate the thoroughness across both this and #235 - the amount of hardware and time you have put into testing this work is well beyond what I could have verified alone. |
Overview
Adds a 128-bit SIMD
vec_dotfor the two ternary formats on x86, in two tiers (SSE2 and SSSE3).Before this PR:
arch-fallback.haliased it to the generic C implementation, so every x86 CPU - including current ones - decoded base-3 trits one weight at a time.The two tiers:
_mm_madd_epi16for the dot. Works on any x86-64._mm_maddubs_epi16for the unsigned x signed multiply-add, which halves the ALU work in the inner loop.The ternary offset (codes are 0..2, values -1..1) is folded into the accumulator as a separate
maddubsagainst ones rather than subtracted per element, which avoids negating -128 activations.Also adds
GGML_SSSE3as a build variant so the runtime dispatcher can score it, matching how the other x86 feature tiers are handled.Results
Measured on an i5-12600K with AVX disabled at compile time, 8 threads. Isolated A/B: same binary, only
libggml-cpu.soswapped, runs alternated to cancel background load, two pairs per configuration.Ternary Bonsai 1.7B PQ2_0 (Q8_K activation path)
Ternary Bonsai 2 27B PTQ1_0
4.5x tg on PQ2_0 and 3.5x on PTQ1_0 over scalar.
Correctness
test-quantize-fnsgains an exact check for both formats. The existingtest_vec_dot_q()compares against a tolerance, which is right for lossy formats but too loose here: a wrong trit ordering still lands inside the error bound. The new check drives 256 bit patterns per format with power-of-two scales, so dequantize-then-dot and the packedvec_dotmust agree bit for bit. Any mismatch in trit order, plane split, or the -1 offset is a hard failure.Passes on SSE2, SSSE3 and AVX2 builds.
Separately, and not part of this PR, I ran the kernels against the generic scalar reference under a property-based harness (theft) plus an exhaustive edge sweep: ~124k unique random inputs per format per ISA tier, and every weight byte 0-255 crossed with the activation extremes (-128, -1, 0, 1, 127), including out-of-range trit bytes (>242 in
qs, >80 inqh) that a corrupt file could carry. No mismatches on any tier. I validated the harness itself by injecting a swapped-activation-half bug, which it caught and shrank to a single non-zero weight byte.Additional information
Also built and run on a dual Xeon E5645 (Westmere, 2010, SSE4.2 ceiling, no AVX at all), which is the hardware tier this mainly serves. Both formats build clean and run there;
-march=nativeconfirms-mavx/-mavx2disabled, so it exercises the SSSE3 path.Caveat on the numbers above: the 12600K figures come from a modern CPU with AVX disabled at compile time, so absolute throughput is higher than period hardware would give. The ratios should hold, since both sides of each A/B face the same memory system. I am arranging a quiet pre-AVX box for clean absolute numbers and will post them here when I have them.
This is independent of #235 (SYCL backend) - no shared files, and it branches from
prismrather than from that work.