Skip to content

x86: SSE2/SSSE3 vec_dot for PTQ1_0 and PQ2_0 - #248

Merged
bri-prism merged 3 commits into
PrismML-Eng:prismfrom
kiljoy001:x86-ternary-sse
Sep 25, 2026
Merged

bri-prism merged 3 commits into
PrismML-Eng:prismfrom
kiljoy001:x86-ternary-sse

Conversation

@kiljoy001

@kiljoy001 kiljoy001 commented Sep 23, 2026 •

Copy link
Copy Markdown

Overview

Adds a 128-bit SIMD vec_dot for the two ternary formats on x86, in two tiers (SSE2 and SSSE3).

Before this PR:

  • PTQ1_0 had no x86 SIMD path at all. arch-fallback.h aliased it to the generic C implementation, so every x86 CPU - including current ones - decoded base-3 trits one weight at a time.
  • PQ2_0 now dots against Q8_K activations, and that kernel is gated on AVX2, so everything below Haswell fell through to the scalar loop.

The two tiers:

  • SSE2 - trit decode by multiply-and-shift, sign-extended _mm_madd_epi16 for the dot. Works on any x86-64.
  • SSSE3 - _mm_maddubs_epi16 for the unsigned x signed multiply-add, which halves the ALU work in the inner loop.

The ternary offset (codes are 0..2, values -1..1) is folded into the accumulator as a separate maddubs against ones rather than subtracted per element, which avoids negating -128 activations.

Also adds GGML_SSSE3 as a build variant so the runtime dispatcher can score it, matching how the other x86 feature tiers are handled.

Results

Measured on an i5-12600K with AVX disabled at compile time, 8 threads. Isolated A/B: same binary, only libggml-cpu.so swapped, runs alternated to cancel background load, two pairs per configuration.

Ternary Bonsai 1.7B PQ2_0 (Q8_K activation path)

build pp128 tg32
scalar 15.04 / 15.12 9.97 / 9.83
SSSE3 83.81 / 82.98 44.91 / 45.21

Ternary Bonsai 2 27B PTQ1_0

build pp64 tg16
scalar 0.71 / 0.72 0.60 / 0.62
SSSE3 2.66 / 2.66 2.11 / 2.10

4.5x tg on PQ2_0 and 3.5x on PTQ1_0 over scalar.

Correctness

test-quantize-fns gains an exact check for both formats. The existing test_vec_dot_q() compares against a tolerance, which is right for lossy formats but too loose here: a wrong trit ordering still lands inside the error bound. The new check drives 256 bit patterns per format with power-of-two scales, so dequantize-then-dot and the packed vec_dot must agree bit for bit. Any mismatch in trit order, plane split, or the -1 offset is a hard failure.

Passes on SSE2, SSSE3 and AVX2 builds.

Separately, and not part of this PR, I ran the kernels against the generic scalar reference under a property-based harness (theft) plus an exhaustive edge sweep: ~124k unique random inputs per format per ISA tier, and every weight byte 0-255 crossed with the activation extremes (-128, -1, 0, 1, 127), including out-of-range trit bytes (>242 in qs, >80 in qh) that a corrupt file could carry. No mismatches on any tier. I validated the harness itself by injecting a swapped-activation-half bug, which it caught and shrank to a single non-zero weight byte.

Additional information

Also built and run on a dual Xeon E5645 (Westmere, 2010, SSE4.2 ceiling, no AVX at all), which is the hardware tier this mainly serves. Both formats build clean and run there; -march=native confirms -mavx/-mavx2 disabled, so it exercises the SSSE3 path.

Caveat on the numbers above: the 12600K figures come from a modern CPU with AVX disabled at compile time, so absolute throughput is higher than period hardware would give. The ratios should hold, since both sides of each A/B face the same memory system. I am arranging a quiet pre-AVX box for clean absolute numbers and will post them here when I have them.

This is independent of #235 (SYCL backend) - no shared files, and it branches from prism rather than from that work.

@bri-prism

Copy link
Copy Markdown
Collaborator

Tested on an Intel laptop: Core Ultra X7 358H (Panther Lake, 16 cores, hybrid), Windows 11, MSYS2 UCRT64 GCC 16.2, CPU only (-ngl 0).

⚠️ Conflicts with current prism

The merge onto prism @ 3b19c377d conflicts in ggml/src/ggml-cpu/arch/x86/quants.c. prism has since rewritten ggml_vec_dot_pq2_0_q8_0 (164c337, "AVX2/AVX-VNNI kernels for PQ2_0 + Q8_K activations"), so on tip PQ2_0 is no longer gated on VNNI. The PQ2_0 SSE tiers still matter for pre-AVX2 CPUs. Two things will need updating in the rebase:

Results on the PR's own base (9a9394a89, the merge-base)

Both arms are ABAB'd: llama-bench -p 64 -n 32 -r 2 -t 16, two rounds with the order reversed (round 1 / round 2, t/s).

SSE-only build (GGML_NATIVE=OFF, AVX/AVX2/FMA/F16C/AVX_VNNI off, GGML_SSE42=ON GGML_SSSE3=ON; I checked the binary contains 0 ymm instructions). test-quantize-fns passes, including the new exact checks:

model base pp64 PR pp64 base tg32 PR tg32
Ternary-Bonsai 1.7B PQ2_0 46.57 / 46.62 81.09 / 78.57 10.85 / 10.74 13.99 / 13.27
2B PTQ1_0 25.06 / 25.79 56.36 / 55.83 7.02 / 6.94 9.58 / 9.27

Native build (AVX2 + AVX-VNNI CPU): the SSSE3 PTQ1_0 kernel is also what an AVX2 build picks up:

model base pp64 PR pp64 base tg32 PR tg32
2B PTQ1_0 27.77 / 27.11 59.46 / 61.72 7.40 / 7.15 9.25 / 9.43
Bonsai 2 27B PTQ1_0 — 3.87 / 3.91 — 1.83 / 1.84

For the 27B, the generic kernel on current prism measures 1.70 / 1.67 pp64 and 1.05 / 1.04 tg32. For comparison on the same box, #250's AVX2+VNNI kernel does 4.65 / 2.00, so on AVX2 hardware #250 is faster, and this PR's value is the pre-AVX2 tiers.

Greedy 128-token output (2B PTQ1_0, temp 0) is byte-identical to its base in the native build.

Tested with Claude Code.

PTQ1_0 had no x86 SIMD path at all - arch-fallback.h aliased it to the
generic C implementation, so every x86 CPU decoded base-3 trits one
weight at a time. PQ2_0 had an x86 kernel, but it is gated on VNNI
(dpbusd), so everything below Ice Lake / Zen 4 fell through to the
scalar loop.

Add a 128-bit path for both, in two tiers:

  SSE2  - trit decode by multiply-and-shift, sign-extended _mm_madd_epi16
          for the dot. Works on any x86-64.
  SSSE3 - _mm_maddubs_epi16 for the unsigned x signed multiply-add, which
          halves the ALU work in the inner loop.

The ternary offset (codes are 0..2, values -1..1) is folded into the
accumulator as a separate maddubs against ones rather than subtracted
per element, which avoids negating -128 activations.

Also add GGML_SSSE3 as a build variant so the dispatcher can score it,
matching how the other x86 feature tiers are handled.

Measured on an i5-12600K with AVX disabled at compile time, 8 threads.
Isolated A/B: same binary, only libggml-cpu.so swapped, runs alternated
to cancel background load, two pairs per configuration.

  Ternary Bonsai 1.7B PQ2_0      pp128          tg32
    scalar                       15.33/15.39    9.49/9.62
    SSE2                         42.91/43.23   25.20/25.35
    SSSE3                        59.49/58.98   32.07/31.42

  Ternary Bonsai 2 27B PTQ1_0    pp64           tg16
    scalar                        0.71/0.72     0.60/0.62
    SSSE3                         2.66/2.66     2.11/2.10

3.3x tg on PQ2_0 and 3.5x on PTQ1_0 over scalar.

Also verified on a dual Xeon E5645 (Westmere, SSE4.2 ceiling, no AVX):
builds clean and runs both formats, which is the hardware tier this is
mainly for.
The existing test_vec_dot_q() compares against a tolerance, which is the
right call for lossy formats but too loose to catch an unpack bug in a
ternary kernel: a wrong trit ordering still lands inside the error bound.

Add a check that is exact instead. Weights and activations are driven
over 256 bit patterns per format, and every scale is a power of two, so
dequantize-then-dot and the packed vec_dot must agree bit for bit. Any
mismatch in trit order, plane split, or the -1 offset shows up as a hard
failure rather than a slightly larger error.

Covers both block counts (1 and 3) so the multi-block loop is exercised.

Passes on SSE2, SSSE3 and AVX2 builds, and on a dual Xeon E5645 where
neither AVX path is available.
PQ2_0 moved to Q8_K activations, so ggml_vec_dot_pq2_0_q8_0 is no longer
what the traits table dispatches and the SSE tier added earlier in this
branch was dead code for normal inference. Add the same 128-bit tier to
ggml_vec_dot_pq2_0_q8_K instead, below the AVX2 path.

Structure mirrors the AVX2 kernel at half the width: the four sub-block
dots of a 128-weight block accumulate in int32 and the block scale is
applied once, rather than per sub-block. The existing unpack already
emits byte b's bit-pair j at element 4b+j, which is the order the Q8_K
activations are in, so no repermutation is needed.

CPUs with AVX2 keep the upstream kernel; this only changes what runs
below it, where PQ2_0 was falling back to the scalar loop.

Also fix the ternary test to build the right activation type per format
(Q8_K for PQ2_0, Q8_0 for PTQ1_0) and expand Q8_K by hand, since it is
an activation-only type with no to_float. Verified the exact comparison
still catches errors: injecting an off-by-one in the accumulator fails
all 512 PQ2_0 cases.

i5-12600K, AVX disabled at compile time, 8 threads, isolated A/B
(same binary, only libggml-cpu.so swapped, alternated, two pairs):

  Ternary Bonsai 1.7B PQ2_0    pp128          tg32
    scalar                     15.04/15.12    9.97/9.83
    SSSE3                      83.81/82.98   44.91/45.21

4.5x tg. Higher than the 3.3x the Q8_0 path gave, because one activation
scale per 256 removes the per-sub-block float chain.

test-quantize-fns passes on SSE2, SSSE3 and AVX2 builds.
@kiljoy001

Copy link
Copy Markdown
Author

Thanks for the thorough review, and for catching this on real hardware with the ymm-instruction check - that is exactly the kind of verification I would not have thought to add myself.

Q8_K activation type

Already fixed, just before your review: aed0a6345 derives the activation type from the format instead of hard-coding block_q8_0, so PQ2_0 builds its Q8_K blocks and PTQ1_0 keeps Q8_0. Current head passes both halves of the exact check on tip. Good to know the PTQ1_0 half was already useful for testing #181/#250 in the meantime.

The three-way overlap on ggml_vec_dot_ptq1_0_q8_0

You are right that this collides with #181 and #250, and I do not think any of the three PRs currently accounts for the other two - each patches the function as if it were the only one touching it. Your own numbers make the shape of the fix clear: #250's AVX2+VNNI path beats this PR's SSE tiers on AVX2 hardware (4.65 vs 3.87-3.91 pp64 on the 27B), which is expected, since this PR was written for pre-AVX2 CPUs and was never meant to compete above that line.

So the tiers are complementary, not competing, and the eventual function should be roughly:

#if AVX2/VNNI     (whichever of #181/#250 lands)
    ...
#elif SSE2        (this PR)
    ...
#else
    ggml_vec_dot_ptq1_0_q8_0_generic(...)
#endif

I would rather not guess at that nesting unilaterally since I do not own #181 or #250. Flagging it here on record, and cross-posting to both so their authors see it before assuming they are alone in that function. Whichever of the three lands first, the next one should expect a conflict in ggml_vec_dot_ptq1_0_q8_0 and arch-fallback.h, and it is probably worth the three of us (or a maintainer) agreeing on landing order rather than each rebasing independently and hoping.

For what it is worth, I would suggest #250 first if its AVX2+VNNI numbers hold up, since that is the tier most machines actually have, then this PR rebased under it for the pre-AVX2 case, then #181 reconciled or dropped depending on how much AVX-VNNI-specific work it does beyond what #250 covers - though that is a call for you three and whoever is merging, not something I should decide from here.

Thanks again for the numbers - byte-identical greedy output on the native build is the check I would have wanted to run myself.

@bri-prism

Copy link
Copy Markdown
Collaborator

Re-tested aed0a6345 (rebased) on the same Core Ultra X7 358H, Windows 11, GCC 16.2. Base = prism @ 0324c6652. Both issues from my review are resolved.

  • Merge: clean onto current prism; the conflict is gone.
  • Test: the exact ternary check (your new test-quantize-fns.cpp) now passes with 0 failures in all four builds: tip + PR, each native (AVX2 + AVX-VNNI) and SSE-only. Tip with only the new test file dropped in also passes, so the Q8_K false failures are gone.
  • SSE-only builds contain 0 ymm instructions (checked).

Performance, SSE-only builds (GGML_NATIVE=OFF, AVX/AVX2/FMA/F16C/AVX_VNNI off, SSE42 + SSSE3 on), llama-bench -ngl 0 -t 8 -p 64 -n 32 -r 2, two rounds with the order reversed, 0 other llama processes (t/s):

model tip pp64 PR pp64 tip tg32 PR tg32
Ternary-Bonsai 1.7B PQ2_0 (now via the Q8_K path) 21.19 / 18.02 87.19 / 78.02 (~4.2x) 9.92 / 9.13 20.15 / 19.80 (~2.1x)
2B PTQ1_0 21.49 / 17.74 43.72 / 42.83 (~2.2x) 8.18 / 7.67 12.82 / 12.61 (~1.6x)

That's lower than your 4.5x tg on the i5-12600K. This is a hybrid Panther Lake laptop part at 8 threads, so I wouldn't read much into the difference.

Native AVX2 build: PQ2_0 takes the unchanged upstream AVX2 kernel (the diff only adds #elif tiers below it), and the numbers are the same within noise: pp64 289 / 276 vs tip 300 / 259, tg32 21.6 / 23.2 vs 24.5 / 24.0.

On landing order, your proposal (#250 for AVX2/VNNI, this PR for pre-AVX2) matches what I measured: #250's kernel is the fastest on AVX2 hardware under both GCC and MSVC. The call belongs to the authors and maintainers.

Tested with Claude Code.

@kiljoy001

Copy link
Copy Markdown
Author

Thank you for the re-test, and for going back to a second device rather than just re-confirming on the same box - that is a better check than I would have run myself.

Good to have both findings closed: clean merge onto current prism, 0 failures across all four configurations including tip-with-just-the-test-file, and 0 ymm instructions confirmed again in the SSE-only build. The Panther Lake numbers being lower than the 12600K's is exactly the kind of hardware variance I would expect and not something I would read into - glad you called that out explicitly rather than letting it sit as an unexplained gap.

And thank you for confirming the landing-order proposal with actual measurements instead of just going along with it - "#250 fastest on AVX2 under both GCC and MSVC" is a real answer to a question I could only guess at from here. That is genuinely useful to whoever ends up doing the three-way merge.

Appreciate the thoroughness across both this and #235 - the amount of hardware and time you have put into testing this work is well beyond what I could have verified alone.

@bri-prism
bri-prism merged commit adfffbe into PrismML-Eng:prism Sep 25, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation ggml testing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants