Benches / catching up
right now compare is reconstruct + matmul MSE on random matrices. throughput is pack/unpack ns/value vs candle Q4_0 / Q8_0.
that's not what people actually measure:
- quality: perplexity (wikitext-2, llama-perplexity) and task scores (lm-eval — mmlu, gsm8k, humaneval). math/code drop first
- speed: tokens/s (decode) + prompt tokens/s. decode is basically memory bandwidth. llama-bench, vllm, mlx
QAT doesn't make PTQ pointless. QAT = fake the rounding while training (expensive, needs data). we are PTQ = compress a finished model in a few minutes. different jobs.
6/7 as implemented are the simple versions of stuff that already exists (k-quants / awq / gptq). more bits everywhere is just a bigger file, not a method.
we're at textbook Q4_0 / Q8_0. field is Q4_K_M / AWQ / GPTQ. we beat candle on pack/unpack on this mac, not on quality, not on gemm.
Benches / catching up
right now
compareis reconstruct + matmul MSE on random matrices.throughputis pack/unpack ns/value vs candle Q4_0 / Q8_0.that's not what people actually measure:
QAT doesn't make PTQ pointless. QAT = fake the rounding while training (expensive, needs data). we are PTQ = compress a finished model in a few minutes. different jobs.
6/7 as implemented are the simple versions of stuff that already exists (k-quants / awq / gptq). more bits everywhere is just a bigger file, not a method.
we're at textbook Q4_0 / Q8_0. field is Q4_K_M / AWQ / GPTQ. we beat candle on pack/unpack on this mac, not on quality, not on gemm.