Skip to content

Benches / catching up #3

Description

@aksheyd

Benches / catching up

right now compare is reconstruct + matmul MSE on random matrices. throughput is pack/unpack ns/value vs candle Q4_0 / Q8_0.

that's not what people actually measure:

  • quality: perplexity (wikitext-2, llama-perplexity) and task scores (lm-eval — mmlu, gsm8k, humaneval). math/code drop first
  • speed: tokens/s (decode) + prompt tokens/s. decode is basically memory bandwidth. llama-bench, vllm, mlx

QAT doesn't make PTQ pointless. QAT = fake the rounding while training (expensive, needs data). we are PTQ = compress a finished model in a few minutes. different jobs.

6/7 as implemented are the simple versions of stuff that already exists (k-quants / awq / gptq). more bits everywhere is just a bigger file, not a method.

  • pick bits / scales from real activations (not just min-max of the weights vs a tolerance)
  • measure error the next layer feels, not only block MSE
  • fused packed matmul so we don't materialize f32
  • report wikitext (or gsm8k) on a small model instead of 1024x1024 random matmul mse

we're at textbook Q4_0 / Q8_0. field is Q4_K_M / AWQ / GPTQ. we beat candle on pack/unpack on this mac, not on quality, not on gemm.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions