Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 14 additions & 4 deletions python/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,18 +25,28 @@ the scales count toward the size: 4-bit codes with one f16 scale per 32 values c
the other schemes return the same `Quantized` type:

- `asymmetric.quantize(weights, bits=8, block=32)` adds a zero-point per block, for values that aren't centered on zero
- `adaptive.quantize(weights, tolerance=0.1 * np.std(weights))` gives each block the fewest bits, from 2 to 8, that round every weight within `tolerance`, in the weights' own units. a tenth of their standard deviation gives about 5 bits a block
- `adaptive.quantize(weights, tolerance=0.1 * weights.std())` gives each block the fewest bits, from 2 to 8, that round every weight within `tolerance`, in the weights' own units. a tenth of their standard deviation gives about 5 bits a block. for a list, use `np.std(weights)`
- `learned.refine(q, weights)` refits each block's scale, and its zero-point if it has one, to lower the mean squared error. it changes `q` in place, so call `q.copy()` first to keep the original
- `learned.alternate(q, weights)` refits too, then rounds each value to the nearest code on its block's new line, and repeats until no code moves. it also changes `q` in place. both can raise the worst error past an adaptive tensor's tolerance
- `Scheme.Q4_32.quantize(weights)` picks a scheme at run time

a block with outliers can need more than 8 bits, which raises `ToleranceTooTightError`. retrying with its `smallest_tolerance` works, but loosens every block, not just that one:

```python
try:
q = adaptive.quantize(weights, tolerance=0.1 * weights.std())
except ToleranceTooTightError as error:
q = adaptive.quantize(weights, tolerance=error.smallest_tolerance)
```

quantized values can be pickled, and compared with `==`. `q.to_bytes()` saves one as bytes, in the same format as the rust crate, and `Quantized.from_bytes(data)` loads it back. to keep it in an `np.savez` or safetensors file, store `np.frombuffer(q.to_bytes(), np.uint8)`.

to save its parts as plain arrays instead, like with `np.savez`, pass them back by name to `Quantized.from_parts`. an adaptive tensor keeps `block_bits` instead of `bits`:
to save its parts as plain arrays instead, like with `np.savez`, pass them back by name to `Quantized.from_parts`. leave out the one that's `None`: `bits` for an adaptive tensor, or `block_bits` for the others:

```python
np.savez("layer.npz", kind=q.kind, shape=q.shape, block=q.block, bits=q.bits,
codes=q.codes, scales=q.scales, zero_points=q.zero_points, scale=q.scale.name)
parts = dict(kind=q.kind, shape=q.shape, block=q.block, bits=q.bits, block_bits=q.block_bits,
codes=q.codes, scales=q.scales, zero_points=q.zero_points, scale=q.scale.name)
np.savez("layer.npz", **{name: part for name, part in parts.items() if part is not None})
q = Quantized.from_parts(**np.load("layer.npz"))
```

Expand Down
9 changes: 4 additions & 5 deletions python/tests/test_quantize.py
Original file line number Diff line number Diff line change
Expand Up @@ -574,16 +574,15 @@ def test_from_parts_rebuilds_parts_saved_with_numpy_as_the_readme_says():
"kind": quantized.kind,
"shape": quantized.shape,
"block": quantized.block,
"bits": quantized.bits,
"block_bits": quantized.block_bits,
"codes": quantized.codes,
"scales": quantized.scales,
"zero_points": quantized.zero_points,
"scale": quantized.scale.name,
}
if quantized.kind == "adaptive":
parts["block_bits"] = quantized.block_bits
else:
parts["bits"] = quantized.bits
rebuilt = Quantized.from_parts(**saved_and_loaded_with_numpy(parts))
saved = {name: part for name, part in parts.items() if part is not None}
rebuilt = Quantized.from_parts(**saved_and_loaded_with_numpy(saved))
assert rebuilt == quantized


Expand Down
Loading