Skip to content

docs: make the README's adaptive and np.savez recipes work on every layer - #103

Merged
cursor[bot] merged 4 commits into
mainfrom
akshey/readme-recipes-fefb
Sep 28, 2026
Merged

cursor[bot] merged 4 commits into
mainfrom
akshey/readme-recipes-fefb

Conversation

@aksheyd

@aksheyd aksheyd commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Two README recipes failed as written on real models.

Adaptive. On layers with outliers, a block can need more than 8 bits, and adaptive.quantize raises ToleranceTooTightError. That's on purpose since #76, but the README didn't say so, so a loop over a model's layers stopped at the first outlier layer. Its np.std(weights) also failed on PyTorch tensors, since NumPy hands the call to Tensor.std, which doesn't take NumPy's arguments.

  • the line now uses weights.std(), which works on NumPy arrays and PyTorch tensors, including ones that require grad or hold bf16. A plain list needs np.std
  • after the list, a four-line retry with the error's smallest_tolerance, which always passes, since it's half an 8-bit step of the widest block. The README says it loosens every block, not just the one with the outlier
try:
    q = adaptive.quantize(weights, tolerance=0.1 * weights.std())
except ToleranceTooTightError as error:
    q = adaptive.quantize(weights, tolerance=error.smallest_tolerance)

On 1,000,000 Student-t weights with 3 degrees of freedom, a common stand-in for heavy tails, the first call raises and the retry works at 0.18 of a standard deviation.

np.savez. For an adaptive tensor, bits=q.bits saved None, which np.load refuses without allow_pickle=True. The recipe now saves every part and leaves out whichever of bits and block_bits is None, so one recipe works for every kind:

parts = dict(kind=q.kind, shape=q.shape, block=q.block, bits=q.bits, block_bits=q.block_bits,
             codes=q.codes, scales=q.scales, zero_points=q.zero_points, scale=q.scale.name)
np.savez("layer.npz", **{name: part for name, part in parts.items() if part is not None})
q = Quantized.from_parts(**np.load("layer.npz"))

Run verbatim through a real file, it round-trips all 12 combinations of kind and scale type, empty tensors included. The existing test now runs this recipe instead of branching on the kind.

#98 changed the next two bullets, so this branch has main merged in, keeping #98's wording. PyPI keeps each release's README as uploaded, so this is worth having in 0.3.0. Merges cleanly with #101. just lint, just test, and just python pass.

Open in Web Open in Cursor 

…ayer

The adaptive line stopped at the first layer with outliers, where a block needs more than 8 bits and ToleranceTooTightError is raised, and its np.std(weights) failed on PyTorch tensors. It now uses weights.std(), which works on arrays and tensors, and a short retry with the error's smallest_tolerance follows the list, saying it loosens every block. The np.savez recipe saved bits=None for an adaptive tensor, which np.load refuses, so it now leaves out whichever part is None, and the test runs that exact recipe for every kind and scale type.
@aksheyd
aksheyd marked this pull request as ready for review September 28, 2026 04:50
@cursor
cursor Bot merged commit 4d38955 into main Sep 28, 2026
10 checks passed
@cursor
cursor Bot deleted the akshey/readme-recipes-fefb branch September 28, 2026 04:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant