docs: recommend f32 scales for asymmetric codes above about 10 bits - #107
Merged
Merged
Conversation
f16 and bf16 store each block's zero-point too, and their rounding grows with the bit width, while the README and docstrings offered 2 to 16 bits with any scale type. On normal weights, asymmetric bf16's worst error is 2.3 times f32's at 11 bits and 71 times at 16, and f16's is 7 times at 16. Symmetric codes are unaffected. The README, asymmetric.quantize, and Scale now say to use f32 scales there.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The README and the
asymmetric.quantizedocstring offered 2 to 16 bits with any scale type, with no caveat. But f16 and bf16 store each block's zero-point too, and at high bit widths their rounding of it outgrows the codes' own rounding.Worst error relative to f32 scales, on 65,536 normal weights with blocks of 32:
On blocks far from zero, like LayerNorm gains around 1, it starts sooner: bf16 is 3× worse at 8 bits and f16 1.5× at 10.
asymmetric.quantizeasymmetric.quantizedocstring says the same, and that blocks far from zero hit it sooner, as the Rustasymmetricdocs already sayScaledocstring says it tooSymmetric codes are unaffected, so
quantize's docstring stays as it is. Adaptive codes stop at 8 bits, and its docstring already covers blocks far from zero.PyPI keeps each release's README as uploaded, so this is worth having in 0.3.0. Merges cleanly with #101, #103, and #106.
just lint,just test, andjust pythonpass.