feat: read PyTorch parameters, bf16 tensors, and uint8 tensors for bytes - #106
Merged
Merged
Conversation
torch.load defaults to weights_only=True, which refuses any global it doesn't trust. Pickles called the static method from_bytes, which pickles as builtins.getattr, so add_safe_globals([Quantized]) didn't help, and allowing getattr would defeat weights_only. Quantized(data) now loads the bytes that to_bytes saved, like from_bytes, and pickles call it, so allowing the class is enough. Pickles made through from_bytes still load, since from_bytes stays.
quantize(linear.weight) raised PyTorch's RuntimeError, since numpy.asarray refuses a tensor that requires grad, and a bf16 weight raised "Got unsupported ScalarType BFloat16", since NumPy has no bfloat16. A PyTorch tensor is now detached, and a floating-point one read as float32, which is exact for bf16, before numpy.asarray reads it. torch is looked up in sys.modules, never imported, and NumPy arrays skip the lookup, so they cost what they did. from_bytes, Quantized(data), and from_parts' codes also take any 1-D uint8 array that numpy.asarray reads, like the tensor that safetensors.torch loads.
This was referenced Sep 28, 2026
cursor
Bot
changed the base branch from
akshey/torch-load-pickles-fefb
to
main
September 28, 2026 04:47
…-fefb # Conflicts: # python/src/quantized/methods.rs
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #101, which adds the
Quantized(data)constructor that this also changes. Merge #101 first. CI only runs on PRs intomain, so update this branch once #101 merges.The first calls a PyTorch user makes failed with PyTorch's errors:
quantize(linear.weight)raisedRuntimeError: Can't call numpy() on Tensor that requires grad, since every module parameter requires grad andnumpy.asarrayrefuses those. feat: accept anything np.asarray reads, like PyTorch tensors #64 let this through on purpose, but it's aRuntimeError, soexcept ValueErrormissed itTypeError: Got unsupported ScalarType BFloat16, since NumPy has no bfloat16Quantized.from_bytesrejected the uint8 tensor thatsafetensors.torch.load_filereturns, andfrom_parts(codes=...)did too, thoughscales=took oneNow:
numpy.asarrayreads it. bf16 to float32 is exact, and fp16 and fp64 tensors give the same codes as before. Complex tensors are still rejectedquantize,matmul,dot,refine,alternate, andfrom_parts' scalessys.modules, never imported, since only a program that imported it can pass a tensor. NumPy arrays skip even that, so they cost what they did: adoton 32 values takes 1.05 µs before and afterfrom_bytes,Quantized(data), andfrom_parts'codesalso take any 1-D uint8 array thatnumpy.asarrayreads. Anything else raisesTypeError: data must be bytes, like to_bytes returns, or a 1-D uint8 arrayChecked with torch 2.14 (CPU), with warnings raised as errors: a
Linearweight, and its bf16, fp16, fp64, int, and bool versions, quantize like their NumPy versions.matmultakes a batch that requires grad,refineandadaptive.quantizetake the parameter, andfrom_bytesandfrom_partstake uint8 tensors. The parameter itself is left as it was. The tests use a stand-in for PyTorch's tensor, so CI doesn't need torch.just lint,just test, andjust pythonpass.