Repository navigation
Add Qwen3.8-27B WebGPU recurrent-state generation example - #590
Closed
RuurdKuiper wants to merge 1 commit into
Closed
RuurdKuiper wants to merge 1 commit into
RuurdKuiper wants to merge 1 commit into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds a standalone Qwen3.8-27B INT4 text-generation example using ONNX Runtime Web's WebGPU EP and JSPI/Blob external-data loading. The example implements the exported graph's ordinary KV and convolution/recurrent state contract, retaining all 128 state tensors on GPU between decoding steps with CPU EP fallback disabled.
Includes a pinned public Hugging Face model pack, a checksum-verifying download script, a minimal UI, measurements and a Chrome smoke test. Model weights (~15.26 GB) remain outside Git and the application build. The local server serves files only. The model is a community conversion from Qwen's Apache-2.0 checkpoint; upstream attribution and conversion notices are included in its model repository. The example source uses Apache-2.0 with its own license file.
Validation:
npm ci,npm run build, and the browser smoke test passed on a 48 GiB Apple Silicon Mac with Chrome 154 / Metal 3 WebGPU. Arithmetic returned 156; the longer response generated 32 tokens, with all 128 states on GPU at every step and CPU EP fallback disabled. The browser test blocks external requests. Exact versions and measurements are included invalidation.json.Scope: batch-one contiguous unpadded text. Requires a current Chromium browser with hardware WebGPU, shader-f16 and WebAssembly JSPI, using an ordinary profile. Other hardware, long contexts and model-quality parity remain unvalidated. The example intentionally pins the tested development runtime because this graph and large Blob loader depend on that runtime's capabilities.
Public model: https://huggingface.co/Ruurd/Qwen3.8-27B-ONNX-WebGPU-INT4
Standalone repository: https://github.com/RuurdKuiper/qwen38-27b-onnx-webgpu
This is submitted as a draft for maintainers' feedback on fit, runtime pinning and the large-download requirement; it does not imply official support for this conversion.