Skip to content

Add Qwen3.8-27B WebGPU recurrent-state generation example - #590

Closed
RuurdKuiper wants to merge 1 commit into
microsoft:mainfrom
RuurdKuiper:add-qwen38-27b-webgpu-example
Closed

RuurdKuiper wants to merge 1 commit into
microsoft:mainfrom
RuurdKuiper:add-qwen38-27b-webgpu-example

Conversation

@RuurdKuiper

Copy link
Copy Markdown

Adds a standalone Qwen3.8-27B INT4 text-generation example using ONNX Runtime Web's WebGPU EP and JSPI/Blob external-data loading. The example implements the exported graph's ordinary KV and convolution/recurrent state contract, retaining all 128 state tensors on GPU between decoding steps with CPU EP fallback disabled.

Includes a pinned public Hugging Face model pack, a checksum-verifying download script, a minimal UI, measurements and a Chrome smoke test. Model weights (~15.26 GB) remain outside Git and the application build. The local server serves files only. The model is a community conversion from Qwen's Apache-2.0 checkpoint; upstream attribution and conversion notices are included in its model repository. The example source uses Apache-2.0 with its own license file.

Validation: npm ci, npm run build, and the browser smoke test passed on a 48 GiB Apple Silicon Mac with Chrome 154 / Metal 3 WebGPU. Arithmetic returned 156; the longer response generated 32 tokens, with all 128 states on GPU at every step and CPU EP fallback disabled. The browser test blocks external requests. Exact versions and measurements are included in validation.json.

Scope: batch-one contiguous unpadded text. Requires a current Chromium browser with hardware WebGPU, shader-f16 and WebAssembly JSPI, using an ordinary profile. Other hardware, long contexts and model-quality parity remain unvalidated. The example intentionally pins the tested development runtime because this graph and large Blob loader depend on that runtime's capabilities.

Public model: https://huggingface.co/Ruurd/Qwen3.8-27B-ONNX-WebGPU-INT4
Standalone repository: https://github.com/RuurdKuiper/qwen38-27b-onnx-webgpu

This is submitted as a draft for maintainers' feedback on fit, runtime pinning and the large-download requirement; it does not imply official support for this conversion.

@RuurdKuiper RuurdKuiper closed this Oct 6, 2026
@RuurdKuiper
RuurdKuiper deleted the add-qwen38-27b-webgpu-example branch October 6, 2026 15:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant