Conversation
… graph The qwen35 MTP draft graph looks the draft token's embedding up with a raw ggml_get_rows on model.tok_embd. Models that fold a Hadamard rotation into their weights list token_embd.weight in prism.hadamard.inverse_weight_names and store its rows rotated; the main graph undoes that in llm_graph_context::build_inp_embd(), the MTP graph did not, so llama_verify_hadamard_graph rejects the draft context. Reproduction: take a Hadamard-folded qwen35 file with a nextn block, for example Ternary-Bonsai-2-27B-PTQ1_0 with the blk.64 tensors of a Qwen 3.8 27B GGUF appended and qwen35.nextn_predict_layers = 1 (tools in github.com/sudoingX/bonsai2-small-gpu), and start llama-server -m model.gguf -ngl 99 -fa on -c 32768 --spec-type draft-mtp --spec-draft-n-max 1 Before this change: llama_verify_hadamard_graph: latent lookup 'mtp_tok_embd-64' consumed by op=RMS_NORM name='norm-64' src0 hint=0 llama_init_from_model: failed to initialize the context: Hadamard-latent table 'token_embd.weight' is read without the inverse transform common_speculative_init_result: failed to create MTP context After: the draft context is created and the head drafts with the trunk's own embedding table (acceptance 0.85 to 0.94 on Python, 0.45 to 0.55 on prose, RTX 3060 12GB), which saves the 682 MiB copy of the donor embedding table that was needed to work around it. Models without prism.hadamard keys take the same path as before.
|
Cross-reference: #210 adds If #210 lands first, this PR becomes a one-line change at |
bri-prism
left a comment
There was a problem hiding this comment.
Agent review: posted by the maintainer's coding agent at their request.
No findings in the embedding inverse-transform change. The translation unit passes a Clang C++17 syntax check against same-base headers.
This is functionally the same fix as #205, with an additional include. Please consolidate the two PRs so only one implementation lands. This was a source/syntax review, not a model-backed MTP execution.
Reviewed commit: 1dc4a579b42905007dba3b3ef42e69fed57a9056.
| // a Hadamard-latent embedding table (prism.hadamard.inverse_weight_names) | ||
| // stores rotated rows; restore the primal basis right after the lookup, | ||
| // the same way llm_graph_context::build_inp_embd() does, so an MTP head | ||
| // that shares the trunk's token_embd reads standard-basis embeddings |
Same fix as #205 — proposing we land one of themThanks for this, and for the cross-reference to #210 — that is the part I would have missed. I compared the two patches line by line. The executable change is identical: the same The maintainer has asked for one of the two to land. I have proposed #205 as the landing point, for the boring reason that it already carries two independent reproductions on unrelated stacks plus the Ada measurements, so no evidence has to be gathered again — not because the change here is any worse. It is not. Your RTX 3060 numbers are better evidence than anything on #205 for the low-VRAM case, so I am folding them in there with attribution:
If you would rather #217 be the landing point, say so and I will move everything over instead. Absent that, I will treat #205 as the one and ask the maintainers to take it. |
|
Agreed, land #205. It has the reproductions, and the change is the same three ops either way. I will close this one when #205 is in, or sooner if a maintainer prefers the queue clean. @renovys opened #230 this morning with the same fix, so that is a third copy worth folding into the same decision. Thanks for carrying the numbers over. Two corrections so the ones on #205 are the current ones, both measured on an RTX 3060 12GB with the branch that carries this fix plus the PTQ1_0 mat-vec of #218, greedy,
That is the case for the fix in one line. Without it the only way to run Acceptance is unchanged by the fix, as expected, since it only corrects what the draft graph reads: 0.92 on Python, 0.79 on bash, 0.56 on prose at a fresh context, 0.65 to 0.73 on a long document between 18K and 120K tokens. I have the card and both merged files here, so if #205 needs anything rechecked on 12GB, ask and I will run it. |
|
#205 landed in |
|
Validated this on a real Hadamard-folded qwen35 with a grafted MTP head and it works — thanks for the fix. Cherry-picked 1dc4a57 onto b10687 (5d80cff) and ran Ternary-Bonsai-2-27B-Uncensored-Gaston (PQ2_0 and PTQ1_0) with a blk.64 nextn head grafted from a Qwen3.8-27B donor. After: the draft context builds and --spec-type draft-mtp runs. RTX 3090, -fa on, greedy, code prompt:
n-max 4 is the sweet spot on PQ2_0 (+31% on code); n-max 5 regresses. On PTQ1_0 the grafted head didn't accelerate decode in my tests. Prose sees little/no gain either quant (low draft acceptance), high on code. Also confirms the fix works for a grafted MTP head (donor embeddings), not just native. The model above is a public GGUF (I'm the author) if it's useful as a test artifact — it's an abliteration of Bonsai-2 with, incidentally, an honest-eval harness that measures why the usual refusal-removal scores come out inflated on reasoning models. Repro command happy to add if wanted. Extra: Posting the repro for the test-artifact record. Optimal --spec-draft-n-max looks hardware/workload dependent (lower values can win once concurrent load and GTT pressure are weighted), so treat these single-gen 3090 numbers as one regime, not universal. Build was prism b10687 (5d80cff) with 1dc4a57 cherry-picked, RTX 3090 24GB CUDA, single generation, greedy, code prompt: MTPllama-cli -m Ternary-Bonsai-2-27B-Uncensored-Gaston-PQ2_0-MTP.gguf baseline: same command on the non-MTP file, drop --spec-type and --spec-draft-n-maxNumbers from the Generation: X t/s line, single-gen on the 3090: PQ2_0 base 67.2, n-max 3 84.5, n-max 4 88.0, n-max 5 74.2 t/s. PTQ1_0 was basically flat (~56 either way), prose prompts saw little to no gain. |

Branch:
pr-hadamard-mtp, one commit on top ofprism(9a9394a89), filesrc/models/qwen35.cpp, 15 insertions.What
llama_model_qwen35::graph_mtplooks the draft token's embedding up with a rawggml_get_rows(model.tok_embd, ...). When the model folds a Hadamard rotation into its weightsand lists
token_embd.weightinprism.hadamard.inverse_weight_names(every Ternary Bonsai 2file does), the rows of that table are stored rotated. The main graph restores the primal basis
right after the lookup in
llm_graph_context::build_inp_embd(); the MTP graph did not. Thischange applies the same
llama_mul_mat_hadamardand sign vector after the lookup when thetable is in
hadamard_inverses. Models withoutprism.hadamard.*keys take the path they tookbefore (the lookup finds nothing).
Why
No PrismML export carries an MTP block today, so the case never came up. It does the moment
someone grafts the Qwen 3.8
blk.64.nextn.*tensors onto Bonsai 2 to get--spec-type draft-mtp(the graft tools and measurements are in github.com/sudoingX/bonsai2-small-gpu). Without the fix the
draft context is refused at load:
The workaround is to ship a second, unrotated embedding table inside the head
(
blk.64.nextn.embed_tokens.weight, a Q4_K copy of the donor'stoken_embd, 682 MiB of VRAM).With the fix the head reads the trunk's own table and the copy is not needed: 9,956 MiB instead
of 10,638 MiB at 131072 context on an RTX 3060 12GB, and 196608 context fits (11,990 MiB) where
the fat file OOMs on the MTP compute buffer.
Reproduction
Ternary-Bonsai-2-27B-PTQ1_0.ggufplus the 15blk.64.*tensors of any Qwen3.8-27B GGUF,with
qwen35.block_countset to 65 andqwen35.nextn_predict_layers = 1(
tools/extract_head.py --no-embed-tokensandtools/merge.pyin the repo above; themerged file is 6,297,658,848 bytes, sha256
1e33c571...5685).llama-server -m Ternary-Bonsai-2-27B-PTQ1_0-mtp-lean.gguf -ngl 99 -fa on -c 32768 -np 1 -ctk q4_0 -ctv q4_0 --jinja --spec-type draft-mtp --spec-draft-n-max 1spec common_specu: adding speculative implementation 'draft-mtp',speculative decoding context initialized.Measured (RTX 3060 12GB, 131072 context, q4_0 K/V, one slot)
The draft quality with the trunk's ternary, Hadamard-rotated embedding table equals the donor's
fp table: acceptance per run 0.85 to 0.94 on Python, 0.45 to 0.55 on prose, 0.73 to 0.79 on bash.
Known limits
qwen35MTP graph is changed. Other architectures with an MTP graph(
qwen3next,glm4-moe,deepseek2, ...) do the same raw lookup; none of them has aHadamard-folded export today, so they are left alone here.
llama_verify_hadamard_graphis the guard that catches the bug, and it now passes on thefile above.