Skip to content

speculative: add opt-in prompt lookup for DFlash2 - #304

Open
bri-prism wants to merge 1 commit into
prismfrom
perf/yukon-followup-20261002
Open

bri-prism wants to merge 1 commit into
prismfrom
perf/yukon-followup-20261002

Conversation

@bri-prism

Copy link
Copy Markdown
Collaborator

Overview

Add opt-in prompt lookup to DFlash2 with LLAMA_DFLASH2_LOOKUP=1. When the committed output matches a prompt span, use its continuation as the draft and skip the neural noise-block forward. The target still verifies every proposed token. Misses, ambiguous continuations and changed prompt prefixes fall back to neural drafting. Disabled by default; target/drafter weight sharing is unchanged.

The first round requires a unique two-token match. Later rounds use the longest matching suffix of 16-64 tokens, bounded by the original prompt and requested draft depth.

Validation

  • Release llama-server build and standalone C++ policy checks passed: unique/ambiguous matches, depth boundaries, prefix mismatch, suffix matches and identical/conflicting ties.
  • M5 Pro, one greedy request, 512 prompt tokens, 129 emitted IDs (128 decode tokens), warmup plus five measured repetitions per arm: median decode 75.72 -> 154.29 tokens/s (2.04x). Median prefill plus decode 3.056 -> 2.186 seconds.
  • All output IDs in every measured run and warmup matched the serial control. Both arms used the same build and retained the same idle/thermal gates.
  • This is a copy-heavy coding prompt, not evidence of general-workload speedup. Multi-request, long-context and mixed-hit stress validation remain pending, so the feature stays opt-in. Timing is paused; no additional performance claims are included.

Requirements

  • Contributing guidelines reviewed; this submission targets the maintained fork.
  • AI usage disclosure: YES. Codex implemented the change and assisted with validation and this description.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant