PoC Proposal: Applying BitState/HSDR to Bonsai 2 Ternary GEMV — Direct 8-Phase Reduction and Reusable Activation State Sets #305
nohara-makoto
started this conversation in
Ideas
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
While experimenting with Bonsai 2 ternary inference, I have been exploring a way to keep activations and products in small logarithmic states and reduce a large number of terms into eight phase aggregates before numerical materialization.
I have been calling this experimental idea BitState/HSDR.
Later, I came across Ye Qiao's CurveFP v2 and realized that many parts of my approach overlap substantially with existing work. In particular, Sections 3.2 and 5 describe not only product generation through integer-index addition and sign XOR, but also signed histogram/POPCOUNT reduction and direct shifting into wide per-phase accumulators.
Because of this overlap, I do not intend to claim novelty for these mathematical ideas themselves.
My professional background is mainly in IT and systems infrastructure. I am not a researcher specializing in numerical formats or CPU/GPU kernel design; I have been exploring this topic as a personal side project and hobby.
My current interest is therefore much more implementation-oriented:
I also do not assume that the complete computation must remain in LNS or in the state domain until the end.
A practical implementation may first reduce many lanes in the state domain, then materialize the reduced result into an INT32 fixed-point representation such as Q15.16, and continue the upper part of the reduction using conventional integer arithmetic.
For the first PoC, I would like to focus on a single batch-1 CPU ternary GEMV path. Performance improvement and real-model quality impact have not yet been measured.
The current materials are published on Hugging Face.
The English Bonsai 2 demo Excel workbook may be the easiest starting point. It shows the reduction from 128 elements into eight phase aggregates and the position at which the group scale is applied.
The Excel workbooks and test vectors are not intended to represent a finished research result. Their main purpose is to provide implementation reference material so that someone writing a scalar reference, SIMD kernel, or GPU kernel can compare intermediate states and expected values.
The baseline activation format currently under consideration is 1 sign bit plus a 7-bit magnitude code:
This uses the same binary-log grid spacing as CurveFP, but it does not assume the same storage format, bias, or scale contract.
1. Start with a direct 8-phase path
Let the ternary weight be
where
d_gis the existing FP16 scale shared by a 128-element group.If
d_gremains outside the inner group reduction, the inner operation only needs ternary zero selection and sign inversion. No LNS magnitude addition with the weight is required.For each non-zero activation code, the following integer accumulation can be used:
The result of the inner reduction is only eight integer aggregates.
POPCOUNT is not mandatory for this path. A phase mask can be used as a routing selector while preserving the E4 weighting as
±(1 << q)during masked or grouped reduction.For real GEMV dimensions such as K=4096 or K=5120, the important cost is therefore not the final eight-value reconstruction, but the cost of routing and reducing the large number of terms inside each scale group into those eight aggregates.
Because Bonsai 2 uses different scales for different groups, each group result must incorporate its own scale before reduction across groups. I do not assume one shared scale across the entire K dimension.
For a more general LNS product path, or for a path in which scale is included in the product coordinate, the required offset must be applied before routing, and an internal coordinate wider than the stored 7-bit code may be necessary.
I would like to evaluate this separately from the factorized ternary path.
The direct 8-phase result also does not necessarily need to return immediately to floating point.
For example, after state-domain reduction within a group, the result could be materialized into an INT32 fixed-point representation such as Q15.16, followed by group-level or higher-level reduction using ordinary integer SIMD.
In this interpretation, the purpose of the state domain is not to replace the entire numerical system.
Instead, it is to delay numerical materialization and move as much dense lane-level reduction as practical into a compact state-domain representation before returning to conventional arithmetic.
2. Compare against reusable activation state sets
Another candidate is to build lane masks for activation magnitude states once for a 128-element activation tile and reuse them across multiple output rows that share the same activation input.
Each row's ternary weights can be represented using a non-zero mask and a negative mask.
From this signed population,
produces the same eight aggregates as the direct path.
The key questions are whether the state-mask construction cost can be amortized across enough output rows, whether real activations contain a sufficiently small number of active states, and how weight decoding, cache footprint, and register pressure affect the result.
For histogram final reduction, I would also like to compare the conventional dyadic shift/add path against a fractional-LNS candidate.
For example,
can be generated using a small count LUT, followed by a balanced reduction tree using fractional-LNS κ± correction LUTs.
The final tree across eight phase values is only three stages, but this does not mean that phase-internal reduction, group reduction, or state construction also become three stages.
Cancellation error and internal precision therefore need to be evaluated independently.
The public
Bonsai2_LNS8_128Excel sheet also contains an example of moving the shared scale into the final fractional coordinate:which is added once to the reduced 128-term result.
The Excel
LOGandPOWERfunctions are only explanatory references. An actual kernel would use LUTs, preconverted parameters, or another low-cost representation.I do not infer runtime performance or real-model quality from algebraic equality checks on synthetic examples.
I also do not consider the pure fractional-LNS path to be a required final design.
If reducing in the state domain and then returning to INT32 fixed-point arithmetic is simpler and faster on real CPUs or GPUs, I would prefer that hybrid path.
3. Implementation and measurement scope
For correctness, I would first start with PQ2_0 or predecoded ternary weights and compare:
After correctness is established, I would integrate PTQ1_0 dense decoding and compare against the current kernel including its real memory footprint and decode cost.
I would like to measure the following components separately:
The final comparison should include decode latency and TG128, including conversion overhead and any model-quality difference.
The current kernels already use SIMD and multiple accumulators, so I do not want to judge performance by comparing against a naive serial reduction.
If the CPU path is promising, I would then like to try Vulkan Compute or CUDA implementations using direct subgroup reduction and ballot-based state-set construction.
A possible GPU structure is:
The goal is to reduce the computational and reduction overhead of decode until it approaches the time required to fetch the weights, and to identify regions where the kernel can become closer to memory-bandwidth-bound.
BitState/HSDR does not necessarily need to replace the entire GEMV computation.
If even part of the overall reduction can be moved into the state domain, and the remaining computation can then continue efficiently using ordinary INT32 arithmetic, that would already be a useful result.
If the approach is slower, I would still like to identify which component dominates:
That information would still be useful for deciding which directions are worth pursuing.
4. Why I am publishing this now
There is also a practical reason why I am publishing these materials at this stage.
Starting in October 2026, my regular work is expected to become considerably busier, so the amount of time I can spend on this PoC may temporarily become quite limited.
Rather than keeping the work private until I can finish a complete implementation, I would prefer to publish the specifications, Excel workbooks, test vectors, and experiment ideas that are already usable.
I may not be able to complete the implementation quickly, and I may sometimes be slow to respond to questions or suggestions.
Even so, feedback such as:
would be extremely valuable to me.
My goal is not to claim ownership of the BitState name or the underlying ideas.
Where the work overlaps with CurveFP or other existing research, I would prefer to treat that prior work as the foundation and focus on producing implementation notes, reproducible examples, and test material that may be useful when applying the ideas to Bonsai 2 and llama.cpp.
5. Language note
English is not my primary working language.
Most of my technical thinking, notes, and spreadsheet development are done in Japanese.
I provide English versions so that the material can be discussed more widely, but some wording may be awkward or may accidentally sound stronger than I intended.
If an English sentence appears to make an excessive novelty claim, or if the technical meaning is unclear, that is not intentional.
Please feel free to point out both wording problems and technical mistakes.
At this stage, my goal is simply:
to leave useful implementation information, numerical examples, test vectors, and experiment ideas for people interested in running existing low-bit and ternary models such as Bonsai 2 efficiently on ordinary CPUs and GPUs.
If even a small part of these materials becomes useful for a future llama.cpp or related implementation, I would consider that a worthwhile outcome.
Public reference materials
I suggest starting with the following materials:
Direct_8PhaseandBonsai2_LNS8_128. Reduction into eight aggregates and scale-first/scale-late handling of the FP16 group scale.推奨_8Phase直集計andBonsai2_LNS8_128. Japanese explanation of the same processing path.A shorter overview is available in the dataset README / Technical Brief v0.5.
Additional materials in the same public folder include:
The five-profile and LNS16 ideas are future exploration material.
For now, I would like to keep the discussion focused on the base74 / S1E4L3 Bonsai 2 PoC.
Questions for implementers
Assuming post-Hadamard activations are reused across multiple output rows, where would direct routing or reusable activation-state construction fit most naturally into the current CPU scheduling/layout?
Does it make sense to begin correctness testing with PQ2_0 or predecoded ternary weights and integrate PTQ1_0 decoding only after the reduction paths are validated?
For AVX2 8-way weighted reduction, or for GPU subgroup-local
A[8]plus workgroup merge, are there layouts or measurement methods that would be particularly useful to try first?For a hybrid path that returns from state-domain reduction to an INT32 fixed-point representation such as Q15.16, are there obvious CPU/GPU disadvantages or more appropriate fixed-point placements that I may be overlooking?
When activation state sets are reused across multiple output rows, what would be a useful llama.cpp-like way to measure the break-even point between state-mask construction cost, register pressure, cache footprint, and row reuse?
References and verification scope
mainare intended to follow future updates.All reactions