LLM·

BitNest — Bit-nested draft and target weights — What does it mean?

The news. On October 2, 2026, researchers from the University of Georgia, Northeastern University, The University of Alabama and XPeng Motors posted BitNest: Bit-Nested Speculative Decoding for Memory-Efficient LLM Inference Acceleration on arXiv. They test it on LLaMA-2-7B, LLaMA-3-8B, Qwen2-7B and Qwen2.5-7B on one NVIDIA RTX A6000, plus NVIDIA Jetson Orin NX edge boards with 16 GB and 8 GB of memory, and report an average 95.2% acceptance rate and 1.48–1.61× end-to-end decoding speedup over 16-bit autoregressive decoding. Read the paper →

Picture a courier who plans routes quickly on a rough city map, while a dispatcher checks each plan against the detailed map before the van leaves. That is speculative decoding: the rough map is the draft model, the detailed map is the target. The usual setup keeps two map books, a small draft model and the full model, and on a small device the second book takes shelf space the trip log needs. The paper's own example: in 16-bit precision a 7B model alone can come close to the memory of a 16 GB device, leaving little room for the KV cache.

BitNest prints one book differently. The detailed map is a base sheet plus a clear overlay laid on top, so the rough map the courier uses is literally the first layer of the dispatcher's map, not a second copy. In weight terms, each 8-bit target weight is a 4-bit base code plus a 4-bit refinement code, and the draft model is the base codes alone. This is a draft-model variant that costs no extra parameters, the same goal as layer-skipping drafts, but the draft here is a lower-precision view of every layer instead of a shorter stack of layers.

Draft Model (1B) — generates K=5 tokens
Paris.Itis
↓ verify all at once↓
Target Model (70B) — one forward pass
✓ Par✓ is✓ .✗ It→The— is
accepted
rejected → corrected
discarded
3 accepted + 1 corrected = 4 tokens from 2 forward passes
↻ repeat until done

The order in which the map is drawn matters. The obvious approach is to make the best 8-bit map and then call its rough strokes the base. The authors point out that quantizing the 8-bit model on its own generally moves the 4-bit solution, so the nesting breaks. BitNest draws the base sheet first and makes it as good as possible, then draws the overlay only to correct it, never touching the base. Concretely, it first applies a function-preserving SpinQuant rotation that redistributes outliers in hidden states and activations, builds the 4-bit base with GPTQ, and then quantizes the leftover residual into a second 4-bit code whose scale is derived from the base scale, so the refinement needs no extra scale metadata. A strong base is what makes the draft agree with the target: across four models and six workloads the paper reports a 95.2% average acceptance rate.

Then the drawers. Memory is fetched in whole bytes, so if each weight's base and overlay bits shared one byte, the draft would haul in the overlay bits it never uses and the cheap draft would not be cheap. BitNest therefore stores the two halves as dual-plane storage: all base codes packed in one stream, all refinement codes in another, each half the size of the 8-bit weights. A draft step reads only the base plane, 4 bits per weight of weight-code traffic, and a verify step reads both. Because the decode step is limited by memory traffic rather than arithmetic, reading half the weight bytes is where most of the draft's saving comes from.

The same trick is applied to the KV cache at long contexts: drafting reads a KV4 base plane, verification reads the full KV8 cache, and the KV entries of rejected draft tokens are thrown away. At short contexts both steps read KV8, because there the KV traffic is small and draft agreement matters more.

Why not just run the 4-bit base as the final model and skip the overlay? The paper's ablations show that a draft good enough to agree with the target is not good enough to be the target. On LLaMA-3-8B, a W4A8 model scores 41.4% on GSM8K math problems, against 50.0% for W8A8; adding the overlay brings the BitNest target back to 49.8%. The KV side shows the same split: on Qwen2.5-7B, verifying with KV4 instead of KV8 lowers GSM8K from 81.6% to 78.4%, even though the draft still matches the target 93.3% of the time. A high acceptance rate only says the draft agrees with the target; it says nothing about whether the target itself is accurate.

A worked example from the paper, holding three things fixed: LLaMA-2-7B, an 8-bit target, and a draft that reads 4-bit weights. Stored the usual way, a separate W4 draft model next to the INT8 target takes 9.38 GiB of weights. Nested, the same pair takes 6.22 GiB, exactly the size of the standalone 8-bit model, because the draft is already inside it. The difference is 9.38 − 6.22 = 3.16 GiB, about a third of the two-model footprint, and on a 16 GB device that is memory the KV cache can use. On the 16 GB Jetson Orin NX, the paper reports BitNest running LLaMA-2-7B-32K at about 1.5× the FP16 decode speed, cutting decode energy from 1.52 to 0.95 J/token, and raising the longest context that fits from 2K tokens (FP16) to 10K.

Method (LLaMA-2-7B, one RTX A6000)DraftFinal modelSpeedup vs each method's 16-bit baseline in its own engine, six workloadsPeak memory at 1K contextSource
QSpecW4A4W4A16 (4-bit weights)1.35–1.46×4.60 GiBTables 1, 3
QuantSpecW4A16W16A161.12–1.23×13.21 GiBTables 1, 3
Draft & Verify (layer skipping)W16A16, fewer layersW16A160.96–1.20×14.21 GiBTables 1, 3
BitNestW4A8 (base plane)W8A8 (both planes)1.50–1.61×7.22 GiBTables 1, 3

The trade is that the final model is 8-bit, not 16-bit, and the draft can only be as parallel as one model running at lower precision. The quality cost is small in the paper's tests (LLaMA-2-7B WikiText-2 perplexity, where lower is better, 5.51 for BitNest against 5.47 for FP16), and QSpec still uses less memory at short contexts because its whole model is 4-bit. The authors also note that BitNest does not use auxiliary drafters such as DFlash, which propose whole blocks of tokens in parallel and may be faster when there is memory to spare for a second model. Bit nesting is a design for the case where memory, not compute, is the limit.

Goes deeper in: LLM Serving → Speculative Decoding → Draft Model Variants

Related explainers

Frequently Asked Questions

Check what you knowMap your AI & GPU knowledge across every track — free, role-based