Inference · Speculative decoding · RTX 4090

Qwen3.8-27B on one RTX 4090

245 tokens per second on the official weights. 322 with better ones.

On the same card and benchmark, at batch size 1, vLLM decodes 200 tokens per second and llama.cpp 102. Part I is the engine, run on the official release weights. Part II re-quantizes the model so that it ends up closer to full precision than that release. Part III is about measuring differences of 0.3%.

Figure 1. Real decoding on one RTX 4090, replayed with its recorded timing from each lane's first token (this post is about decode speed; llama.cpp's first token takes about 0.55 s, the others' under 0.2 s). Each tape segment is one verify round, and its width is the number of tokens that round produced. The first lane decodes one token per forward pass; the others use speculative decoding, with the engines and weights of Figure 2. Each prompt is the median prompt of its Spec-Bench category by our tokens per round. Lanes can produce different text: different weights, or kernels of different shapes, break near-ties differently.
Figure 2. Decode speed by engine, all measured in one session on one RTX 4090 with the same streaming client (batch 1, greedy, thinking off). Spec-Bench is 480 prompts with 256 output tokens each. The single prompts add structured output and contexts up to 32K tokens; each is the median of 3 runs, and one near-tie can move a bar by several percent. Every engine runs a public 4-bit build of the official Qwen3.8-27B checkpoint with a DFlash2 drafter: llama.cpp a4d880f with unsloth's UD-Q4_K_XL and up to 7 drafts; vLLM 0.27.1 with sidnaZ's single-user RTX 4090 recipe (AutoRound W4A16, 7 drafts); Cinference-4090, the engine we started from, and our engine, both on the NInfer release container. The last bar adds our weights. Hover over a bar for tokens per round and round time.

The problem: one 24 GB card, one user

At batch size 1, every forward pass reads every weight once. Qwen3.8-27B at about 4 bits per weight is 13–14 GB, and the 4090 streams 0.93–1.0 TB/s from DRAM in practice. That caps plain autoregressive decoding near 70 tokens per second. We measured 46 tok/s in llama.cpp and 54 in Cinference (the first lane of Figure 1).

Speculative decoding changes the unit of work. A small drafter proposes a tree of 15 candidate tokens. The target model scores all 16 positions in one forward pass, which still reads the weights once. It then keeps the longest branch that matches its own greedy choices, plus one token of its own. Every emitted token is the target's own greedy choice, so the drafter can change speed but never the output. If a round accepts τ tokens on average:

tok/s = τ / tround,    tround ≈ bytesround / bandwidth + tsmall kernels + tidle

So there are three things to work on: the tokens a round accepts, the bytes it reads, and the time DRAM spends not streaming. Figure 3 is the map of this post: up is more tokens, left is shorter rounds.

Figure 3. Tokens per round against time per round on Spec-Bench; each dashed ray is a line of constant tok/s. From Cinference-4090, the verify tree moves straight up, and kernel work and fewer bytes move left. The diamonds are the other engines: vLLM's 7-draft chain accepts 4.07 tokens per round in 20.3 ms, llama.cpp 4.09 in 40.0 ms, and our final configuration 5.43 in 16.8 ms. By task type splits our configurations into the six Spec-Bench task groups: round time hardly depends on the task, while tokens per round ranges from 3 to 9. Hover over points and arrows for details.

The model is a hybrid: 48 Gated DeltaNet (linear-attention) layers and 16 full-attention layers, hidden size 5,120. The drafter is DFlash2, which reads the target's hidden states and proposes a whole block of tokens at once. The engine is Cinference, a descendant of NInfer, in jram4's port to Ada (sm_89). We merged upstream's verify trees into that port and added 53 commits. The card is one RTX 4090 at its 450 W limit, on PCIe Gen3 with CUDA 12.8.

Part I

The engine, on the official weights

Everything in this part runs the released container unchanged: target and drafter weights byte for byte as published. It covers the first two arrows in Figure 3.

More tokens per round: verify trees

The cheapest win was already in upstream's code. Instead of one chain of drafts, the drafter's per-position candidates are grown best-first into a 16-node tree, and the target verifies the whole tree in one pass with a tree attention mask. A 16-node tree accepts 20% more tokens per round than a 16-token chain at the same cost. Against the 7-token chain of Cinference-4090, our defaults accept 29% more tokens per round at the same round time, even though each round now verifies 16 positions instead of 8.

Figure 4. Real verify trees, from a dump of the final configuration on the same four prompts. Each bar in the strip is one round, and its height is the tokens that round produced. Click a bar or press play. Green nodes matched the target's greedy choice and were accepted. The dashed orange node is the target's own token, which every round emits. Gray branches were drafted but rejected. Code and math accept long branches. Prose mostly accepts short ones, and there the side branches do the work. Below the tree: the text so far, then the accepted tokens and the bonus token.

The drafter was trained to propose blocks of 8, and we ask it for 15. A tree temperature of 1.25 on the drafter's scores spreads the 15 nodes a little wider than its raw probabilities would. Wider trees do not pay on this card. Replaying recorded trees offline, a 32-node tree accepts 11% more tokens, but 32-column GEMMs cost 12–39% more per shape, because the activations no longer fit the 16-column tiles.

Where a round goes

Figure 5. One verify round, kernel by kernel: the median round of an nsys trace on a code prompt, for three builds. Oct 2 is our first tree build on the official weights. Oct 3 adds Ada-tuned GEMM schedules. Oct 6 is the final configuration, with our weights. Drag across a strip to zoom, or use the presets. 4 target layers shows one repeating block of three GDN layers and one attention layer, where the small kernels and the DRAM-idle slivers between GEMMs become visible. Hover over the legend to highlight a kernel class. nsys adds a few percent to every kernel.

A round has a fixed shape (Figure 5). It commits the previous round's recurrent state, runs the drafter (five layers) and its proposal head, grows the tree, verifies it through the target's 64 layers, scores the result with the LM head, and accepts. In the Oct 3 trace, 20.4 of 22.9 ms went to weight streaming: target and drafter GEMMs and the two LM heads. The rest went to hundreds of small kernels: norms, the recurrent GDN update, attention. Those kernels are latency-bound and leave DRAM idle while they run. Hopper can overlap one kernel's tail with the next kernel's start (programmatic dependent launch); Ada cannot, so every kernel boundary is a small bubble.

Streaming kernels

Small-T GEMM schedules tuned for Ada, single launches that cover both input projections of a layer, and evict-first hints on weight loads bring the weight GEMMs to 930–1,000 GB/s. The tuning alone took the median round from 25.3 to 22.9 ms (Oct 2 → Oct 3 in Figure 5). The largest GEMM, the Q4 gate/up projection, streams 100 MB in 98 µs. The 248k × 5,120 LM head is read through a 3-bit copy that keeps 32 candidates per position, and only those are rescored exactly. That screen runs at 927 GB/s, the card's measured streaming rate, and its outputs were identical on all 480 prompts. The drafter's proposal head uses the same trick.

L2 fills

A latency-bound kernel is DRAM-idle time, and the obvious ways to use it fail on Ada. prefetch.global.L2 does not populate L2, and prefetching from a side stream slows whatever kernel it lands next to. What works is giving each latency-bound kernel a few extra CTAs. After the kernel's own loads are issued, they pull the first megabytes of the next GEMM's weights into L2 with ordinary cached loads (Figure 6).

Figure 6. L2 fills, schematic and not to scale. A latency-bound kernel such as an RMSNorm leaves DRAM idle for a few microseconds. Extra CTAs in the same launch read the first megabytes of the next GEMM's weights into L2 (the 4090 has 72 MB), so the GEMM starts warm. Results are bit-identical, because the fill only loads data.

Fills on five kernel types cut round time by about 0.5%. The accounting was subtle in two ways. First, size matters: doubling the fills made rounds slower, because the stretches they ride on were already full. Second, deleting a small kernel also deletes its fill window. Our faster attention kernel at first showed no gain, because the old kernel had been carrying the fill for the next projection. Folding the RMSNorms into neighbouring GEMMs removed 80 launches per round and saved no time at all, for the same reason.

Small kernels, and one CMake property

The GDN convolution moved into the input projection's epilogue (−0.4%, bit-exact). Verify attention was rewritten to split rows across warps and CTAs: 17 µs became 7 µs at the median Spec-Bench context, and it reproduces the old kernel's split partials bit for bit. A chunked (WY-form) recurrent update takes 6 µs instead of 12–20. That one is not bit-exact: it misrounds 0.02% of outputs by one BF16 ulp, against 0.01% for the staged kernel it replaced.

The engine compiled every CUDA file as relocatable device code (-rdc=true), although nothing calls across translation units. Compiled whole-program, 541 of 2,801 kernels use fewer registers. The gate/up GEMM drops from 84 registers to 72, which lets three 256-thread CTAs share an SM instead of two. That is −0.5% round time with identical outputs. It also explained an earlier negative result: an epilogue fusion that had lost 1.5% to the same register cliff.

A cheaper cache and state

KV-cache values move from NVFP4 to INT8. Keys stay FP8, after a Hadamard rotation. Against a BF16 cache, INT8 values measure as lossless, where NVFP4 had cost about 0.0015 KL and 0.002 top-1, and verify attention gets faster too. The recurrent GDN state is stored in FP16 between rounds. In the final round (Figure 5, Oct 6) the GPU is idle for 0.29 ms of 17.0, mostly while the host prepares the next round.

Part II

Better weights

Part I left a round that is almost all weight streaming. From there, the only big lever left is fewer bytes. But a lower-bit model that is worse than the official one would not count. So we first pinned down what "no worse" means.

The reference is FP32 compute over the BF16 checkpoint, streamed layer by layer on the same card. We score 46,059 next-token predictions in 44 windows of 2,048 tokens, half Wikipedia prose and half CPython source, none of it overlapping any calibration data. We track three numbers: KL64 (KL from the reference to the candidate on the reference's top-64 tokens, which carry 98% of its mass), top-1 agreement and perplexity. A change ships only if every number, on prose and on code, is at least as good as the release's.

GPTQ, then spend the headroom

The release quantizes the target with round-to-nearest at 4 and 5 bits, and that leaves a lot on the table. GPTQ with act-order and a per-group clip search, at the release's own bit allocation, cuts KL64 by more than a third. Everything after that spends this headroom on bytes:

  1. All projections at Q4, with 8-bit group scales. This saves 1.8 GB per round. Larger calibration sets (96, then 384 windows) won back about a third of the KL that all-Q4 cost.
  2. 3 bits where they are cheap. Sensitivity is very uneven (Figure 7). Late MLP down projections cost 6× more KL per byte saved than early ones, and code is 4× more sensitive than prose. MLP gate/up in layers 0–31 and down in layers 0–15 now run at Q3. They stay in the existing Q4 container as codes restricted to [−4, 3] and are repacked into 3-bit tiles at load, so no new file format was needed. The Q4 copy stays resident for prefill (see Limitations).
  3. End-to-end scale tuning. With the integer codes frozen, we train only the group scales, a per-row offset and the RMSNorm gains (4.6 M parameters) to match the BF16 model's output distribution on held-out text, in the spirit of EfficientQAT's E2E-QP. The engine's exact KV-cache codec runs inside the training loop. Held-out KL drops 10.5%, most of it on code.
Figure 7. Where 3 bits are cheap. Each cell moves one projection group in one range of layers from 4 to 3 bits, in a simulation of the all-Q4 model. It shows the KL cost and the bytes saved per round. Darker cells buy more bytes per unit of KL. ★ marks our choice. Click cells to build your own allocation. The meter adds the costs, which we measured to be close to additive, and compares them with the KL budget the all-Q4 model leaves under the release. Bytes are not time: the GDN projections' 3-bit routes lost bandwidth, so only the MLP cells turn bytes into speed.

Step by step, the final model reads 19% fewer target bytes per round than the release, at lower KL and perplexity:

target GB per roundKL64 proseKL64 codeKL64top-1PPL
release (RTN Q4/Q5)14.370.02900.10950.06680.92503.294
GPTQ, at the release's bit allocation14.37––0.04210.9425–
all projections Q4, 8-bit group scales12.57––0.06040.9328–
+ Q3 MLP gate/up, layers 0–1512.20––0.0629––
+ end-to-end scale tuning, Q3 gate/up 0–3111.82––0.0652––
final: + tuning v5, Q3 MLP down 0–15, INT8 V cache, FP16 state11.640.02650.10180.06190.92503.213

Engine fidelity against the FP32 reference over 46,059 predictions (reference perplexity 3.146). Lower KL and PPL are better. Bytes are the target weights read per verify round, from the bit allocation. The two Q3 steps were scored on KL only; the final model has 42,606 top-1 agreements against the release's 42,605.

Quality is also a speed variable. A noisier target agrees less often with a drafter distilled from the BF16 model. An RTN all-Q4 target accepted 5–7% fewer tokens per round than the release, which cancelled most of its byte savings. So every quantization step was judged on round time and tokens per round together, never on bytes alone.

A drafter for the new target

A LoRA fine-tune of the drafter on 8,000 sequences generated by the quantized target gave +2.3% tok/s. Higher rank, more epochs, soft labels and new data all landed at the same place. Measured teacher-forced on a chain, that fine-tune raised acceptance by 6.8%. Measured in the engine, tokens per round rose only 1.4%. The tree recovers most of what a weak chain loses, so chain metrics overstate drafter changes about threefold. Every drafter decision after that was measured in the engine.

The drafter scores only a 131,072-token shortlist of the 248,320-token vocabulary, and a target token outside the shortlist ends the draft. Re-ranking the shortlist with the target's own outputs cut those misses from 0.96% to 0.35% of tokens, worth 2% more tokens per round. The drafter's own precision is a free parameter, since it can only change acceptance. Q4 and Q3 drafter projections paid for themselves once they ran on kernels that stream at full bandwidth.

Checking the path that actually runs

The fidelity numbers above score text through the prefill path. A speculative decoder spends its life elsewhere, in 16-token verify passes that read and write recurrent state between rounds, and some changes, such as FP16 state storage, exist only there. So we also measured the decode path directly. We run forced decoding on Spec-Bench prompts and dump the final hidden state at every accepted position. We then apply the BF16 head offline, exactly as the reference does, and score against an FP32 reference over the same text. Over about 40,000 positions, FP16 state moved KL64 from 0.02125 to 0.02079 and perplexity from 1.1942 to 1.1930.

The same tool exposed a bias in quality bars. Any perturbation that is not bit-exact loses 10–25 top-1 agreements out of about 46,000, whatever its size: rounding the state to 5 mantissa bits cost fewer than rounding it to FP16. Near a tie, the engine agrees with the reference more often than chance, because the two computations are close, and added noise of any size re-randomizes those ties. A rule like "top-1 must not drop" therefore rejects every harmless change. We used non-inferiority against the release instead, and treated the top-1 headroom as a budget that each non-bit-exact change spends.

Part III

Measuring 0.3%

Late in the project most wins were 0.3–1% of round time, and speculative decoding makes effects that small hard to see. A change that is not bit-exact flips a few near-tie tokens. From the first flip on, the generated text differs, and different text has different acceptance. A perturbation we knew to be null read as +0.44% tokens per round. Four practices made small effects measurable:

  1. Forced text. The engine can replace its greedy choice with a reference token at each position. Every configuration then walks the same 480 texts, and each round costs exactly what it would anyway. Validate the instrument first: forcing a configuration onto its own reference must reproduce its outputs. Our first version reproduced 0 of 482 because of an off-by-one, which would have passed for a plausible 3% drafter regression.
  2. Alternate configurations on one GPU, in a drift-cancelling order. Repeats of the same configuration drift by 0.05–0.6% within a job, as large as the effects. The SM clock steps from 2,715 to 2,685 MHz as the card warms, and a run that starts on an idle GPU reads fast. We kept builds off the box (a niced rebuild once cost 0.9%), held the other GPU's load constant and discarded a warm-up run. Then we ordered runs A B B A B A A B, which is orthogonal to linear and quadratic drift, and fitted the drift away (Figure 8).
  3. Paired per-prompt bootstrap. Comparing the same prompt across runs cancels prompt difficulty, giving intervals 5–13× tighter than unpaired ones. Next to it we report the run-level interval from the drift fit, which is the honest one.
  4. Explain every result. Each end-to-end number was checked against nsys per-call kernel times. Host-side costs came from unprofiled logs, because tracing inflates them several-fold.
Figure 8. The last adopted change, measured: an 8 MB versus a 4 MB L2 fill before verify attention, in eight forced-text Spec-Bench runs. Over the eight runs the card drifts by about 0.2%, more than the effect. Both analyses estimate B to be 0.16% faster, but the interval that ignores the drift is about twice as wide as the one from the drift fit. Drift removed shows the fit: the two configurations separate into two tight lines.

What did not work

Besides the dead ends described above:

idearesultwhy
Trellis quantization (EXL3-style)not adoptedMatching our scalar quality needs ≈3.75 bits per weight, at most 9% fewer bytes, and the decoder is slower.
Entropy-coded weightsclosedDecoding needs ≈2 T codes/s; the GPU decodes 1.5–1.65 T/s.
GDN or attention work on a second stream±0The kernels share DRAM and stretch each other one for one.
Four CTAs per SM for gate/up4 µs slower per callMore CTAs mean more activation reloads and more L2 contention.
2-bit LM-head screenrejectedTop-1 recall is 0.987.
FP32 residual streamtop-1 unchangedThe top-1 loss comes from BF16 projection inputs and outputs, not the residual.
SM-balanced grids, NUMA pinning, split CUDA graphs±0

Limitations

Code and data

Code: github.com/zedong-peng/qwen38-4090-decode. Weights: huggingface.co/zedongpeng/Qwen3.8-27B-NInfer-4090. The engine changes are a series of 53 small patches on the Cinference Ada port, each research feature behind an environment flag that is off by default. The tools for forced-text A/Bs, prefill and decode-path fidelity, nsys round accounting, the microbenchmarks and the data behind every figure on this page are in the repository. REPRODUCE.md goes step by step from a fresh clone to each Spec-Bench number: our engine on the official weights and on ours, and the vLLM and llama.cpp recipes. The configuration behind our final numbers:

NINFER_TWO_LEVEL_HEAD=1 NINFER_HEAD_SCREEN=3 NINFER_SCREEN_TILED=1 \
NINFER_PROPOSAL_SCREEN=3 NINFER_TREE_TEMPERATURE=1.25 \
NINFER_LS8=1 NINFER_Q3_SHADOW=1 NINFER_Q3_TILED=1 NINFER_DRAFT_CONV_Q4=1 \
NINFER_KV_VI8=1 NINFER_GDN_STATE_F16=1 NINFER_GDN_CHUNKED_RECORD=1 \
NINFER_CONV_EPILOGUE=1 NINFER_ATTN_ROWSPLIT=3 \
NINFER_L2_FILL=121 NINFER_L2_FILL_NORM_KB=1536 NINFER_L2_FILL_POST_KB=768 \
NINFER_L2_FILL_GATED_KB=768 NINFER_L2_FILL_RECORD_KB=2048 \
NINFER_L2_FILL_ATTN_KB=4096 \
ninfer-serve MODEL.ninfer --spec dflash2 --draft-tokens 15 --lm-head-draft \
  --verify-tree --kv-dtype k8v4 --max-context 32768 --kv-capacity 32768 \
  --max-concurrency 1 --prefill-chunk 1024 --temperature 0 --no-thinking \
  --no-prefix-reuse

Acknowledgments

This builds directly on NInfer (Neroued), Cinference (satellitedown) and jram4's Ada port; on DFlash2 (z-lab); on Spec-Bench (Xia et al.); and on the community llama.cpp and vLLM recipes for this model, which were our starting point.

A note on process: an AI coding agent (Claude Code) ran this project autonomously on the box for four days. It designed and ran the experiments, wrote the kernels and kept the research log, while the author set the goals, constraints and quality bar. Every number here comes from logged runs.

Citation

@misc{peng2026qwen38rtx4090,
  author = {Zedong Peng},
  title  = {Qwen3.8-27B on One RTX 4090},
  year   = {2026},
  month  = {oct},
  note   = {Blog post}
}