Every layer of Qwen2.5-0.5B written by hand in PyTorch, then made fast: a KV cache, batched decoding, CUDA-graph replay, INT8 quantization, and a fused CUDA kernel. Each optimization was benchmarked on its own on a T4, including the two that didn't help. The result matches Hugging Face token-for-token and runs 2.1× faster at batch 32, 4.2× faster at batch 1.
When a language model writes a reply, it produces one token at a time, and for every single token it has to read the entire model out of GPU memory — about 1 GB for this one. On a T4 that read alone takes ~3 ms; the actual math takes ~15 µs. The GPU spends over 95% of its time waiting on memory, not computing. So making inference fast isn't about doing math faster. It's about reading fewer bytes per token, or getting more tokens out of every read. This project is me building that from nothing and measuring what each trick is actually worth.
That 200× ratio sets the ceiling: ~300 tokens/sec at batch 1, from bandwidth alone. Every optimization below either reads fewer bytes (INT8, fused kernels) or serves more tokens per read (KV cache, batching). And one more cost shows up only when you measure: a 24-layer decode step is ~300 tiny kernel launches, and at small batch the GPU finishes each one faster than Python can issue the next. That's what CUDA graphs fix.
Tokens per second, aggregate across the batch, 128 generated tokens, median of 3 runs on a Google Colab T4. Log scale — hover a bar for the exact value.
| Configuration | Batch 1 | Batch 8 | Batch 32 | What changed |
|---|---|---|---|---|
Hugging Face generate, fp16 | 28.6 | 240 | 919 | The baseline everyone starts from |
| Ours, no cache, fp32 | 34.2 | 114 | 131 | Recomputes every token each step — quadratic |
| Ours + KV cache, fp16 | 36.0 | 282 | 1,076 | Linear. The single biggest win |
| Ours + CUDA-graph decode | 120 | 806 | 1,929 | Whole decode step replayed as one launch |
| Ours + INT8 weight-only | 29.5 | 228 | 879 | Smaller weights, but slower — see below |
| Ours + fused RMSNorm | 31.2 | 265 | 990 | 5.8× faster kernel, invisible end-to-end |
| vLLM, fp16 | — | — | 3,449 | Paged KV, fused attention, continuous batching |
Without it, every new token recomputes attention over the full sequence, so cost grows quadratically — 131 tok/s at batch 32 and falling. Caching each layer's keys and values in preallocated buffers turns it linear: 1,076. This is the optimization every real engine has, and the one everything else stacks on.
Eager batch-1 decode was 36 tok/s against a ~300 tok/s bandwidth ceiling, so memory wasn't the bottleneck there — Python was. Capturing one decode step with torch.cuda.graph and replaying it as a single launch gives 3.3× at batch 1, 2.9× at batch 8, 1.8× at batch 32. The gain shrinks as batch grows because the weight read starts to dominate. Making it work meant removing every Python-side integer from the step: a fixed mask over the whole cache, index_copy_ at a device-side position, positions derived on-device.
Each decode step reads the ~1 GB of weights once. At batch 32 that one read serves 32 sequences, so throughput goes 120 → 1,929 for nearly the same memory traffic. This is why serving systems care so much about keeping batches full.
Per-row INT8 cut weights from 988 MB to 631 MB with no perplexity loss on WikiText-2 (18.07 → 17.99, inside noise). But throughput dropped from 1,076 to 879. Dequantizing int8 → fp16 on the fly in PyTorch writes a temporary fp16 copy of every weight matrix and then reads it back for the GEMM — an extra full pass that costs more bandwidth than the int8 read saves. The win only materializes with a fused dequant-GEMM kernel that reads int8 and multiplies in one pass, which is what TensorRT-LLM does. INT8 is a kernel problem, not a storage problem.
My CUDA RMSNorm replaces PyTorch's three ops (square-mean, rsqrt, multiply) — three launches and three memory round-trips — with one read and one write, reaching 128 GB/s on a prefill-sized block. In isolation it's 5.8× faster. End-to-end, nothing moves: RMSNorm is under 2% of a decode step for an 896-wide model, and 5.8× on 2% is Amdahl's law. Profile before you optimize. The kernel worth writing next is attention or the dequant-GEMM.
vLLM's remaining lead comes from three things this engine doesn't have: paged KV memory (no padding waste across different-length sequences), fused attention kernels (QKᵀ, softmax, and ×V in one pass), and continuous batching (admit new requests as old ones finish). Each is a known next step, not a mystery.
RMSNorm is y = x / sqrt(mean(x²) + ε) · w. PyTorch runs it as three kernels: square-and-mean, rsqrt, multiply — each reading and writing the full activation tensor. My kernel does it in one pass: 256 threads per block, one block per row; each thread sums the squares of its slice, a __shfl_down_sync reduction collapses each warp, shared memory collapses the warps, and every thread writes its normalized outputs. One HBM read, one write. Compiled as a PyTorch C++ extension with torch.utils.cpp_extension.load, templated over float, half, and bfloat16, and verified against PyTorch to 7e-4 in fp16.
The microbenchmark reports achieved bandwidth so the number means something: 128 GB/s is 40% of the T4's peak, respectable for a reduction with scalar loads. Vectorized half2 loads would push it higher. At decode shapes the kernel is launch-bound, which is exactly why it doesn't move the end-to-end number.
| Shape | PyTorch | Fused | Speedup | Bandwidth |
|---|---|---|---|---|
| decode b=1 · 1×1×896 | 111 µs | 19 µs | 5.8× | launch-bound |
| decode b=32 · 32×1×896 | 125 µs | 19 µs | 6.6× | 6 GB/s |
| prefill · 32×128×896 | 643 µs | 115 µs | 5.6× | 128 GB/s · 40% of peak |
Hugging Face is used for exactly three things: downloading the weights, running the tokenizer, and being the reference I test against. Everything else is in the repo. Click a row.
Embedding → 24 × (RMSNorm → grouped-query attention with RoPE → residual → RMSNorm → SwiGLU MLP → residual) → final norm → tied unembedding. Hidden size 896, MLP 4,864, 14 query heads sharing 2 key/value heads, head dim 64, vocab 151,936.
The details that had to be exactly right to match Hugging Face to 7.5e-5: RoPE uses the "halves" layout (rotating pairs i, i+32, not adjacent pairs); q/k/v projections carry a bias but o_proj doesn't; RMSNorm accumulates in fp32 even when the model runs in fp16; the output head shares the embedding matrix. engine/model.py
Preallocated [layers, batch, kv_heads, max_len, head_dim] buffers. prefill() runs the whole prompt once and fills them; decode() runs one new token and appends its key and value in place. Grouped-query attention makes the cache 7× smaller than full multi-head would be. engine/model.py → KVCache
Prompts of different lengths are left-padded into one tensor. A boolean attention mask keeps pad tokens from ever being attended to, and per-row RoPE positions make each sequence count from its own first real token. The test asserts that batched output equals running each prompt alone — it caught a real bug on the first run. engine/model.py → prefill / decode
A CUDA graph records every kernel launch in a decode step once, then replays the sequence as a single launch with no CPU in the loop. The catch: every tensor must keep a fixed address and shape. So the decoder owns static token, position, and mask buffers; writes the cache with index_copy_ at a device-side position tensor; and attends over the full cache with a float -inf mask that hides the unfilled slots. Between replays it updates those buffers in place. Verified token-for-token against eager decode. engine/graph.py → GraphDecoder
The bug worth mentioning: the first benchmark run crashed because prefill didn't reset the cache position between runs, so the second prefill wrote at slot 8 instead of 0. One line to fix, found only because the benchmark reuses the decoder.
Each of the 168 linear layers stores int8 weights plus one fp16 scale per output row, dequantized on the fly in forward. Embedding and output head stay fp16, which is why 988 MB became 631 MB rather than exactly half. Perplexity on WikiText-2 (40 windows of 512 tokens) was unchanged. Throughput dropped 18% for the reason explained above. engine/quant.py → QuantLinear
Described in full above. The engine's RMSNorm module routes to the kernel when a flag is on and the tensor is on CUDA, so the same code runs on CPU for tests. csrc/rmsnorm.cu, engine/kernels.py
Logits vs Hugging Face across four prompts (7.5e-5 max error). Cached greedy == uncached greedy == Hugging Face greedy, token for token. Batched left-padded output == each prompt alone. CUDA-graph decode == eager decode. INT8 keeps 100% of greedy tokens on the test prompt. Fused kernel vs PyTorch RMSNorm to 7e-4. Plus a --tiny mode that builds a random 2-layer Qwen so the whole suite runs on a laptop CPU in seconds. tests/test_engine.py
The mismatch that wasn't a bug: Hugging Face's greedy output differed from mine even with logits matching to 1e-4. Qwen's generation_config.json silently sets repetition_penalty=1.1 even for greedy decoding. Disable it and the outputs are identical.
Every configuration × batch {1, 8, 32}, 128 new tokens, same prompts, median of 3 after a warm-up, written to CSV. A kernel microbenchmark that reports µs, speedup, and achieved GB/s against the T4's peak. Perplexity on WikiText-2 for fp16 vs INT8 on identical windows. The same model through vLLM for an honest reference. One Colab notebook runs all of it top to bottom. bench/, run_all.ipynb
I'd like to hear what you'd have done differently — which kernel you'd write first, what I'm measuring wrong, what a strong systems engineer should build next. Blunt is better.