← Back to home
Flagship project · Oct 2026 · PyTorch · CUDA · C++

An LLM inference engine,
built from scratch.

Every layer of Qwen2.5-0.5B written by hand in PyTorch, then made fast: a KV cache, batched decoding, CUDA-graph replay, INT8 quantization, and a fused CUDA kernel. Each optimization was benchmarked on its own on a T4, including the two that didn't help. The result matches Hugging Face token-for-token and runs 2.1× faster at batch 32, 4.2× faster at batch 1.

1,929 tok/s
DECODE THROUGHPUT · BATCH 32 · T4
4.2×
FASTER THAN HUGGING FACE AT BATCH 1
5.8×
FUSED RMSNORM KERNEL VS PYTORCH
7.5e-5
MAX LOGIT ERROR VS HUGGING FACE
In plain terms

When a language model writes a reply, it produces one token at a time, and for every single token it has to read the entire model out of GPU memory — about 1 GB for this one. On a T4 that read alone takes ~3 ms; the actual math takes ~15 µs. The GPU spends over 95% of its time waiting on memory, not computing. So making inference fast isn't about doing math faster. It's about reading fewer bytes per token, or getting more tokens out of every read. This project is me building that from nothing and measuring what each trick is actually worth.

01 — The bottleneck

Decode is memory-bound.

Weights read per token
~1 GB
494M params × 2 bytes (fp16)
÷
T4 memory bandwidth
320 GB/s
HBM2, the physical ceiling
=
Time moving bytes
~3 ms
per token, at batch 1
vs
Time doing math
~15 µs
~1 GFLOP at 65 TFLOPS

That 200× ratio sets the ceiling: ~300 tokens/sec at batch 1, from bandwidth alone. Every optimization below either reads fewer bytes (INT8, fused kernels) or serves more tokens per read (KV cache, batching). And one more cost shows up only when you measure: a 24-layer decode step is ~300 tiny kernel launches, and at small batch the GPU finishes each one faster than Python can issue the next. That's what CUDA graphs fix.

02 — What each optimization is worth

Decode throughput, configuration by configuration.

Tokens per second, aggregate across the batch, 128 generated tokens, median of 3 runs on a Google Colab T4. Log scale — hover a bar for the exact value.

vLLM, the production reference, reaches 3,449 tok/s at batch 32 — 1.8× our best. Before CUDA graphs that gap was 3.3×.
ConfigurationBatch 1Batch 8Batch 32What changed
Hugging Face generate, fp1628.6240919The baseline everyone starts from
Ours, no cache, fp3234.2114131Recomputes every token each step — quadratic
Ours + KV cache, fp1636.02821,076Linear. The single biggest win
Ours + CUDA-graph decode1208061,929Whole decode step replayed as one launch
Ours + INT8 weight-only29.5228879Smaller weights, but slower — see below
Ours + fused RMSNorm31.22659905.8× faster kernel, invisible end-to-end
vLLM, fp16——3,449Paged KV, fused attention, continuous batching
03 — What the numbers say

Three wins, two honest negatives.

WIN · 8.2× at batch 32

The KV cache is the whole game

Without it, every new token recomputes attention over the full sequence, so cost grows quadratically — 131 tok/s at batch 32 and falling. Caching each layer's keys and values in preallocated buffers turns it linear: 1,076. This is the optimization every real engine has, and the one everything else stacks on.

WIN · 3.3× at batch 1

CUDA graphs beat the launch overhead

Eager batch-1 decode was 36 tok/s against a ~300 tok/s bandwidth ceiling, so memory wasn't the bottleneck there — Python was. Capturing one decode step with torch.cuda.graph and replaying it as a single launch gives 3.3× at batch 1, 2.9× at batch 8, 1.8× at batch 32. The gain shrinks as batch grows because the weight read starts to dominate. Making it work meant removing every Python-side integer from the step: a fixed mask over the whole cache, index_copy_ at a device-side position, positions derived on-device.

WIN · 16× from batch 1 → 32

Batching is the bandwidth argument made visible

Each decode step reads the ~1 GB of weights once. At batch 32 that one read serves 32 sequences, so throughput goes 120 → 1,929 for nearly the same memory traffic. This is why serving systems care so much about keeping batches full.

NEGATIVE · 18% slower

INT8 shrank the weights and made it slower

Per-row INT8 cut weights from 988 MB to 631 MB with no perplexity loss on WikiText-2 (18.07 → 17.99, inside noise). But throughput dropped from 1,076 to 879. Dequantizing int8 → fp16 on the fly in PyTorch writes a temporary fp16 copy of every weight matrix and then reads it back for the GEMM — an extra full pass that costs more bandwidth than the int8 read saves. The win only materializes with a fused dequant-GEMM kernel that reads int8 and multiplies in one pass, which is what TensorRT-LLM does. INT8 is a kernel problem, not a storage problem.

NEGATIVE · inside noise

The fused kernel is 5.8× faster and invisible

My CUDA RMSNorm replaces PyTorch's three ops (square-mean, rsqrt, multiply) — three launches and three memory round-trips — with one read and one write, reaching 128 GB/s on a prefill-sized block. In isolation it's 5.8× faster. End-to-end, nothing moves: RMSNorm is under 2% of a decode step for an 896-wide model, and 5.8× on 2% is Amdahl's law. Profile before you optimize. The kernel worth writing next is attention or the dequant-GEMM.

04 — The CUDA kernel

One block per row, one warp-shuffle reduction.

RMSNorm is y = x / sqrt(mean(x²) + ε) · w. PyTorch runs it as three kernels: square-and-mean, rsqrt, multiply — each reading and writing the full activation tensor. My kernel does it in one pass: 256 threads per block, one block per row; each thread sums the squares of its slice, a __shfl_down_sync reduction collapses each warp, shared memory collapses the warps, and every thread writes its normalized outputs. One HBM read, one write. Compiled as a PyTorch C++ extension with torch.utils.cpp_extension.load, templated over float, half, and bfloat16, and verified against PyTorch to 7e-4 in fp16.

The microbenchmark reports achieved bandwidth so the number means something: 128 GB/s is 40% of the T4's peak, respectable for a reduction with scalar loads. Vectorized half2 loads would push it higher. At decode shapes the kernel is launch-bound, which is exactly why it doesn't move the end-to-end number.

ShapePyTorchFusedSpeedupBandwidth
decode b=1 · 1×1×896111 µs19 µs5.8×launch-bound
decode b=32 · 32×1×896125 µs19 µs6.6×6 GB/s
prefill · 32×128×896643 µs115 µs5.6×128 GB/s · 40% of peak
05 — What's in the engine

Every piece, written by hand.

Hugging Face is used for exactly three things: downloading the weights, running the tokenizer, and being the reference I test against. Everything else is in the repo. Click a row.

The model

Embedding → 24 × (RMSNorm → grouped-query attention with RoPE → residual → RMSNorm → SwiGLU MLP → residual) → final norm → tied unembedding. Hidden size 896, MLP 4,864, 14 query heads sharing 2 key/value heads, head dim 64, vocab 151,936.

The details that had to be exactly right to match Hugging Face to 7.5e-5: RoPE uses the "halves" layout (rotating pairs i, i+32, not adjacent pairs); q/k/v projections carry a bias but o_proj doesn't; RMSNorm accumulates in fp32 even when the model runs in fp16; the output head shares the embedding matrix. engine/model.py

Preallocated [layers, batch, kv_heads, max_len, head_dim] buffers. prefill() runs the whole prompt once and fills them; decode() runs one new token and appends its key and value in place. Grouped-query attention makes the cache 7× smaller than full multi-head would be. engine/model.py → KVCache

Prompts of different lengths are left-padded into one tensor. A boolean attention mask keeps pad tokens from ever being attended to, and per-row RoPE positions make each sequence count from its own first real token. The test asserts that batched output equals running each prompt alone — it caught a real bug on the first run. engine/model.py → prefill / decode

The optimizations

A CUDA graph records every kernel launch in a decode step once, then replays the sequence as a single launch with no CPU in the loop. The catch: every tensor must keep a fixed address and shape. So the decoder owns static token, position, and mask buffers; writes the cache with index_copy_ at a device-side position tensor; and attends over the full cache with a float -inf mask that hides the unfilled slots. Between replays it updates those buffers in place. Verified token-for-token against eager decode. engine/graph.py → GraphDecoder

The bug worth mentioning: the first benchmark run crashed because prefill didn't reset the cache position between runs, so the second prefill wrote at slot 8 instead of 0. One line to fix, found only because the benchmark reuses the decoder.

Each of the 168 linear layers stores int8 weights plus one fp16 scale per output row, dequantized on the fly in forward. Embedding and output head stay fp16, which is why 988 MB became 631 MB rather than exactly half. Perplexity on WikiText-2 (40 windows of 512 tokens) was unchanged. Throughput dropped 18% for the reason explained above. engine/quant.py → QuantLinear

Described in full above. The engine's RMSNorm module routes to the kernel when a flag is on and the tensor is on CUDA, so the same code runs on CPU for tests. csrc/rmsnorm.cu, engine/kernels.py

The proof

Logits vs Hugging Face across four prompts (7.5e-5 max error). Cached greedy == uncached greedy == Hugging Face greedy, token for token. Batched left-padded output == each prompt alone. CUDA-graph decode == eager decode. INT8 keeps 100% of greedy tokens on the test prompt. Fused kernel vs PyTorch RMSNorm to 7e-4. Plus a --tiny mode that builds a random 2-layer Qwen so the whole suite runs on a laptop CPU in seconds. tests/test_engine.py

The mismatch that wasn't a bug: Hugging Face's greedy output differed from mine even with logits matching to 1e-4. Qwen's generation_config.json silently sets repetition_penalty=1.1 even for greedy decoding. Disable it and the outputs are identical.

Every configuration × batch {1, 8, 32}, 128 new tokens, same prompts, median of 3 after a warm-up, written to CSV. A kernel microbenchmark that reports µs, speedup, and achieved GB/s against the T4's peak. Perplexity on WikiText-2 for fp16 vs INT8 on identical windows. The same model through vLLM for an honest reference. One Colab notebook runs all of it top to bottom. bench/, run_all.ipynb

06 — What I'd do next

In order of expected payoff.

1

Fused INT8 dequant-GEMM kernel

Turns the 18% INT8 loss into the ~35% bandwidth win the smaller weights promise. The single most valuable kernel for this model.

2

Fused attention

FlashAttention-style: QKᵀ, softmax, and ×V in one kernel without materializing the score matrix. The largest remaining piece of the vLLM gap.

3

Paged KV cache + continuous batching

Fixed-size cache blocks so different-length sequences don't waste padded slots, then admit new requests as old ones finish.

4

Run it on an RTX 5070 Ti

Blackwell has ~3× the T4's bandwidth. Batch-1 should scale almost linearly; batch-32 less, because launch overhead doesn't scale with bandwidth. A clean way to test the whole model of where time goes.

Work on inference, GPU software, or ML systems?

I'd like to hear what you'd have done differently — which kernel you'd write first, what I'm measuring wrong, what a strong systems engineer should build next. Blunt is better.