KV-Cache Paging, AWQ INT4 Quantization, and Tensor Parallelism Across Dual-GPU Workstations

Distributed Architecture Takeaway

Standard Transformer inference wastes up to 80% of GPU memory due to KV-cache fragmentation. By adapting virtual memory paging principles from operating systems, PagedAttention (vLLM) enables 4x to 8x higher batch concurrency on local workstations.

Empirical Architecture Comparison: HuggingFace Baseline vs. vLLM with PagedAttention and AWQ

Serving EngineTokens / Second (Single Stream)Throughput at 32 Concurrent StreamsGPU VRAM Footprint (70B Model)
Vanilla HuggingFace Transformers (FP16)18.4 tokens/secOOM (Out Of Memory Crash at 6 streams)140 GB (Requires 2x 80GB A100s)
HuggingFace + bitsandbytes INT48.2 tokens/sec (Slow dequant)110 total tokens/sec38 GB VRAM
vLLM + PagedAttention (FP16)42.1 tokens/sec620 total tokens/sec140 GB (Efficient non-fragmented VRAM)
vLLM + AWQ INT4 + Tensor Parallelism (TP=2)68.5 tokens/sec1,420 total tokens/sec39 GB total (Runs on 2x RTX 3090/4090!)

1. The KV-Cache Bottleneck in Autoregressive Generation

During the autoregressive generation phase, an LLM generates tokens sequentially. To avoid recomputing past key and value vectors at every step, past projections are cached in GPU memory (the Key-Value Cache). The memory required for a single request scales linearly with sequence length: $$\text{Memory}_{KV} = 2 \times 2 \times n_{\text{layers}} \times n_{\text{heads}} \times d_{\text{head}} \times \text{seq\_len} \times \text{precision}$$ For a 70B parameter model with a 4,096 context window, a single user session consumes over 2.4 GB of VRAM solely for its KV-cache. In standard memory allocators, contiguous memory pre-allocation causes severe internal and external memory fragmentation.

2. PagedAttention: Operating System Virtual Memory for GPUs

Developed by Woosuk Kwon et al. at UC Berkeley, PagedAttention manages the KV-cache using physical memory blocks analogous to OS virtual memory pages. Rather than allocating contiguous physical VRAM for maximum context lengths, PagedAttention allocates fixed-size memory blocks (e.g., 16 tokens per block). A software block table maps logical token positions to non-contiguous physical GPU VRAM slots, reducing memory waste from 80% down to under 4% and enabling dramatic increases in batch size.

3. Activation-Aware Weight Quantization (AWQ)

Quantizing models from FP16 (16-bit floating point) to INT4 (4-bit integer) slashes memory footprints by 75%. However, naive round-to-nearest quantization degrades model reasoning. Activation-Aware Weight Quantization (AWQ) recognizes that not all weights are equally critical: $$\text{Salience}(W) \propto \|X \cdot W\|$$ By protecting the 1% most salient weight channels (those corresponding to high-magnitude activation features) and quantizing the remaining 99% to INT4, AWQ preserves FP16 perplexity scores while delivering 2.5x faster matrix multiplication kernels on Tensor Cores.

4. Tensor Parallelism Across Multi-GPU Nodes

When running 70B models locally on multi-GPU developer workstations (e.g., two 24GB RTX 3090s), Tensor Parallelism ($TP=2$) splits individual weight matrices across GPUs via Megatron-LM tensor slicing. Column-parallel linear layers split attention projection heads ($W_q, W_k, W_v$) across GPU 0 and GPU 1, followed by an `All-Reduce` collective communication step over high-speed PCIe 4.0 / NVLink, achieving blistering real-time inference without cloud API dependencies.