KV-Cache Paging, AWQ INT4 Quantization, and Tensor Parallelism Across Dual-GPU Workstations
Distributed Architecture Takeaway
Standard Transformer inference wastes up to 80% of GPU memory due to KV-cache fragmentation. By adapting virtual memory paging principles from operating systems, PagedAttention (vLLM) enables 4x to 8x higher batch concurrency on local workstations.
Empirical Architecture Comparison: HuggingFace Baseline vs. vLLM with PagedAttention and AWQ
| Serving Engine | Tokens / Second (Single Stream) | Throughput at 32 Concurrent Streams | GPU VRAM Footprint (70B Model) |
|---|---|---|---|
| Vanilla HuggingFace Transformers (FP16) | 18.4 tokens/sec | OOM (Out Of Memory Crash at 6 streams) | 140 GB (Requires 2x 80GB A100s) |
| HuggingFace + bitsandbytes INT4 | 8.2 tokens/sec (Slow dequant) | 110 total tokens/sec | 38 GB VRAM |
| vLLM + PagedAttention (FP16) | 42.1 tokens/sec | 620 total tokens/sec | 140 GB (Efficient non-fragmented VRAM) |
| vLLM + AWQ INT4 + Tensor Parallelism (TP=2) | 68.5 tokens/sec | 1,420 total tokens/sec | 39 GB total (Runs on 2x RTX 3090/4090!) |