Local LLM

Reserving 128k Context Costs Nothing. Filling It Costs Two Thirds.

Two RTX 3090s, eleven models with identical digests, 528 requests and no errors. Reserving a 128k KV cache is nearly free; filling it costs 10 to 66 percent of throughput — with nothing offloaded to the CPU on ten of eleven models. Plus: an attached monitor costs you 1.8 GB of usable VRAM.