Intel tests presented at the OCP APAC Summit 2026 that moving the key-value cache used in large language model (LLM) inference from GPU memory to system DRAM can raise serving throughput and support more concurrent requests when VRAM is the bottleneck,...
The article requires paid subscription.
Subscribe Now