
As generative AI shifts toward complex, long-context agentic workflows, the primary scaling bottleneck has moved from model training to inference deployment. Traditional architectures, constrained by local GPU HBM, force a 'recomputation tax,' where GPUs waste valuable cycles recalculating conversation history every time context limits are hit. This creates latency, spikes Time-to-First-Token (TTFT), and limits concurrency.Join our conversation with Anat Heilper, Director of AI Architecture at VAST Data, Yihua Cheng, CTO & Co-Founder, Tensormesh and moderator Calvin Nieh, VAST Data Technical Alliance Marketing Manager, to move beyond these limitations. Together, we will dive into the engineering reality of deploying KV cache offloading, sharing technical insights from VAST’s collaboration with GPU and KV cache framework partners, as well as global cloud providers and AI-forward enterprises.
Learn how to transform KV cache from limited and expensive local RAM into a high-value, persistent data asset using VAST’s Disaggregated Shared-Everything (DASE) architecture.
In this technical session, you will learn how to: