Tensormesh and AMD Collaborate to Empower Fewer GPUs to 
Serve More Models

SAN FRANCISCO – July 23, 2026 – Tensormesh, the company pioneering caching-accelerated inference optimization for enterprise AI, today announced a collaboration with AMD through which Tensormesh KV cache solution and AMD virtual memory offering will work together to allow more models to be served on fewer GPUs while retaining high KV cache hit rates and throughput even with oversubscribed high-bandwidth memory (HBM). Tensormesh is working with AMD, leveraging its GPU technology and AMD Live Context Virtualization components, and is tested using Dell servers with 8x AMD/ATI accelerators (MI355) GPUs and Dell storage. Tensormesh integrated LMCache coordinates KV cache management.

For users, running more models on the same set of GPUs means lower costs. LMCache users can now reuse the infrastructure they have built to virtually expand their GPUs' capacity. In addition, customers can reuse KV cache chunks stored for short-term memory virtualization later for prefix or non-prefix KV cache matching, maximizing system efficiency.

The Results

This new approach is far more efficient than building GPUs with more memory, which has led to the industry’s current memory supply crisis and increased the number of GPUs. Enterprises running AI over large document sets can now have a better experience and a lower bill. Testing showed:

  • Near-instant responses at scale: On a 300 GB document set (Kimi-K2.6), time-to-first-token dropped from 3.4 seconds to under half a second, a nearly 7x improvement, thanks to reusing cached KV data from DRAM and NFS instead of recomputing it from scratch.
  • No slowdown as usage grows: Output throughput held steady at ~48 tokens per second regardless of workload size, while unoptimized inference throughput fell by nearly 40% under the same load.
  • Twice the model density: The same hardware doubled the model density.

“This powerful new collaboration builds on AMD’s recent strategic investment in Tensormesh and expands the capabilities of AMD GPUs,” explained Junchen Jiang, Tensormesh CEO and LMCache co-creator. “Together, we’re greatly enhancing the memory that inference engines can access for model weights and KV cache, using all of the memory resources on each node.”

“We recognize Tensormesh and LMCache as KV cache management leaders,” said Anush Elangovan, AMD’s vice president of AI software. “And we’re thrilled to announce AMD’s breakthrough in virtual GPU memory management, amplified by LMCache and Tensormesh.”

The companies are sharing performance results from their collaboration at this week’s AMD Advancing AI 2026, a month before the public release of the LMCache open-source code.

Additional Resources:

  • Fixing AI’s Most Expensive Problem with Junchen Jiang, Tensormesh CEO
  • LMCache: Build The Foundation of AI Memory Tensor with KV Cache Infrastructure
  • Persistent KV Cache: Own Your Context Caching Lifecycle

About Tensormesh

Tensormesh is the leader in caching-accelerated inference optimization for enterprise AI. Founded by faculty, PhD researchers and alumni from the University of Chicago, UC Berkeley,  and Carnegie Mellon, and led by Junchen Jiang, University of Chicago faculty member and co-creator of LMCache, Tensormesh builds on years of academic research in distributed systems and AI infrastructure. The company has raised $24.5 million in total funding and is backed by Valley Capital Partners, NVentures, AMD Ventures, CoreWeave, and Laude Ventures.

Media Contact
press@tensormesh.ai

Recent Blog Posts

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua.Lorem ipsum dolor sit amet.

Name

Position
July 1, 2026

Designing AI Infrastructure Products for Developers

Read article

June 24, 2026

Persistent KV Cache: Own Your Context Caching Lifecycle

Read article

June 17, 2026

Fighting the Amnesia Tax: The Hidden Cost of Open-Weight LLM Serving

Read article

June 10, 2026

Run Open-Weight LLMs in Claude Code via Tensormesh Serverless Inference

Read article

June 2, 2026

Run Open-Weight LLMs in Your AI Agent with Codex CLI & Tensormesh Serverless Inference

Read article

May 28, 2026

Fixing AI's Most Expensive Problem — Junchen Jiang, Tensormesh CEO

Read article

May 27, 2026

Tensormesh Raises $20M from Investors Including AMD Ventures, CoreWeave, NVentures, Launches Tensormesh Inference to Fix AI’s Most Expensive Problem

Read article

May 20, 2026

KV Cache isn't just Cache, it's Memory: A Guide for LLM & Agent Devs

Read article

May 13, 2026

The AI Agent Metrics That Actually Matter: Beyond Tokens and Latency

Read article

May 6, 2026

Tensormesh Inference: Cheaper LLM Inference for AI Agents

Read article

April 29, 2026

Agentic AI Inference Cost: How LLM Agent Loops Break Caching and Drain Your Budget

Read article

April 28, 2026

Inside Tensormesh: Meet our CTO and Chief Scientist

Read article

April 22, 2026

Enterprise AI Vendor Lock-In: What It Costs When Your Provider Pulls Access

Read article

April 15, 2026

Introducing Tensormesh Beta 2.2: Serverless Inference & $0 Cached Input Tokens

Read article

April 8, 2026

How We Optimized Redis for LLM KV Cache: 0.3 GB/s to 10 GB/s

Read article

February 25, 2026

Introducing Tensormesh Beta 2: One-Click LLM Deployment, New UI & Real-Time Cost Savings

Read article

February 18, 2026

Agent Skills Caching with CacheBlend: Achieving 85% Cache Hit Rates for LLM Agents

Read article

February 11, 2026

Beyond Prefix Caching: How Non-Prefix Caching Achieves 25x Better Hit Rates for AI Agents

Read article

February 4, 2026

The Open Source Revolution: Why Open-Weight AI Models Are Redefining the Future

Read article

January 28, 2026

LMCache's Production-Ready P2P Architecture: Powers Tensormesh's 5-10x Cost Reduction

Read article

January 21, 2026

The Document Reprocessing Problem: How LLMs Waste 93% of Your GPU Budget

Read article

January 15, 2026

Building Tensormesh: A conversation with the CEO (Junchen Jiang)

Read article

January 7, 2026

The Hidden Metric That's Destroying Your AI Agent's Performance & Budget

Read article

December 17, 2025

LMCache Storage ROI Calculator: When KV Cache Storage Reduces AI Inference Costs

Read article

December 10, 2025

AI Inference Costs in 2025: The $255B Market's Energy Crisis and Path to Sustainable Scaling

Read article

December 3, 2025

New Hugging Face Integration: Access 300,000+ AI Models with Real-Time Performance Monitoring

Read article

November 26, 2025

The AI Inference Throughput Challenge: Scaling LLM Applications Efficiently

Read article

November 19, 2025

Solving AI Inference Latency: How Slow Response Times Cost You Millions in Revenue

Read article

November 13, 2025

GPU Cost Crisis: How Model Memory Caching Cuts AI Inference Costs Up to 10×

Read article

October 23, 2025

Tensormesh Emerges From Stealth to Slash AI Inference Costs and Latency by up to 10x

Read article

October 21, 2025

Comparing LLM Serving Stacks: Introduction to Tensormesh Benchmark

Read article