Tensormesh, AMD Partner to Let Fewer GPUs Serve More AI Models
Tensormesh, a company focused on caching-accelerated inference optimization for enterprise AI, has announced a collaboration with AMD through which Tensormesh's KV cache solution and AMD's virtual memory offering will work together to allow more models to run on fewer GPUs, while maintaining high KV cache hit rates and throughput even under oversubscribed high-bandwidth memory (HBM). Tensormesh is working with AMD by leveraging its GPU technology alongside AMD Live Context Virtualization components, and the setup has been tested using Dell servers equipped with eight AMD/ATI MI355 accelerator GPUs and Dell storage. Tensormesh's integrated LMCache handles KV cache management across the system.
For users, running more models on the same set of GPUs translates directly into lower costs. LMCache users can reuse infrastructure they've already built to effectively expand their GPUs' capacity, and customers can also reuse KV cache chunks originally stored for short-term memory virtualization for later prefix or non-prefix KV cache matching, further improving system efficiency.
The Results: This approach offers a more efficient alternative to simply building GPUs with more memory, a trend that has contributed to the industry's current memory supply crunch and driven up GPU demand. Enterprises running AI over large document sets can now expect both better performa...
