Solidigm and NVIDIA Rebuild Storage's Oldest Rule Around AI's KV Cache

Solidigm and NVIDIA Rebuild Storage's Oldest Rule Around AI's KV Cache

Infrastructure
NVIDIA and Solidigm are building storage that's allowed to lose data, because recalculating it is cheaper than protecting it.
By Shashi Bellamkonda · August 1, 2026
$10,000
per terabyte of high bandwidth memory, per NVIDIA's Kevin Deierling
600kW
projected power draw per AI rack by 2027
40 years
how long storage held to one rule, now being rewritten
Key Takeaway
NVIDIA and Solidigm are co-designing solid-state drives that sit inside the GPU memory hierarchy instead of behind it, and the entire justification rests on a single premise: a dropped byte of AI context costs a recompute, not a data loss event.

Ten thousand dollars a terabyte. That is the price NVIDIA's Kevin Deierling puts on the fast memory chip that sits next to every processor running an AI model today. Deierling said it during a conversation with Solidigm's Greg Matson (Deierling, qtd. in Kennedy, LinkedIn, 2026), and the price tag explains why two companies that build almost nothing together are now designing hardware side by side.

The subject was the key-value cache, known as KV cache, the mechanism a large language model uses to avoid recalculating every prior word in a conversation each time it generates the next one. Every chatbot exchange, every long document an AI agent works through, holds that running context in memory so the system does not reread the whole thing from scratch on each reply. Longer conversations, longer documents, more concurrent users, all of it grows the cache. Grow it past what high bandwidth memory can hold, and the industry has settled on one answer: push the overflow onto flash, storage built from the same solid-state technology as a laptop drive, running at a fraction of high bandwidth memory's cost.

The Cache That's Allowed to Forget

Storage has operated by one rule for four decades. Never lose a byte. RAID arrays, checksums, replication, the entire discipline of enterprise storage engineering exists to guarantee that whatever gets written comes back exactly as written.

KV cache breaks that rule on purpose.

Drop a cached key-value pair and the model does not corrupt a customer record. It recomputes the tensor and moves on, at a cost measured in graphics processing unit cycles instead of a data loss incident. A dropped byte costs a few seconds of processing time, not a permanent record gone.

That single fact opens design space that has not existed in storage engineering before, and Solidigm and NVIDIA are building directly into it (Solidigm, 2026). It is not the only way to attack the same cost problem. shashi.co covered Google's TurboQuant squeezing the same cache down to a sixth of its size through compression alone, no new hardware required. Solidigm and NVIDIA are solving the problem from the opposite direction, moving the overflow to cheaper media instead of shrinking it.

Two Vendors, One Rack

NVIDIA calls the result an inference context memory storage platform, or ICMS. It treats a pod of solid-state drives as an extension of the memory hierarchy, positioned past high bandwidth memory and dynamic random access memory but closer to the graphics processing unit than a traditional storage array (Solidigm, 2026). Solidigm builds the drives. NVIDIA's BlueField data processing units front the pool and decide where each piece of cached context lives.

Getting there took more than a software interface. At GTC 2025, Solidigm demonstrated liquid-cooled solid-state drives built to survive inside the same cooling loop as the graphics processing units, solving two problems standard drives were never built for: hot-swapping a drive without breaking a sealed liquid loop, and cooling a device that exposes only one face to a cold plate (ServeTheHome, 2025).

The resulting drive, the D7-PS1010, ships in two form factors, a 15-millimeter air-cooled version for NVIDIA's hybrid-cooled NVL72 racks and a 9.5-millimeter liquid-cooled version for fully liquid-cooled deployments, so an operator standardizes on one drive family regardless of which cooling architecture a given rack uses (Solidigm Newsroom, 2025).

Storage, compute, and cooling get specified together now, by the same team, often against the same vendor's roadmap.

The Bill Nobody Separates Anymore

The liquid loop is not optional at this scale, and it is the reason cooling and storage stopped being separate purchasing decisions. Data center racks are on pace to draw roughly 600 kilowatts by 2027 (ServeTheHome, 2025). At that density, air alone cannot move enough heat.

Fans get removed. Storage, compute, and cooling get specified together now, by the same team, often against the same vendor's roadmap. shashi.co's earlier look at how Nebius bought its inference stack one layer at a time traced the same pattern from the compute side. The pattern is repeating on the storage side now.

That is a different procurement conversation than buying a shelf of drives used to be. It also means a storage refresh cycle and a graphics processing unit refresh cycle, historically negotiated on separate timelines with separate vendors, now share a thermal envelope and, increasingly, a joint roadmap.

What We Don't Know Yet

Some of what Deierling and Matson described has not been independently benchmarked outside NVIDIA and Solidigm's own materials. Neither company has published third-party throughput or latency numbers for context-memory deployments running production inference at scale. Samsung, Micron, and Kioxia are pursuing versions of the same recomputable-cache approach, so whether Solidigm's specific design becomes the default, or one of several competing architectures, is not yet decided.

Key Takeaway
The KV cache is turning storage vendors into co-designers of the GPU rack rather than suppliers to it. For enterprises buying inference capacity, the storage decision and the graphics processing unit decision are converging into one contract.
CIO/CTO Viability Question
If your storage vendor now sits inside the same design loop as your graphics processing unit vendor, is your next storage refresh actually a graphics processing unit platform decision, and does your contract reflect that dependency?

Common Questions

What is KV cache in AI models?
KV cache is short for key-value cache, the working memory a model uses to hold a conversation's context so it does not recalculate every prior word on each new reply. Longer conversations and more concurrent users grow the cache.

Why is high bandwidth memory so expensive for AI inference?
It sits next to the processor and is built for speed over capacity, which makes it costly and supply-constrained. Deierling put the cost at roughly $10,000 per terabyte, far above solid-state flash doing a similar job at lower speed.

What is NVIDIA's inference context memory storage platform?
ICMS treats a pod of solid-state drives as an added layer in the memory hierarchy, sitting past high bandwidth memory and system memory but closer to the processor than a traditional storage array. NVIDIA's BlueField data processing units manage where cached context lives within that pool.

Are liquid-cooled solid-state drives necessary for AI data centers?
At roughly 600 kilowatts per rack, air cooling cannot remove enough heat on its own. Solidigm builds its D7-PS1010 in both liquid-cooled and air-cooled versions so operators can standardize on one drive family regardless of a rack's cooling architecture.

Does storing KV cache on flash risk losing AI conversation data?
No. Losing a cached key-value pair does not delete a permanent record. The model recalculates that piece of context, at a cost of a few extra seconds rather than a real data loss event.

Kennedy, Patrick. "Thinking Needs Memory." LinkedIn, https://lnkd.in/ey7VegvE. Accessed 1 Aug. 2026.
ServeTheHome. "Storage for the AI Factory Era: A Discussion." ServeTheHome, 14 May 2026, https://www.servethehome.com/storage-for-the-ai-factory-era-solidigm-nvidia-an-interview/.
ServeTheHome. "This Is the Solidigm Liquid-Coolable NVMe SSD Design." ServeTheHome, 24 Mar. 2025, https://www.servethehome.com/this-is-the-solidigm-liquid-coolable-nvme-ssd-design-nvidia-gtc/.
Solidigm. "Solidigm Develops One of the World's First Liquid-Cooled Enterprise SSDs for AI Deployments." Solidigm Newsroom, 18 Mar. 2025, https://news.solidigm.com/en-WW/248022-solidigm-develops-one-of-the-world-s-first-liquid-cooled-enterprise-ssds-for-ai-deployments/.
Solidigm. "ICMS Unlocks Greater AI Scale When Solidigm SSDs Store Context in KV Cache." Solidigm, 27 Jan. 2026, https://www.solidigm.com/products/technology/icmsp-ai-inference-is-flash-storage-problem.html.
VentureBeat. "Liquid-Cooled AI Systems Expose the Limits of Traditional Storage Architecture." VentureBeat, 24 Mar. 2026, https://venturebeat.com/infrastructure/liquid-cooled-ai-systems-expose-the-limits-of-traditional-storage.
Disclaimer: This blog reflects my personal views only. Content does not represent the views of my employer, Info-Tech Research Group. AI tools may have been used for brevity, structure, or research support. Please independently verify any information before relying on it.