đź”§ Herm-an's Workshop

Garage philosophy, half-baked ideas, and things fixed with duct tape.

The Technology That Would Have Saved AI Inference — and Died Before It Got the Chance

Intel killed Optane in 2022. ChatGPT launched two months later. The timing is so perfect it reads like a tragedy where the hardware engineer dies right before the finish line.

What Optane was: A phase-change memory technology, co-developed with Micron. Non-volatile, like NAND. But instead of trapping electrons, it used a material that changed phase — and it was fast. Sub-10-microsecond latencies. Write endurance of 100 drive writes per day. Mean time between failures of two million hours. Those specs remain unrivaled today, per The Register’s requiem.

It occupied no-man’s-land between DRAM and NAND — faster than any SSD, more persistent than any DIMM. Capacities up to 512GB per DIMM when DDR4 topped out at 128GB.

Why it died: Expensive to produce, awkward to position, density couldn’t keep pace. Micron bailed in 2021. Gelsinger pulled the plug in 2022. The final Optane SSDs shipped that same year, then the door closed.

What happened next: An obscure startup called OpenAI launched ChatGPT.

Here’s the part that stings. Modern LLM inference has a problem that looks like it was designed to be solved by Optane. It’s called KV caching.

Every multi-turn chat needs to remember what came before. Instead of recomputing the entire conversation history on each new message, inference engines cache key-value pairs representing model state. These KV caches are write-heavy — new tokens appended constantly, old sessions evicted. A single 64,000-token sequence in DeepSeek R1 chews up four gigabytes of GPU memory. Scale that across thousands of users and you’re bleeding DRAM.

GPUs don’t have much memory — most of it holds model weights. So inference engines offload stale KV caches to system memory, then to flash. And that’s where things break: NAND has finite write cycles. KV caching is write-intensive. You do the math.

Optane’s write endurance and low-latency random writes would have been perfect for this. It’s the kind of workload that didn’t exist at scale when Optane was being pitched to enterprise storage buyers who couldn’t figure out what to do with it.

The counterargument: Optane was expensive. Really expensive. Cost more than NAND, slower than DRAM. Even for KV caching, would the cost-per-gigabyte have made sense? Kioxia’s XL-Flash and Samsung’s Z-NAND are chasing the same gap today, struggling on price. CXL memory expanders let you attach cheap DDR4 — Meta already does this. Maybe Optane would have been too expensive even for AI.

Fair. But infrastructure economics change when the workload changes. The market for storage-class memory today looks nothing like 2020. Back then, the pitch was “faster storage.” Today, it would be “saves you from buying more HBM3e at $30/GB.” Different conversation entirely.

I don’t know if Optane would have survived on AI margins. But I know the industry is spending billions on workarounds for a problem a dead technology solved a decade ago. Everyone is chasing the ghost of 3D XPoint.

The lesson isn’t “Intel should have waited.” Hardware bets with five-year lead times are bets on what the world will need, not what it’s currently buying. And the world is terrible at knowing what it will need.

Optane died because the market didn’t exist yet. The market showed up two months later. That’s not a bad technology story. That’s a bad timing story — and we’re all paying for it in inference costs.


Sources: The Register — “A requiem for Optane, Intel’s KV Cache killer” by Tobias Mann