
Cache-hit price cut to RMB 0.025/M and 95%+ hit rates shift KV to SSD; NAND demand set to soar.
In our previous article we discussed DeepSeek V4’s architectural customization on non-NVIDIA hardware and the first round of API price cuts at 75% off. This article focuses on V4’s second round of cuts: DeepSeek separately took the input cache-hit tier further down to 1/10 of list, stacked on top of the 75% off from the previous round, with the floor at ¥0.025 per million tokens. This widens the cache hit / cache miss spread from 1/12 to 1/120 (cache hit ¥0.025 vs. cache miss ¥3). DeepSeek V4’s real-world cache hit rate in agent settings has reached 95%+, and based on our research, DeepSeek’s current SSD configuration and utilization have stepped up materially versus before. Behind this is V4 compressing KV cache size to 10% of V3.2’s, plus DeepSeek’s accumulated engineering work on SSD-based KV cache, which together migrate KV cache from expensive, capacity-limited DRAM / HBM onto larger and cheaper SSD at scale. We believe DeepSeek V4’s cache-hit repricing implies upside for SSD, with NAND demand set to grow exponentially.
DeepSeek V4 has executed two consecutive rounds of price cuts since launch. Round 1 was a limited-time 75% off: input (cache hit, ¥1 → ¥0.25), input (cache miss, ¥12 → ¥3), and output (¥24 → ¥6) all stepped to 25% of list in lockstep. Round 2 separately took the input (cache hit) tier further down to 1/10 of list, then stacked the limited-time 75% off on top, bringing this tier to a floor of ¥0.025. Cache miss (¥3) and output (¥6) were left unchanged, signaling a structural shift in cache-hit cost.
This report is available to subscribers. Sign in or subscribe to read the full analysis.