
DeepSeek's EP+MoE training reduces reliance on NVLink/TP, eroding NVDA's hardware moat even as CUDA stays strong.
Disclaimer: The content provided in this newsletter is for informational purposes only and does not constitute investment advice. We are not registered investment advisors, and nothing in this newsletter should be construed as a recommendation to buy or sell any securities. Always do your own research and consult with a licensed financial professional before making any investment decisions.
In our previous article Research|DeepSeek: This isn’t a 'Sputnik moment' for U.S. AI, we discussed how DeepSeek has trained a top-tier open-source large model using approximately 2000 chips at a training cost of just 5 million US dollars and the impact on AI industry. This achievement has led to doubts about Nvidia's dominant market position. Specifically, in terms of hardware, the demand for Nvidia's NVLink technology has been significantly impacted, if the model will not further scale the size parameter and demand for KV cache exponentially. However, on the software side, Nvidia's CUDA ecosystem, the most fundamental moat, remains extremely robust compared with other 3rd party AI chips. We will analyze these two aspects in detail. For model scaling, as we talked about before, frontier labs and hyperscalers will keep pouring CapEx into it, hoping the scaling laws hold. Self-play RL, multi-modality and computer use may bring us more excitement in 2025, but the verification of these methods require multiple 100,000 GB200/300 clusters.
This report is available to subscribers. Sign in or subscribe to read the full analysis.