
RL and inference scaling outpace pretraining, driving sustained GPU demand; NVDA moat strengthens.
The rapid evolution of artificial intelligence (AI) models has led to an unprecedented surge in computational demands, challenging the capacities of even the most advanced tech infrastructures. Industry leaders like Microsoft and OpenAI are experiencing overwhelming demand for AI services, often surpassing their current computational resources. Microsoft's CEO, Satya Nadella, noted that Azure's OpenAI service usage has more than doubled in the past six months, with AI business expected to reach an annual revenue run rate of $10 billion by the second quarter, marking it the fastest-growing business in Microsoft's history. Similarly, OpenAI's CEO, Sam Altman, acknowledged that a lack of compute capacity is delaying the release of the company's products. This escalating demand underscores the critical need for scalable, efficient computing solutions to support the next generation of AI advancements.
However, there are several recent debate on whether the scaling law has ended and we’ve reached the plateau of AGI, e.g., the information’s recent articles on o1 and reasoning. However, those articles lack in-depth discussion in the technology behind it. To fully understand the GPU demand, we need to understand how the reasoning AI models are evolving and how that will affect Nvidia’s moat.
This report is available to subscribers. Sign in or subscribe to read the full analysis.