TSLAGOOGLNVDA

Grok 4 Dominates Benchmarks Through the Power of Scaling Law, and the Potential ChatGPT Moment for Waymo

FUNDA·2025年7月11日

XAI scales Grok 4 with 200k H100s as Waymo begins NYC robotaxi tests, setting up a Tesla showdown.

Two major stories have been making waves in the AI community over the past couple of days:

First, Elon Musk's XAI has unveiled Grok 4 just half a year after its previous release. Grok 4 epitomizes the "brute force elegance" of scaling laws, trained on an unprecedented cluster of over 200,000 H100 GPUs. According to Tony from XAI, the computational resources dedicated to Grok 4 have exponentially surpassed those of Grok 3. Particularly impressive are XAI’s parallel reasoning and multi-agent reasoning capabilities, vividly demonstrating the power of test-time inference scaling. The parallel reasoning system clearly showcases the immense bandwidth requirements for cluster-based inference, encompassing task decomposition for agents, concurrent processing, scheduling, and memory pooling, etc.

Additionally, XAI has publicly stated that half of their compute power is now dedicated to large-scale, verifiable reinforcement learning (RL)—internally dubbed as "RL is the new pre-training"—setting the stage for future AGI exploration. Interestingly, RL’s compute demand is now approaching or even exceeding that of traditional pre-training, a shift we highlighted as early as the second half of last year. Recently, OpenAI’s technical lead Noam Brown also emphasized that mid-training, primarily driven by RL, is becoming the new scaling frontier. With the growing accumulation of high-quality data through the RL-driven flywheel in specialized scenarios, the previously perceived ceiling on training data is unlikely to pose short-term constraints.

Image

Continue reading with FUNDA

This report is available to subscribers. Sign in or subscribe to read the full analysis.