
CPO Is a Game Changer, While OIO Holds Greater Long-Term Promise
With OpenAI's release of the new O1 and O3 models, along with the diversified computational demands of the Orion model, the focus for computing resources has expanded beyond training to include post-training and inference. In particular, the post-training phase has shifted from the relatively lighter demands of SFT+RLHF to the more resource-intensive RL+CoT combined with synthetic data processing, requiring substantially greater computational power. This shift has significantly amplified the overall demand for computing resources.
The current application of Reinforcement Learning (RL) in model training has driven a sharp rise in computational demand. Unlike the previous generation of human-machine adversarial models, today’s AI-to-AI adversarial interactions have significantly escalated resource requirements. Post-training computational demands now surpass those of the pre-training phase, posing new challenges for existing computing solutions. Additionally, extended RL models, designed to push performance boundaries, demand even greater computational power. As a result, the need for computing resources continues to grow exponentially, with no clear upper limit in sight. Experimental setups frequently require tens of thousands of GPUs, placing unprecedented demands on cluster architecture and design.

Figure1: o1 performance smoothly improves with both train-timeand test-time compute
This report is available to subscribers. Sign in or subscribe to read the full analysis.