
K3 still needs HBM-scale nodes; smaller KV makes offload practical, boosting DRAM and NAND usage.
Over the weekend, a peer firm's Twitter commentary on Kimi K3 and its implications for memory demand set off a wave of discussion. We agree with the direction of their core call: Kimi K3 is not a lightweight model that "meaningfully reduces memory and networking demand," but rather a very large MoE with more efficient attention memory, heavier model weights, and wider Expert Parallelism. K3's total parameter count reaches 2.8T; it activates 16 of 896 experts per token and uses KDA, Gated MLA, Stable LatentMoE, and MXFP4 weights. Moonshot officially recommends deploying it on a high-bandwidth supernode with at least 64 accelerators. It therefore still needs plenty of HBM, scale-up networking, and GPUs. In terms of performance, it can be considered firmly in the top tier. However, Moonshot itself admits overall capability and user experience still trail the strongest closed-source models, so the hype should be kept in check. Meanwhile, as our earlier report noted, K3's token efficiency is relatively low, making it more expensive than GPT-5.6 Sol on real long-horizon tasks, with no clear cost-performance edge.
This report is available to subscribers. Sign in or subscribe to read the full analysis.